# asr

Automatic speech recognition is the task of transcribing audio into text.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Integrate Deepgram Flux with Twilio's Conversation Relay

DevFeed: [Integrate Deepgram Flux with Twilio's Conversation Relay](<https://devfeed.tech/articles/integrate-deepgram-flux-with-twilio-s-conversation-relay-16090.md>)

Original publisher: [Read original article](<https://www.twilio.com/en-us/blog/developers/tutorials/integrations/deepgram-flux-twilio-conversation-relay>)

Author: Dhruv Patel

Published: 2026-09-08T00:00:00Z

Content type: tutorial

Language: en

Sources: [Twilio Blog](<https://devfeed.tech/sources/twilio-blog.md>)

Topics: [flux](<https://devfeed.tech/topics/flux.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Tutorial](<https://devfeed.tech/topics/tutorial.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Node.js](<https://devfeed.tech/topics/node-js.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [developer-insights](<https://devfeed.tech/tags/developer-insights.md>), [flux](<https://devfeed.tech/tags/flux.md>), [latency](<https://devfeed.tech/tags/latency.md>), [node](<https://devfeed.tech/tags/node.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>), [websocket](<https://devfeed.tech/tags/websocket.md>)

### AI overview

This tutorial shows how to integrate Deepgram Flux with Twilio's Conversation Relay in Node.js. Flux combines speech transcription and turn detection, with configurable end-of-turn confidence, and is described as reducing response latency and false interruptions.

### Source excerpt

Integrate Deepgram Flux with Twilio ConversationRelay in Node.js for faster turn detectio, and tunable end-of-turn control.

## Home Assistant 2026.9 adds Modbus device integrations, Cloud voice processing, and Matter network mapping

DevFeed: [Home Assistant 2026.9 adds Modbus device integrations, Cloud voice processing, and Matter network mapping](<https://devfeed.tech/articles/2026-9-there-s-room-on-this-bus-16691.md>)

Original publisher: [Read original article](<https://www.home-assistant.io/blog/2026/09/02/release-20269/>)

Author: Franck Nijhof

Published: 2026-09-02T00:00:00Z

Content type: release

Language: en

Sources: [Home Assistant](<https://devfeed.tech/sources/home-assistant.md>)

Topics: [Home Assistant](<https://devfeed.tech/topics/home-assistant.md>), [Release notes](<https://devfeed.tech/topics/release-notes.md>), [Modbus](<https://devfeed.tech/topics/modbus.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Matter](<https://devfeed.tech/topics/matter.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [cloud](<https://devfeed.tech/tags/cloud.md>), [core](<https://devfeed.tech/tags/core.md>), [home-assistant](<https://devfeed.tech/tags/home-assistant.md>), [matter](<https://devfeed.tech/tags/matter.md>), [modbus](<https://devfeed.tech/tags/modbus.md>), [release](<https://devfeed.tech/tags/release.md>), [release-notes](<https://devfeed.tech/tags/release-notes.md>), [security](<https://devfeed.tech/tags/security.md>), [speech](<https://devfeed.tech/tags/speech.md>), [ui](<https://devfeed.tech/tags/ui.md>)

### AI overview

Home Assistant 2026.9 introduces Modbus integrations that can share a connection, new Cloud voice processing focused on speech-to-text, Security dashboard alerts and favorites, shared sources for History and Activity, a Matter network map, keyboard and sound navigation for charts, and 13 new integrations.

### Source excerpt

Home Assistant 2026.9! 🎉 I'll be honest: Modbus has never been the flashiest part of Home Assistant. It quietly powers solar inverters, heat pumps, and energy meters, but only if you're comfortable hand-writing your own register map in YAML, a real wall that's kept a huge category of devices out of reach. This release finally makes room on the bus: integrations that already know how a device talks over Modbus, sharing a single connection instead of fighting over it, so you simply pick your device in the UI like anything else in Home Assistant. It won't show up in a screenshot, but modernizing Modbus is hands-down my favorite change this release. 🔧 I'm also really looking forward to jumping on the new Cloud voice processing myself and putting it through its paces. We're a household with a non-English main language, and if you've ever run a private voice assistant in one, you know exactly how fast accents and background noise trip things up. This new speech-to-text engine specifically targets those pain points, so I'm very curious to see how it holds up against my own accent. 😉 And there's plenty more: active alerts and favorites for the Security dashboard, a shared Sources panel for History and Activity, Activity that finally explains why something changed, a brand new network map for Matter, charts you can now navigate by keyboard and sound, and 13 new integrations! 🚀 One more thing before you dive in: we're at IFA Berlin this weekend! It's one of the largest consumer electronics expos out there, running September 4 to 8, 2026, and the Open Home Foundation has its own booth in the Smart Home hall: Hall 1.2, stand 1.2-153. If you're around, come find us and say hi. 👋 Enjoy the release! ../Frenck Home Assistant Cloud Seeing through the Cloud Test the new voice processing Float in to be a Cloud subscriber Active alerts and favorites for the Security dashboard A new Sources pane for History and Activity From what changed to why it changed Insights into your Matter world

## The Open ASR Leaderboard Adds Its First Global South Language

DevFeed: [The Open ASR Leaderboard Adds Its First Global South Language](<https://devfeed.tech/articles/the-open-asr-leaderboard-adds-its-first-global-south-language-7411.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-global-south>)

Author: Eric Bezzam; Shobhit Banga; Manas Dhir; Bhaskar Singh; Manmeet Kaur; Aaditya Pareek; Walecha; Sagar Jain; Hanuman Sidh; Vanshika Chhabra

Published: 2026-08-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [contributors](<https://devfeed.tech/tags/contributors.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [devices](<https://devfeed.tech/tags/devices.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

The Open ASR Leaderboard introduces Monsoon evaluation sets for Hindi in India, expanding coverage beyond European languages and testing how recognition performance varies across populations and conditions. The sets use public and private splits, speaker-disjoint data, detailed speaker attributes, and variation in geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and transcript validity.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Intelligent transcription with Gemini 3.5 Transcribe

DevFeed: [Intelligent transcription with Gemini 3.5 Transcribe](<https://devfeed.tech/articles/intelligent-transcription-with-gemini-3-5-transcribe-6190.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/intelligent-transcription-with-gemini-3-5-transcribe/>)

Author: Diego Melendo Casado

Published: 2026-08-26T17:01:00Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [API](<https://devfeed.tech/topics/api.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [android](<https://devfeed.tech/tags/android.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developers](<https://devfeed.tech/tags/developers.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [google-ai](<https://devfeed.tech/tags/google-ai.md>), [latency](<https://devfeed.tech/tags/latency.md>), [macos](<https://devfeed.tech/tags/macos.md>), [model](<https://devfeed.tech/tags/model.md>), [none](<https://devfeed.tech/tags/none.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Google introduces Gemini 3.5 Transcribe, a speech-to-text model for accurate, formatted transcription in noisy and specialized audio. Developers can use it through the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform for voice agents, real-time captioning, and post-call analytics. It supports real-time streaming and pre-recorded audio processing, custom vocabulary, speaker attribution, word-level timestamps, and more than 85 languages.

### Source excerpt

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

## Measuring benchmark optimization in speech recognition

DevFeed: [Measuring benchmark optimization in speech recognition](<https://devfeed.tech/articles/measuring-benchmark-optimization-in-speech-recognition-7104.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asr-benchmark-optimization>)

Author: Theo Lebryk; Eric Bezzam; Alice; David Ayllon; Jakub Piotr Cłapa; Jens Madsen; Panagiotis Tzirakis

Published: 2026-08-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmark overfitting machine learning](<https://devfeed.tech/topics/benchmark-overfitting-machine-learning.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [errors](<https://devfeed.tech/tags/errors.md>), [measurement](<https://devfeed.tech/tags/measurement.md>), [model](<https://devfeed.tech/tags/model.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

The article examines benchmark optimization, or "benchmaxxing," in speech recognition. It presents three tests and evaluates 11 open-source ASR models, finding that some reproduced benchmark transcripts even when the audio contradicted them. The research also uses model ensembles and human annotations to identify and validate likely benchmark errors.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Fish Audio models now available on Vercel AI Gateway for free

DevFeed: [Fish Audio models now available on Vercel AI Gateway for free](<https://devfeed.tech/articles/fish-audio-models-now-available-on-vercel-ai-gateway-for-free-934.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/fish-audio-models-now-available-on-ai-gateway-for-free>)

Author: Jerilyn Zheng

Published: 2026-08-19T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [free](<https://devfeed.tech/tags/free.md>), [launch](<https://devfeed.tech/tags/launch.md>), [models](<https://devfeed.tech/tags/models.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Fish Audio's text-to-speech and transcription models are available on Vercel AI Gateway. The models are free for 30 days, with AI SDK 7 support for speech generation and transcription, including timestamped segments and word-level timing.

### Source excerpt

Fish Audio's audio models are now available on AI Gateway. To celebrate the launch, every Fish Audio model is free on AI Gateway for the next 30 days, through September 18. Capability Regular Through September 18 Text-to-speech $15.00 per million characters Free Speech-to-text $0.36 per hour of audio Free Four models from Fish Audio are available, including their latest text-to-speech model: fish-audio/s2.1-pro (text-to-speech): Built for low-latency streaming; clones a voice from a reference recording. fish-audio/transcribe-1 (transcription): Returns the text along with the duration of the audio and timestamped segments, down to individual words. fish-audio/s2-pro (text-to-speech): Covers around eighty languages and takes inline tags, plain-language directions written into the text itself, so you can change how a single word or phrase is delivered instead of setting one style for the whole request. fish-audio/s1 (text-to-speech): Reads text that can carry markers for emotion, tone, and sound effects. How to use models during the offer period Using the standard model name (i.e., fish-audio/s2.1-pro) is free, but will automatically begin billing when the offer period ends. To ensure you aren't billed after the free period, add the -free suffix to the standard name, and the model will stop serving when the offer ends (i.e., fish-audio/s2.1-pro-free). Speech and transcription ship in the current AI SDK 7 release. Text-to-speech Generate spoken audio from text with generateSpeech and write the result: Speech-to-text Transcribe recordings into text with transcribe. The audio can be a buffer, a base64 string, or a URL: Each segment carries the text and its start and end time in seconds, down to individual words. Playground You can also try the Fish Audio models without writing any code. Open the models list, click into a model, and send text or audio to hear or read the result in your browser. For a full overview of how to utilize audio models, refer to the speech quickst

## Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS

DevFeed: [Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS](<https://devfeed.tech/articles/build-low-latency-multilingual-voice-agents-open-weights-full-deployment-control-with-nvidia-magpie-tts-7386.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents>)

Author: Maryam Motamedi; Mikyas Desta; Jason Li; Jason Roche

Published: 2026-08-10T16:25:36Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [NVIDIA NIM](<https://devfeed.tech/topics/nvidia-nim.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [automation](<https://devfeed.tech/tags/automation.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [data](<https://devfeed.tech/tags/data.md>), [developers](<https://devfeed.tech/tags/developers.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-nim](<https://devfeed.tech/tags/nvidia-nim.md>), [open](<https://devfeed.tech/tags/open.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

This developer article presents NVIDIA Magpie Multilingual TTS as an open-weights, 364M-parameter text-to-speech model for building low-latency multilingual voice applications. It explains how a self-managed cascaded ASR, TTS, and LLM architecture can provide deployment control, tuning, privacy, and predictable performance across 12 languages.

### Source excerpt

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS Every voice interaction has a latency budget. By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retrieving context, and generating a response. Text-to-speech (TTS) is the final step -- and the one users notice most. If speech generation is slow, the whole experience feels slow.

## AI Model Risk Intelligence Know Which Models You Can Trust Before You Deploy

DevFeed: [AI Model Risk Intelligence Know Which Models You Can Trust Before You Deploy](<https://devfeed.tech/articles/ai-model-risk-intelligence-know-which-models-you-can-trust-before-you-deploy-8253.md>)

Original publisher: [Read original article](<https://snyk.io/blog/why-we-rebuilt-evo-ai-model-risk-scoring/>)

Author: Ranko Cupovic

Published: 2026-08-04T04:00:00Z

Content type: article

Language: en

Sources: [Blog RSS Feed | Snyk](<https://devfeed.tech/sources/blog-rss-feed-snyk.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Adversarial attacks](<https://devfeed.tech/topics/adversarial-attacks.md>), [Security](<https://devfeed.tech/topics/security.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [pii](<https://devfeed.tech/topics/pii.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [application-security](<https://devfeed.tech/tags/application-security.md>), [asr](<https://devfeed.tech/tags/asr.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [awareness](<https://devfeed.tech/tags/awareness.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [blog](<https://devfeed.tech/tags/blog.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [developer](<https://devfeed.tech/tags/developer.md>), [devops](<https://devfeed.tech/tags/devops.md>), [enablement](<https://devfeed.tech/tags/enablement.md>), [model](<https://devfeed.tech/tags/model.md>), [pii](<https://devfeed.tech/tags/pii.md>), [safety](<https://devfeed.tech/tags/safety.md>), [security](<https://devfeed.tech/tags/security.md>), [snyk-platform](<https://devfeed.tech/tags/snyk-platform.md>), [snyk-security-intel](<https://devfeed.tech/tags/snyk-security-intel.md>)

### AI overview

Evo's AI model risk score combines attack success rate and attack impact into a 0-1000 score. It is based on adversarial testing against standard system-prompt hardening and breaks risk down by attacker goals such as PII extraction, system-prompt extraction, and insecure code generation to support deployment decisions.

### Source excerpt

AI model risk depends on how a model is deployed. Learn how Evo combines adversarial testing, attack impact, and deployment context to help teams compare models and enforce policy.

## OpenAI sells a $230 keypad to run AI agents. I built one for $28 on Linux.

DevFeed: [OpenAI sells a $230 keypad to run AI agents. I built one for $28 on Linux.](<https://devfeed.tech/articles/openai-sells-a-230-keypad-to-run-ai-agents-i-built-one-for-28-on-linux-25179.md>)

Original publisher: [Read original article](<https://www.ivanmorgillo.com/2026/07/29/openai-codex-micro-28-dollar-linux-macropad/>)

Author: Ivan Morgillo

Published: 2026-07-29T10:00:00Z

Content type: tutorial

Language: en

Sources: [Ivan Morgillo](<https://devfeed.tech/sources/ivan-morgillo.md>)

Topics: [codex](<https://devfeed.tech/topics/codex.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>)

Tags: [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [ai-coding-agents](<https://devfeed.tech/tags/ai-coding-agents.md>), [asr](<https://devfeed.tech/tags/asr.md>), [codex](<https://devfeed.tech/tags/codex.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [linux](<https://devfeed.tech/tags/linux.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [openai](<https://devfeed.tech/tags/openai.md>)

### AI overview

The article explains how the author recreated the core workflow of OpenAI's Codex Micro for $28 on Linux using a low-cost macropad. The setup supports voice dictation through a local speech-to-text server and one-press commands for interacting with coding agents.

### Source excerpt

OpenAI's first hardware is a $230 keypad for driving AI coding agents. I rebuilt the part that matters -- voice dictation and one-press agent commands -- for $28, on Linux, with a cheap Chinese macropad.

## Adding new quick commands to a smart speaker without degrading existing commands

DevFeed: [Adding new quick commands to a smart speaker without degrading existing commands](<https://devfeed.tech/articles/article-24871.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1061968/>)

Author: khaymon (Яндекс)

Published: 2026-07-23T09:00:04Z

Content type: tutorial

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [яндекс](<https://devfeed.tech/topics/tag-4004cf5948d3.md>), [asr](<https://devfeed.tech/topics/asr.md>), [cpu](<https://devfeed.tech/topics/cpu.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [continual-learning](<https://devfeed.tech/tags/continual-learning.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [keyword-spotting](<https://devfeed.tech/tags/keyword-spotting.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [speech-processing](<https://devfeed.tech/tags/speech-processing.md>), [tag-355bb785df82](<https://devfeed.tech/tags/tag-355bb785df82.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>), [tag-8483d32db3b1](<https://devfeed.tech/tags/tag-8483d32db3b1.md>), [tag-967d8467ce56](<https://devfeed.tech/tags/tag-967d8467ce56.md>), [tag-9c1e53a23032](<https://devfeed.tech/tags/tag-9c1e53a23032.md>), [tag-a1312fd2c7ff](<https://devfeed.tech/tags/tag-a1312fd2c7ff.md>), [tag-d14eb265d33e](<https://devfeed.tech/tags/tag-d14eb265d33e.md>), [tag-d346fb5ae499](<https://devfeed.tech/tags/tag-d346fb5ae499.md>)

### AI overview

The article discusses Yandex's compact on-device neural model for recognizing Alice quick commands. It covers adding track-rating and Bluetooth commands while aiming to preserve performance on existing commands and limit device resource use.

### Source excerpt

Чтобы дать команду умной колонке, не обязательно говорить активационное слово "Алиса": есть быстрые команды -- короткие фразы, с помощью которых можно управлять музыкой, громкостью или умным домом. Например, чтобы переключить трек, достаточно просто сказать "дальше", а чтобы убавить звук -- "тише". Весь список команд можно посмотреть в настройках вашего аккаунта в приложении "Дом с Алисой". Быстрые команды удобнее не только пользователям, но и системе: запросы через слово "Алиса" требуют обращения к модели распознавания речи ASR, которой из-за её размеров необходимы серверные вычислительные ресурсы, а модель быстрых команд устроена гораздо компактнее. Она работает прямо на устройстве, а значит, ограничена вычислительными ресурсами самой колонки -- её CPU и оперативной памятью. Из-за этого модель нельзя сильно увеличить: ей приходится оставаться компактной, зато запрос обрабатывается быстрее. За распознавание быстрых команд отвечает нейросеть. Её архитектура почти полностью совпадает с решением для наушников Яндекс Дропс, которое подробно описал в своей статье Григорий Афанасенко. Разница в основном в масштабе: наша модель весит всего от 0,5 до 1,5 МБ в зависимости от железа конкретного устройства. Со временем перед нами встала задача добавить к базовым командам "лайк" и "дизлайк" для управления треками, а также команды "включи блютус" и "выключи блютус". Особенно это актуально для Станции Стрит, которую часто берут с собой на природу, где нет интернета. Но главным было гарантировать абсолютное отсутствие ухудшения на уже запущенных командах и не слишком сильно увеличивать потребление ресурсов на устройстве. Читать далее

## AI Gateway now supports streaming transcription

DevFeed: [AI Gateway now supports streaming transcription](<https://devfeed.tech/articles/ai-gateway-now-supports-streaming-transcription-802.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/ai-gateway-now-supports-streaming-transcription>)

Author: Jerilyn Zheng

Published: 2026-07-22T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Streaming](<https://devfeed.tech/topics/streaming.md>), [asr](<https://devfeed.tech/topics/asr.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [latency](<https://devfeed.tech/tags/latency.md>), [openai](<https://devfeed.tech/tags/openai.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [streams](<https://devfeed.tech/tags/streams.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [voice](<https://devfeed.tech/tags/voice.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

AI Gateway now supports beta streaming transcription through the AI SDK's streamTranscribe function. Applications can stream audio as it is captured and receive partial and final transcript updates with low latency, enabling use cases such as live captioning, voice input, and voice-enabled agents.

### Source excerpt

AI Gateway now supports streaming transcription. Previously, transcription required a complete audio file and returned the full transcript in a single response. Now you can stream audio in as it's captured and get transcript updates back as the model produces them, keeping latency low for uses like live captioning and voice input. Streaming transcription is in beta and available through the AI SDK's streamTranscribe function with any streaming-capable transcription model. The example below streams raw PCM audio to openai/gpt-realtime-whisper and prints each transcript delta as it arrives. The result stream also carries partial and final transcripts: The same code works across providers: swap the model string to use xai/grok-stt or any other streaming-capable transcription model. Streaming transcription also makes it easy to add a voice mode to an agent. Stream the user's speech to a transcription model and pass the live text to your agent. The agent itself does not change: it still receives text, so this works with any text-based agent. For agents that speak back, pair it with speech generation, or use realtime voice for full two-way conversation. For more information on streaming transcription on AI Gateway, see the documentation. Read more

## Introducing Real World VoiceEQ: Measuring the human quality of voice AI

DevFeed: [Introducing Real World VoiceEQ: Measuring the human quality of voice AI](<https://devfeed.tech/articles/introducing-real-world-voiceeq-measuring-the-human-quality-of-voice-ai-7454.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/real-world-voiceeq>)

Author: David Ayllon; Alice; Jeff Brooks; Franc Camps Febrer; Jakub Piotr Cłapa; Theo Lebryk; Jens Madsen; Olya Ossipova; Sharath Rao; Hoon Shin

Published: 2026-07-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [asr](<https://devfeed.tech/topics/asr.md>), [human feedback](<https://devfeed.tech/topics/human-feedback.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [human-feedback](<https://devfeed.tech/tags/human-feedback.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [speech](<https://devfeed.tech/tags/speech.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>)

### AI overview

Real World VoiceEQ is a benchmark for evaluating the human quality of voice AI beyond latency and word error rate. It measures how voice systems recognize, produce, and respond to acoustic information such as tone, emotion, speaker identity, and background context across ASR, TTS, speech-to-speech, and speech understanding. The benchmark covers more than 40 voice models, 15+ evaluation dimensions, and more than 60 metrics, using over 1 million human ratings.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Meeting Notes - A Free Desktop App for Tracking 1:1s

DevFeed: [Meeting Notes - A Free Desktop App for Tracking 1:1s](<https://devfeed.tech/articles/meeting-notes-a-free-desktop-app-for-tracking-1-1s-37543.md>)

Original publisher: [Read original article](<https://deanhume.com/meeting-notes-desktop-app-windows/>)

Author: Dean Hume

Published: 2026-06-29T10:16:48Z

Content type: article

Language: en

Sources: [Dean Hume](<https://devfeed.tech/sources/dean-hume.md>)

Topics: [App](<https://devfeed.tech/topics/app.md>), [meetings](<https://devfeed.tech/topics/meetings.md>), [Local-First](<https://devfeed.tech/topics/local-first.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Markdown](<https://devfeed.tech/topics/markdown.md>)

Tags: [app](<https://devfeed.tech/tags/app.md>), [autosave](<https://devfeed.tech/tags/autosave.md>), [desktop](<https://devfeed.tech/tags/desktop.md>), [free](<https://devfeed.tech/tags/free.md>), [meetings](<https://devfeed.tech/tags/meetings.md>), [notes](<https://devfeed.tech/tags/notes.md>), [offline](<https://devfeed.tech/tags/offline.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [recording](<https://devfeed.tech/tags/recording.md>), [speech-to-text](<https://devfeed.tech/tags/speech-to-text.md>), [technical-leadership](<https://devfeed.tech/tags/technical-leadership.md>), [technical-program-manager](<https://devfeed.tech/tags/technical-program-manager.md>), [whisper](<https://devfeed.tech/tags/whisper.md>), [writing](<https://devfeed.tech/tags/writing.md>)

### AI overview

The article introduces Meeting Notes, a free, open-source Windows desktop app for organizing 1:1 notes by person. It supports topic tags, Markdown, autosave, voice recording, and offline speech-to-text using Whisper, with data and audio kept on the device.

### Source excerpt

Meeting Notes is a free offline desktop app for 1:1s. Organise notes by person, tag topics, and transcribe meetings on-device with Whisper.

## Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World

DevFeed: [Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World](<https://devfeed.tech/articles/introducing-the-ffasr-leaderboard-benchmarking-asr-in-the-real-world-7196.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ffasr-leaderboard>)

Author: Daniel Gert Nielsen; Shivam Saini; Alessia Milo; Georg Götz; Eric Bezzam

Published: 2026-06-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [community](<https://devfeed.tech/tags/community.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [development](<https://devfeed.tech/tags/development.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

Treble Technologies and Hugging Face introduce the open, community-driven FFASR Leaderboard, a benchmark for evaluating automatic speech recognition models in realistic far-field acoustic conditions. It measures performance across factors such as reverberation, background noise, microphone distance, and low signal-to-noise ratios, while also showing the tradeoff between recognition accuracy and speed.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Advancing voice intelligence with new models in the API

DevFeed: [Advancing voice intelligence with new models in the API](<https://devfeed.tech/articles/advancing-voice-intelligence-with-new-models-in-the-api-6280.md>)

Original publisher: [Read original article](<https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api>)

Published: 2026-05-07T10:00:00Z

Content type: release

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [developers](<https://devfeed.tech/tags/developers.md>), [launch](<https://devfeed.tech/tags/launch.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [product](<https://devfeed.tech/tags/product.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [tools](<https://devfeed.tech/tags/tools.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

OpenAI is introducing three audio models in its API: GPT-Realtime-2 for more capable conversational voice interactions, GPT-Realtime-Translate for live speech translation, and GPT-Realtime-Whisper for streaming speech-to-text. The models are designed to support voice applications that can listen, reason, translate, transcribe, use tools, and take action in real time.

### Source excerpt

Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.

## Adding Benchmaxxer Repellant to the Open ASR Leaderboard

DevFeed: [Adding Benchmaxxer Repellant to the Open ASR Leaderboard](<https://devfeed.tech/articles/adding-benchmaxxer-repellant-to-the-open-asr-leaderboard-7413.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-private-data>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Sergio Bruccoleri; Jeanine Sinanan-Singh; Casey Ford; Guanbo Wang; Yukai Huang; Ke Li; Yufeng Hao

Published: 2026-05-06T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The article announces private English ASR datasets from Appen and DataoceanAI for the Open ASR Leaderboard. Keeping the datasets private is intended to reduce benchmark-specific optimization and test-set contamination while preserving a high-quality evaluation of multiple speech-recognition tasks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

DevFeed: [Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents](<https://devfeed.tech/articles/introducing-nvidia-nemotron-3-nano-omni-long-context-multimodal-intelligence-for-documents-audio-and-video-agents-7395.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence>)

Author: Tuomas Rintamaki; Amala Sanjay Deshmukh; Nabin Mulepati; Collin McCarthy; Pritam Biswas; Arushi Goel; Alexandre Milesi; Danial Mohseni Taheri; Kateryna Chumachenko; Isabel Hulseman; Zhehuai Chen; Kara

Published: 2026-04-28T15:58:57Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [asr](<https://devfeed.tech/topics/asr.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [Mamba](<https://devfeed.tech/topics/mamba.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [moe](<https://devfeed.tech/topics/moe.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>)

Tags: [alternatives](<https://devfeed.tech/tags/alternatives.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [speed](<https://devfeed.tech/tags/speed.md>)

### AI overview

NVIDIA introduces Nemotron 3 Nano Omni, an omni-modal model for document analysis, image reasoning, speech recognition, long audio-video understanding, computer use, and general reasoning. It combines a hybrid Mamba-Transformer Mixture-of-Experts backbone with vision and audio encoders, supports long multimodal contexts, and reports strong benchmark accuracy, throughput, reasoning speed, and system efficiency.

### Source excerpt

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents - NVIDIA Nemotron 3 Nano Omni is a new omni-modal understanding model built for real-world document analysis, multiple image reasoning, automatic speech recognition, long audio-video understanding, agentic computer use, and general reasoning. - It extends the Nemotron multimodal line from a strong vision-language system to a broader text + image + video + audio model.

## WAXAL: A large-scale open resource for African language speech technology

DevFeed: [WAXAL: A large-scale open resource for African language speech technology](<https://devfeed.tech/articles/waxal-a-large-scale-open-resource-for-african-language-speech-technology-6927.md>)

Original publisher: [Read original article](<https://research.google/blog/waxal-a-large-scale-open-resource-for-african-language-speech-technology/>)

Published: 2026-03-06T20:06:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [data](<https://devfeed.tech/topics/data.md>), [Google](<https://devfeed.tech/topics/google.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [africa](<https://devfeed.tech/tags/africa.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [google](<https://devfeed.tech/tags/google.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [research](<https://devfeed.tech/tags/research.md>), [resources](<https://devfeed.tech/tags/resources.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google Research introduces WAXAL, an open-access speech dataset covering 27 Sub-Saharan African languages. The release includes approximately 1,846 hours of transcribed ASR data and more than 565 hours of high-fidelity TTS recordings under a CC-BY-4.0 license.

### Source excerpt

Natural Language Processing

## HarborFM - Self-Hosted Podcast Creator

DevFeed: [HarborFM - Self-Hosted Podcast Creator](<https://devfeed.tech/articles/harborfm-self-hosted-podcast-creator-10731.md>)

Original publisher: [Read original article](<https://noted.lol/harborfm/>)

Author: Logan Rickert

Published: 2026-02-20T18:37:34Z

Content type: article

Language: en

Sources: [Noted](<https://devfeed.tech/sources/noted.md>)

Topics: [Open Source](<https://devfeed.tech/topics/open-source.md>), [App](<https://devfeed.tech/topics/app.md>), [RSS Feed](<https://devfeed.tech/topics/rss-feed.md>), [PWA](<https://devfeed.tech/topics/pwa.md>), [WebRTC](<https://devfeed.tech/topics/webrtc.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [podcast](<https://devfeed.tech/tags/podcast.md>), [rss](<https://devfeed.tech/tags/rss.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [self-hosted-media-streaming-audio-streaming](<https://devfeed.tech/tags/self-hosted-media-streaming-audio-streaming.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [vision](<https://devfeed.tech/tags/vision.md>), [webrtc](<https://devfeed.tech/tags/webrtc.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

HarborFM is presented as a free, open-source, self-hosted podcast creation tool. It lets users assemble episodes from reusable audio segments, publish per-show RSS feeds, record remote guests through WebRTC, collaborate on episodes, manage private feeds, and view analytics. It also supports podcast chapters, Whisper-based transcription, and AI summaries.

### Source excerpt

Take back your podcast data! HarborFM is a free and open source platform to create and host your podcasts.

## Voxtral transcribes at the speed of sound.

DevFeed: [Voxtral transcribes at the speed of sound.](<https://devfeed.tech/articles/voxtral-transcribes-at-the-speed-of-sound-7134.md>)

Original publisher: [Read original article](<https://mistral.ai/news/voxtral-transcribe-2/>)

Published: 2026-02-04T16:00:00Z

Content type: article

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Security](<https://devfeed.tech/topics/security.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [apache](<https://devfeed.tech/tags/apache.md>), [arabic](<https://devfeed.tech/tags/arabic.md>), [audio](<https://devfeed.tech/tags/audio.md>), [batch](<https://devfeed.tech/tags/batch.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cost](<https://devfeed.tech/tags/cost.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [performance](<https://devfeed.tech/tags/performance.md>), [security](<https://devfeed.tech/tags/security.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Mistral introduces Voxtral Transcribe 2, a family of speech-to-text models comprising Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications. The release highlights speaker diarization, word-level timestamps, multilingual transcription in 13 languages, configurable sub-200 ms latency, streaming transcription, and open Apache 2.0 weights for edge deployment.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR

DevFeed: [Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR](<https://devfeed.tech/articles/next-generation-medical-image-interpretation-with-medgemma-1-5-and-medical-speech-to-text-with-medasr-6842.md>)

Original publisher: [Read original article](<https://research.google/blog/next-generation-medical-image-interpretation-with-medgemma-15-and-medical-speech-to-text-with-medasr/>)

Published: 2026-01-13T20:57:16Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Medical imaging](<https://devfeed.tech/topics/medical-imaging.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kaggle](<https://devfeed.tech/topics/kaggle.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [asr](<https://devfeed.tech/tags/asr.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [health](<https://devfeed.tech/tags/health.md>), [health-bioscience](<https://devfeed.tech/tags/health-bioscience.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [kaggle](<https://devfeed.tech/tags/kaggle.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [medical-imaging](<https://devfeed.tech/tags/medical-imaging.md>)

### AI overview

Google Research describes MedGemma 1.5 4B, an updated open medical generative AI model with improved support for medical imaging, text, medical records, and 2D images. The article also presents MedASR, an open medical speech-to-text model for dictation that can pair with MedGemma for advanced reasoning. The models are available for research and commercial use through Hugging Face and Vertex AI, with a related medical AI hackathon on Kaggle.

### Source excerpt

Generative AI

## Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks

DevFeed: [Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks](<https://devfeed.tech/articles/open-asr-leaderboard-trends-and-insights-with-new-multilingual-long-form-tracks-7409.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Vaibhav Srivastav

Published: 2025-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [trends](<https://devfeed.tech/tags/trends.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The Open ASR Leaderboard expands its evaluation with multilingual and long-form transcription tracks. The article highlights accuracy advantages from combining Conformer encoders with LLM decoders, throughput advantages from CTC and TDT decoders, Whisper as a multilingual baseline, and the effects of fine-tuning on specialized performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Real-time speech-to-speech translation

DevFeed: [Real-time speech-to-speech translation](<https://devfeed.tech/articles/real-time-speech-to-speech-translation-6852.md>)

Original publisher: [Read original article](<https://research.google/blog/real-time-speech-to-speech-translation/>)

Published: 2025-11-19T09:59:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [asr](<https://devfeed.tech/tags/asr.md>), [data](<https://devfeed.tech/tags/data.md>), [google](<https://devfeed.tech/tags/google.md>), [ml](<https://devfeed.tech/tags/ml.md>), [personalization](<https://devfeed.tech/tags/personalization.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Google researchers introduce an end-to-end speech-to-speech translation model that translates speech in real time while preserving the original speaker's voice, with a two-second delay. The system uses a streaming architecture, time-synchronized training data, and a scalable data-acquisition pipeline to support more languages and improve conversational naturalness.

### Source excerpt

Algorithms & Theory

## Voice Cloning with Consent

DevFeed: [Voice Cloning with Consent](<https://devfeed.tech/articles/voice-cloning-with-consent-7562.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/voice-consent-gate>)

Author: Margaret Mitchell; Lucie-Aimée Kaffee

Published: 2025-10-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [ethics](<https://devfeed.tech/tags/ethics.md>), [guide](<https://devfeed.tech/tags/guide.md>), [speech](<https://devfeed.tech/tags/speech.md>), [systems](<https://devfeed.tech/tags/systems.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-cloning](<https://devfeed.tech/tags/voice-cloning.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

This article presents a voice consent gate for voice cloning. The system requires a speaker to generate and speak a consent phrase, uses automatic speech recognition to verify it, and only then permits a text-to-speech voice-cloning model to run. The design makes consent a traceable and auditable prerequisite for AI action while addressing both the risks and beneficial uses of voice cloning.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

[Next page](<https://devfeed.tech/topics/asr.md?cursor=WyIyMDI1LTEwLTI4VDAwOjAwOjAwKzAwOjAwIiwgIjMwNjIyZWUwLTVlOWYtNGIzNi05MDY0LTNlMWQ4YzA2NzkyNSJd>)