# transcription

Published articles for transcription.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Integrate Deepgram Flux with Twilio's Conversation Relay

DevFeed: [Integrate Deepgram Flux with Twilio's Conversation Relay](<https://devfeed.tech/articles/integrate-deepgram-flux-with-twilio-s-conversation-relay-16090.md>)

Original publisher: [Read original article](<https://www.twilio.com/en-us/blog/developers/tutorials/integrations/deepgram-flux-twilio-conversation-relay>)

Author: Dhruv Patel

Published: 2026-09-08T00:00:00Z

Content type: tutorial

Language: en

Sources: [Twilio Blog](<https://devfeed.tech/sources/twilio-blog.md>)

Topics: [flux](<https://devfeed.tech/topics/flux.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Tutorial](<https://devfeed.tech/topics/tutorial.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Node.js](<https://devfeed.tech/topics/node-js.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [developer-insights](<https://devfeed.tech/tags/developer-insights.md>), [flux](<https://devfeed.tech/tags/flux.md>), [latency](<https://devfeed.tech/tags/latency.md>), [node](<https://devfeed.tech/tags/node.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>), [websocket](<https://devfeed.tech/tags/websocket.md>)

### AI overview

This tutorial shows how to integrate Deepgram Flux with Twilio's Conversation Relay in Node.js. Flux combines speech transcription and turn detection, with configurable end-of-turn confidence, and is described as reducing response latency and false interruptions.

### Source excerpt

Integrate Deepgram Flux with Twilio ConversationRelay in Node.js for faster turn detectio, and tunable end-of-turn control.

## The Four Vendor Relationships Commonly Involved in a Production Voice Feature

DevFeed: [The Four Vendor Relationships Commonly Involved in a Production Voice Feature](<https://devfeed.tech/articles/how-assemblyai-collapsed-four-vendors-into-a-single-api-key-17932.md>)

Original publisher: [Read original article](<https://read.bytesizeddesign.com/p/how-assemblyai-collapsed-needing-four-vendors-to-summarize-a-phone-call>)

Author: Byte-Sized Design

Published: 2026-08-28T21:47:43Z

Content type: article

Language: en

Sources: [Byte-Sized Design](<https://devfeed.tech/sources/byte-sized-design.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [pii](<https://devfeed.tech/topics/pii.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [api](<https://devfeed.tech/tags/api.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [llm](<https://devfeed.tech/tags/llm.md>), [pii](<https://devfeed.tech/tags/pii.md>), [pii-redaction](<https://devfeed.tech/tags/pii-redaction.md>), [production](<https://devfeed.tech/tags/production.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

The article describes how a production voice feature typically involves four vendor relationships: transcription, PII redaction, an LLM provider, and compliance review.

### Source excerpt

TLDR A production voice feature in 2026 typically ships with four vendor relationships: a transcription API, a PII redaction step somebody built in a sprint, an LLM provider, and a compliance review stretched across all of it.

## Intelligent transcription with Gemini 3.5 Transcribe

DevFeed: [Intelligent transcription with Gemini 3.5 Transcribe](<https://devfeed.tech/articles/intelligent-transcription-with-gemini-3-5-transcribe-6190.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/intelligent-transcription-with-gemini-3-5-transcribe/>)

Author: Diego Melendo Casado

Published: 2026-08-26T17:01:00Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [API](<https://devfeed.tech/topics/api.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [android](<https://devfeed.tech/tags/android.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developers](<https://devfeed.tech/tags/developers.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [google-ai](<https://devfeed.tech/tags/google-ai.md>), [latency](<https://devfeed.tech/tags/latency.md>), [macos](<https://devfeed.tech/tags/macos.md>), [model](<https://devfeed.tech/tags/model.md>), [none](<https://devfeed.tech/tags/none.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Google introduces Gemini 3.5 Transcribe, a speech-to-text model for accurate, formatted transcription in noisy and specialized audio. Developers can use it through the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform for voice agents, real-time captioning, and post-call analytics. It supports real-time streaming and pre-recorded audio processing, custom vocabulary, speaker attribution, word-level timestamps, and more than 85 languages.

### Source excerpt

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

## Gemini 3.5 Transcribe now available on AI Gateway

DevFeed: [Gemini 3.5 Transcribe now available on AI Gateway](<https://devfeed.tech/articles/gemini-3-5-transcribe-now-available-on-ai-gateway-944.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/gemini-3-5-transcribe-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-08-26T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Vercel announces that Google's Gemini 3.5 Transcribe models are available through AI Gateway for both complete recordings and live audio. The models support multilingual detection across more than 85 languages, custom vocabulary, speaker identification, word-level timestamps, and streaming transcription through AI SDK 7.

### Source excerpt

Gemini 3.5 Transcribe from Google is now available on AI Gateway for recorded and live audio: google/gemini-3.5-transcribe transcribes a complete audio file in one request. google/gemini-3.5-transcribe-live transcribes audio over a WebSocket and returns text as the audio arrives. Both models automatically detect more than 85 languages, including when a speaker switches languages. You can also provide custom vocabulary to improve the transcription of names, technical terms, and uncommon spellings. The model for complete recordings can also identify speakers and return word-level timestamps. Streaming transcription is available in AI SDK 7. Install the latest AI SDK and AI Gateway provider: Transcribe live audio Use streamTranscribe with a ReadableStream of raw audio chunks. Set inputAudioFormat to match the audio being sent: Transcribe a complete recording Use transcribe to send a complete audio file and receive the finished transcript: Try Gemini 3.5 Transcribe Live in the model playground, browse all transcription models, or read the speech quickstart. Read more

## Measuring benchmark optimization in speech recognition

DevFeed: [Measuring benchmark optimization in speech recognition](<https://devfeed.tech/articles/measuring-benchmark-optimization-in-speech-recognition-7104.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asr-benchmark-optimization>)

Author: Theo Lebryk; Eric Bezzam; Alice; David Ayllon; Jakub Piotr Cłapa; Jens Madsen; Panagiotis Tzirakis

Published: 2026-08-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmark overfitting machine learning](<https://devfeed.tech/topics/benchmark-overfitting-machine-learning.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [errors](<https://devfeed.tech/tags/errors.md>), [measurement](<https://devfeed.tech/tags/measurement.md>), [model](<https://devfeed.tech/tags/model.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

The article examines benchmark optimization, or "benchmaxxing," in speech recognition. It presents three tests and evaluates 11 open-source ASR models, finding that some reproduced benchmark transcripts even when the audio contradicted them. The research also uses model ensembles and human annotations to identify and validate likely benchmark errors.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Fish Audio models now available on Vercel AI Gateway for free

DevFeed: [Fish Audio models now available on Vercel AI Gateway for free](<https://devfeed.tech/articles/fish-audio-models-now-available-on-vercel-ai-gateway-for-free-934.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/fish-audio-models-now-available-on-ai-gateway-for-free>)

Author: Jerilyn Zheng

Published: 2026-08-19T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [free](<https://devfeed.tech/tags/free.md>), [launch](<https://devfeed.tech/tags/launch.md>), [models](<https://devfeed.tech/tags/models.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Fish Audio's text-to-speech and transcription models are available on Vercel AI Gateway. The models are free for 30 days, with AI SDK 7 support for speech generation and transcription, including timestamped segments and word-level timing.

### Source excerpt

Fish Audio's audio models are now available on AI Gateway. To celebrate the launch, every Fish Audio model is free on AI Gateway for the next 30 days, through September 18. Capability Regular Through September 18 Text-to-speech $15.00 per million characters Free Speech-to-text $0.36 per hour of audio Free Four models from Fish Audio are available, including their latest text-to-speech model: fish-audio/s2.1-pro (text-to-speech): Built for low-latency streaming; clones a voice from a reference recording. fish-audio/transcribe-1 (transcription): Returns the text along with the duration of the audio and timestamped segments, down to individual words. fish-audio/s2-pro (text-to-speech): Covers around eighty languages and takes inline tags, plain-language directions written into the text itself, so you can change how a single word or phrase is delivered instead of setting one style for the whole request. fish-audio/s1 (text-to-speech): Reads text that can carry markers for emotion, tone, and sound effects. How to use models during the offer period Using the standard model name (i.e., fish-audio/s2.1-pro) is free, but will automatically begin billing when the offer period ends. To ensure you aren't billed after the free period, add the -free suffix to the standard name, and the model will stop serving when the offer ends (i.e., fish-audio/s2.1-pro-free). Speech and transcription ship in the current AI SDK 7 release. Text-to-speech Generate spoken audio from text with generateSpeech and write the result: Speech-to-text Transcribe recordings into text with transcribe. The audio can be a buffer, a base64 string, or a URL: Each segment carries the text and its start and end time in seconds, down to individual words. Playground You can also try the Fish Audio models without writing any code. Open the models list, click into a model, and send text or audio to hear or read the result in your browser. For a full overview of how to utilize audio models, refer to the speech quickst

## Grok Voice Think Fast 2.0 now available on AI Gateway

DevFeed: [Grok Voice Think Fast 2.0 now available on AI Gateway](<https://devfeed.tech/articles/grok-voice-think-fast-2-0-now-available-on-ai-gateway-977.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/grok-voice-think-fast-2-0-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-29T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [API](<https://devfeed.tech/topics/api.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [latency](<https://devfeed.tech/tags/latency.md>), [model](<https://devfeed.tech/tags/model.md>), [playground](<https://devfeed.tech/tags/playground.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [server](<https://devfeed.tech/tags/server.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tool](<https://devfeed.tech/tags/tool.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Grok Voice Think Fast 2.0 from xAI is now available through AI Gateway as a speech-to-speech model. It processes audio input and output while reasoning in parallel, using fewer reasoning tokens to reduce tool-call delays. The article also describes transcription performance in noisy and compressed audio conditions and access through the AI SDK realtime API.

### Source excerpt

Grok Voice Think Fast 2.0 from xAI is now available on AI Gateway. It is a speech-to-speech voice model that takes audio in and audio out, improving on the previous Grok Voice model in reasoning, transcription accuracy, and conversation. The model reasons in parallel with speech, so it can think through a query while talking without adding latency. It has also been trained to use fewer reasoning tokens than before, so tool calls fire sooner, often before the end of the agent's first sentence. Transcription holds up in real-world conditions, including background noise and telephony compression. Use xai/grok-voice-think-fast-2.0 through the AI SDK's realtime API. Mint a short-lived token on the server so your API key never reaches the client: See the realtime documentation to build a voice agent with Grok Voice Think Fast 2.0. Try out the model in the AI Gateway playground. Read more

## AI Gateway now supports streaming transcription

DevFeed: [AI Gateway now supports streaming transcription](<https://devfeed.tech/articles/ai-gateway-now-supports-streaming-transcription-802.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/ai-gateway-now-supports-streaming-transcription>)

Author: Jerilyn Zheng

Published: 2026-07-22T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Streaming](<https://devfeed.tech/topics/streaming.md>), [asr](<https://devfeed.tech/topics/asr.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [latency](<https://devfeed.tech/tags/latency.md>), [openai](<https://devfeed.tech/tags/openai.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [streams](<https://devfeed.tech/tags/streams.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [voice](<https://devfeed.tech/tags/voice.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

AI Gateway now supports beta streaming transcription through the AI SDK's streamTranscribe function. Applications can stream audio as it is captured and receive partial and final transcript updates with low latency, enabling use cases such as live captioning, voice input, and voice-enabled agents.

### Source excerpt

AI Gateway now supports streaming transcription. Previously, transcription required a complete audio file and returned the full transcript in a single response. Now you can stream audio in as it's captured and get transcript updates back as the model produces them, keeping latency low for uses like live captioning and voice input. Streaming transcription is in beta and available through the AI SDK's streamTranscribe function with any streaming-capable transcription model. The example below streams raw PCM audio to openai/gpt-realtime-whisper and prints each transcript delta as it arrives. The result stream also carries partial and final transcripts: The same code works across providers: swap the model string to use xai/grok-stt or any other streaming-capable transcription model. Streaming transcription also makes it easy to add a voice mode to an agent. Stream the user's speech to a transcription model and pass the live text to your agent. The agent itself does not change: it still receives text, so this works with any text-based agent. For agents that speak back, pair it with speech generation, or use realtime voice for full two-way conversation. For more information on streaming transcription on AI Gateway, see the documentation. Read more

## VideoIO, PCM Mixing And Timed Whisper Captions

DevFeed: [VideoIO, PCM Mixing And Timed Whisper Captions](<https://devfeed.tech/articles/videoio-pcm-mixing-and-timed-whisper-captions-19662.md>)

Original publisher: [Read original article](<https://www.codenameone.com/blog/videoio-audio-mixer-whisper/>)

Author: Shai Almog

Published: 2026-07-08T00:00:00Z

Content type: release

Language: en

Sources: [CodeName One](<https://devfeed.tech/sources/codename-one.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [cross-platform](<https://devfeed.tech/topics/cross-platform.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [cross-platform](<https://devfeed.tech/tags/cross-platform.md>), [encoding](<https://devfeed.tech/tags/encoding.md>), [playback](<https://devfeed.tech/tags/playback.md>), [release](<https://devfeed.tech/tags/release.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [video](<https://devfeed.tech/tags/video.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

Codename One's release adds cross-platform video encoding and decoding, sample-accurate PCM mixing, and timed Whisper transcription segments with SRT and VTT subtitle output. It can encode app-rendered frames, decode exact frames and audio where supported, and package native integrations across several platforms.

### Source excerpt

Codename One can now generate and inspect real video: encode app-rendered frames, decode exact frames, mix PCM, and turn Whisper timestamps into subtitles.

## Build realtime voice agents on AI Gateway

DevFeed: [Build realtime voice agents on AI Gateway](<https://devfeed.tech/articles/build-realtime-voice-agents-on-ai-gateway-772.md>)

Original publisher: [Read original article](<https://vercel.com/blog/realtime-voice-agents-on-ai-gateway>)

Author: Kevin Dawkins

Published: 2026-06-29T07:00:00Z

Content type: tutorial

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [observability](<https://devfeed.tech/tags/observability.md>), [openai](<https://devfeed.tech/tags/openai.md>), [routing](<https://devfeed.tech/tags/routing.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [tools](<https://devfeed.tech/tags/tools.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Vercel AI Gateway adds beta support for realtime voice, text-to-speech, and speech-to-text in AI SDK 7. The article explains browser-based realtime sessions, token-based authentication, interruption handling, tool calls, and direct audio input/output models.

### Source excerpt

AI Gateway now supports audio/voice. You can add realtime voice, text to speech, and speech to text with the same calls you already use for text, image, and video, routed through AI Gateway alongside every other modality. Audio launches with models from OpenAI and xAI. Each call gets the same provider routing, observability, spend controls, and bring-your-own-key support you already use for your other models. These capabilities are in beta and available in AI SDK 7. Capability How it works Use it for Realtime voice Live audio in and out, for streaming, low-latency session Two-way voice agents and live conversation Text to speech Text in, audio file out, single request Voiceovers, spoken responses, audio versions of written content Speech to text Recorded audio in, text out, single request Transcribing voice notes, call recordings Getting started Realtime, speech, and transcription model are supported on AI SDK 7. Realtime voice agents Realtime turns your app into something a user can hold a conversation with. When they speak, the model responds right away. Because it replies in the moment instead of waiting for a full turn, users can interrupt and talk over it the way they would with a person. It fits voice assistants, customer support agents, hands-free tools, and anywhere a user would rather talk than type. What sets it apart from chaining models together is that a single realtime model hears audio and produces audio directly, instead of running a speech-to-text, then language model, then text-to-speech pipeline. In the browser, the useRealtime hook manages the WebSocket connection, microphone capture, and audio playback. The connection is authenticated with your AI Gateway credential, so you mint a short-lived token on the server and hand the browser only that token. Your API key never reaches the client. Add a route that mints the token: Then connect from a client component: The hook captures the microphone, streams the audio to the model through AI Gateway, and

## Realtime voice, speech, and transcription now supported on AI Gateway

DevFeed: [Realtime voice, speech, and transcription now supported on AI Gateway](<https://devfeed.tech/articles/realtime-voice-speech-and-transcription-now-supported-on-ai-gateway-1067.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/realtime-voice-speech-and-transcription-now-supported-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-06-29T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [real-time](<https://devfeed.tech/topics/real-time.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [browser](<https://devfeed.tech/topics/browser.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>), [observability](<https://devfeed.tech/topics/observability.md>), [App](<https://devfeed.tech/topics/app.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [app](<https://devfeed.tech/tags/app.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [code](<https://devfeed.tech/tags/code.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [playground](<https://devfeed.tech/tags/playground.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

AI Gateway adds beta support for realtime voice and audio models through AI SDK 7, enabling voice agents, text-to-speech, and speech-to-text. Developers can use the realtime example, quickstart, or browser playground, with observability, spend controls, bring-your-own-key support, and no markup or platform fees.

### Source excerpt

AI Gateway now supports voice and audio models. You can build realtime voice agents, generate speech from text, and transcribe audio to text. This provides the same observability, spend controls, and bring-your-own-key support as text, image, and video models in AI Gateway, with no markup or platform fees. These capabilities are in beta and available via AI SDK 7. With realtime support, a single model takes audio in and audio out, so a user can talk and hear a reply back in near real time instead of waiting on a chain of separate models. Capability What it does Realtime voice agents Model listens to the user, works out a response, and speaks it back in a live, low-latency conversation. It can call your tools mid-conversation to look something up or take an action. The useRealtime hook handles microphone capture and playback. Text to speech Generate spoken audio from text, with a selectable voice and output format such as MP3. Use it for voiceovers, audio versions of written content, and spoken responses. Speech to text Transcribe recordings into text, from a file buffer, base64 string, or URL. Use it for voice notes or other transcriptions. Two ways to get started: Follow the realtime example below or the realtime quickstart to add a voice agent to your app. Use the playground. Talk to a realtime model in the browser, no code required, in the AI Gateway Playground. Realtime example A voice agent has two pieces: a server route that mints a short-lived token, so your API key never reaches the client, and a browser component that connects with it. Add the token route: Then connect from the browser. The useRealtime hook fetches that route and manages the WebSocket connection, microphone capture, and audio playback: Playground You can also try audio models without writing any code. Open the models page, click into a model, and interact with it right in the browser: Talk to a realtime model to hold a voice conversation Send text and have a transcription model read it back

## A New Framework for Evaluating Voice Agents (EVA)

DevFeed: [A New Framework for Evaluating Voice Agents (EVA)](<https://devfeed.tech/articles/a-new-framework-for-evaluating-voice-agents-eva-7048.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/eva>)

Author: Tara Bogavelli; Gabrielle Gauthier Melancon; Katrina Stankiewicz; Nifemi Bamgbose; Hoang Nguyen; Raghav Mehndiratta; Hari Subramani; Fanny Riols

Published: 2026-03-24T02:01:52Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Bot](<https://devfeed.tech/topics/bot.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [demo](<https://devfeed.tech/tags/demo.md>), [eval](<https://devfeed.tech/tags/eval.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

EVA is an end-to-end framework for evaluating conversational voice agents across both task accuracy and conversational experience. It scores complete multi-turn spoken conversations, includes an airline dataset of 50 scenarios, and reports benchmark results for cascade and audio-native systems. The article highlights a recurring tradeoff: stronger task completion can coincide with worse user experience.

### Source excerpt

Conversational voice agents present a distinct evaluation challenge: they must simultaneously satisfy two objectives -- accuracy (completing the user's task correctly and faithfully) and conversational experience (doing so naturally, concisely, and in a way appropriate for spoken interaction).

## Building a Local Voice Dictation Device with Raspberry Pi, Whisper, and Ollama

DevFeed: [Building a Local Voice Dictation Device with Raspberry Pi, Whisper, and Ollama](<https://devfeed.tech/articles/i-built-my-own-wisprflow-fully-local-under-50-and-it-types-into-any-computer-25154.md>)

Original publisher: [Read original article](<https://blog.droidchef.dev/i-built-my-own-wisprflow-fully-local-under-50-and-it-types-into-any-computer/>)

Author: Ishan Khanna

Published: 2026-03-16T20:55:45Z

Content type: tutorial

Language: en

Sources: [Ishan Khanna](<https://devfeed.tech/sources/ishan-khanna.md>)

Topics: [Raspberry Pi](<https://devfeed.tech/topics/raspberry-pi.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Ollama](<https://devfeed.tech/topics/ollama.md>), [ASGI](<https://devfeed.tech/topics/asgi.md>), [Python](<https://devfeed.tech/topics/python.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [CircuitPython](<https://devfeed.tech/topics/circuitpython.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [audio](<https://devfeed.tech/tags/audio.md>), [circuitpython](<https://devfeed.tech/tags/circuitpython.md>), [fastapi](<https://devfeed.tech/tags/fastapi.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [llms](<https://devfeed.tech/tags/llms.md>), [ollama](<https://devfeed.tech/tags/ollama.md>), [python](<https://devfeed.tech/tags/python.md>), [raspberry-pi](<https://devfeed.tech/tags/raspberry-pi.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [whisper](<https://devfeed.tech/tags/whisper.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

This tutorial describes a local voice dictation device built with a Raspberry Pi Zero W, Raspberry Pi Pico, and an INMP441 microphone. Audio is sent over Wi-Fi to a Windows PC for Whisper transcription and Ollama text cleanup, then returned through a USB keyboard interface via a KVM switch. The author reports about $40 in hardware costs and under 700 milliseconds of end-to-end latency.

### Source excerpt

I spend most of my day talking to AI agents in the terminal. Claude Code, ChatGPT, aider -- you name it. And every time I have to type out a long, detailed prompt explaining what I want refactored, I think: why am I typing this when I could just say

## WAXAL: A large-scale open resource for African language speech technology

DevFeed: [WAXAL: A large-scale open resource for African language speech technology](<https://devfeed.tech/articles/waxal-a-large-scale-open-resource-for-african-language-speech-technology-6927.md>)

Original publisher: [Read original article](<https://research.google/blog/waxal-a-large-scale-open-resource-for-african-language-speech-technology/>)

Published: 2026-03-06T20:06:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [data](<https://devfeed.tech/topics/data.md>), [Google](<https://devfeed.tech/topics/google.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [africa](<https://devfeed.tech/tags/africa.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [google](<https://devfeed.tech/tags/google.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [research](<https://devfeed.tech/tags/research.md>), [resources](<https://devfeed.tech/tags/resources.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google Research introduces WAXAL, an open-access speech dataset covering 27 Sub-Saharan African languages. The release includes approximately 1,846 hours of transcribed ASR data and more than 565 hours of high-fidelity TTS recordings under a CC-BY-4.0 license.

### Source excerpt

Natural Language Processing

## How Descript engineers multilingual video dubbing at scale

DevFeed: [How Descript engineers multilingual video dubbing at scale](<https://devfeed.tech/articles/how-descript-engineers-multilingual-video-dubbing-at-scale-6375.md>)

Original publisher: [Read original article](<https://openai.com/index/descript>)

Published: 2026-03-06T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [llms](<https://devfeed.tech/tags/llms.md>), [openai](<https://devfeed.tech/tags/openai.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [scale](<https://devfeed.tech/tags/scale.md>), [startup](<https://devfeed.tech/tags/startup.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [video](<https://devfeed.tech/tags/video.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

Descript redesigned its video-translation pipeline with OpenAI reasoning models to generate dubbed content that preserves meaning while fitting the original timing. The company reports higher dubbed-video exports and improved duration adherence after rollout.

### Source excerpt

Using OpenAI reasoning models, Descript unlocked automatic localization of large content libraries without losing timing or meaning.

## One-Shot Any Web App with Gradio's gr.HTML

DevFeed: [One-Shot Any Web App with Gradio's gr.HTML](<https://devfeed.tech/articles/one-shot-any-web-app-with-gradio-s-gr-html-7227.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/gradio-html-one-shot-apps>)

Author: yuvraj sharma; hysts; Freddy Boulton

Published: 2026-02-18T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [web applications](<https://devfeed.tech/topics/web-applications.md>), [modern web development](<https://devfeed.tech/topics/modern-web-development.md>), [HTML5 and CSS3 tricks](<https://devfeed.tech/topics/html5-and-css3-tricks.md>), [React UI animations](<https://devfeed.tech/topics/react-ui-animations.md>)

Tags: [3d](<https://devfeed.tech/tags/3d.md>), [app](<https://devfeed.tech/tags/app.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [claude](<https://devfeed.tech/tags/claude.md>), [community](<https://devfeed.tech/tags/community.md>), [css](<https://devfeed.tech/tags/css.md>), [gradio](<https://devfeed.tech/tags/gradio.md>), [html](<https://devfeed.tech/tags/html.md>), [html5](<https://devfeed.tech/tags/html5.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [object-detection](<https://devfeed.tech/tags/object-detection.md>), [python](<https://devfeed.tech/tags/python.md>), [react](<https://devfeed.tech/tags/react.md>), [spaces](<https://devfeed.tech/tags/spaces.md>), [speech](<https://devfeed.tech/tags/speech.md>), [three-js](<https://devfeed.tech/tags/three-js.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [vibe-coding](<https://devfeed.tech/tags/vibe-coding.md>)

### AI overview

The article shows how Gradio's gr.HTML can create interactive, single-file Python web apps without a separate frontend build step. Examples include timers, a kanban board, visual ML result viewers, 3D controls, and live speech transcription.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Voxtral transcribes at the speed of sound.

DevFeed: [Voxtral transcribes at the speed of sound.](<https://devfeed.tech/articles/voxtral-transcribes-at-the-speed-of-sound-7134.md>)

Original publisher: [Read original article](<https://mistral.ai/news/voxtral-transcribe-2/>)

Published: 2026-02-04T16:00:00Z

Content type: article

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Security](<https://devfeed.tech/topics/security.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [apache](<https://devfeed.tech/tags/apache.md>), [arabic](<https://devfeed.tech/tags/arabic.md>), [audio](<https://devfeed.tech/tags/audio.md>), [batch](<https://devfeed.tech/tags/batch.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cost](<https://devfeed.tech/tags/cost.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [performance](<https://devfeed.tech/tags/performance.md>), [security](<https://devfeed.tech/tags/security.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Mistral introduces Voxtral Transcribe 2, a family of speech-to-text models comprising Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications. The release highlights speaker diarization, word-level timestamps, multilingual transcription in 13 languages, configurable sub-200 ms latency, streaming transcription, and open Apache 2.0 weights for edge deployment.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## AI model choices 2026-01

DevFeed: [AI model choices 2026-01](<https://devfeed.tech/articles/ai-model-choices-2026-01-25215.md>)

Original publisher: [Read original article](<https://kau.sh/blog/ai-model-choices/>)

Author: Kaushik Gopal

Published: 2026-01-13T19:39:48Z

Content type: opinion

Language: en

Sources: [Kaushik Gopal's Site](<https://devfeed.tech/sources/kaushik-gopal-s-site.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [coding](<https://devfeed.tech/tags/coding.md>), [generation](<https://devfeed.tech/tags/generation.md>), [image](<https://devfeed.tech/tags/image.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>), [writing](<https://devfeed.tech/tags/writing.md>)

### AI overview

The author shares a current set of preferred AI models for different tasks, including planning and writing, coding and tool calling, learning, quick questions, image generation, and voice transcription.

### Source excerpt

Which AI model do I use? This is a common question I get asked, but models evolve so rapidly that I never felt like I could give an answer that would stay relevant for more than a month or two. This year, I finally feel like I have a stable set of model choices that consistently give me good results. I'm jotting it down here to share more broadly and to trace how my own choices evolve over time. GPT 5.2 (High) for planning and writing, including plans Opus 4.5 for anything coding, task automation, and tool calling Gemini's range of models for everything else: Gemini 3 (Thinking) for learning and understanding concepts (underrated) Gemini 3 (Flash) for quick fire questions Nano Banana (obv) for image generation NVIDIA's Parakeet for voice transcription

## Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks

DevFeed: [Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks](<https://devfeed.tech/articles/open-asr-leaderboard-trends-and-insights-with-new-multilingual-long-form-tracks-7409.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Vaibhav Srivastav

Published: 2025-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [trends](<https://devfeed.tech/tags/trends.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The Open ASR Leaderboard expands its evaluation with multilingual and long-form transcription tracks. The article highlights accuracy advantages from combining Conformer encoders with LLM decoders, throughput advantages from CTC and TDT decoders, Whisper as a multilingual baseline, and the effects of fine-tuning on specialized performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## How mpathic built better ML workflows by switching from Elasticsearch to ClickHouse Cloud

DevFeed: [How mpathic built better ML workflows by switching from Elasticsearch to ClickHouse Cloud](<https://devfeed.tech/articles/how-mpathic-built-better-ml-workflows-by-switching-from-elasticsearch-to-clickhouse-cloud-5430.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/mpathic-better-ml-elastic-to-clickhouse-migration>)

Author: ClickHouse

Published: 2025-10-10T16:20:55Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [elasticsearch](<https://devfeed.tech/topics/elasticsearch.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [audio](<https://devfeed.tech/tags/audio.md>), [build-faster](<https://devfeed.tech/tags/build-faster.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [elasticsearch](<https://devfeed.tech/tags/elasticsearch.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [ml](<https://devfeed.tech/tags/ml.md>), [model-training](<https://devfeed.tech/tags/model-training.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

mpathic uses AI to analyze therapy-session audio for patient safety and compliance in clinical drug trials. The article explains how noisy recordings require speaker segmentation, fingerprinting, transcription, classification, model training, and analytics, and describes the company's move from Elasticsearch and EC2-based pipelines to ClickHouse Cloud for faster, leaner ML workflows.

### Source excerpt

Learn why Elasticsearch was holding mpathic back, and how switching to ClickHouse Cloud helped them build faster, leaner ML workflows.

## Mistral releases Voxtral open speech understanding models

DevFeed: [Mistral releases Voxtral open speech understanding models](<https://devfeed.tech/articles/voxtral-7138.md>)

Original publisher: [Read original article](<https://mistral.ai/news/voxtral/>)

Published: 2025-07-15T12:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [release](<https://devfeed.tech/tags/release.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Mistral announces Voxtral, a family of open speech understanding models available in 24B and 3B variants. The models support transcription, semantic understanding, audio Q&A and summarization, with deployment options for production, local, and edge use.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Blazingly fast whisper transcriptions with Inference Endpoints

DevFeed: [Blazingly fast whisper transcriptions with Inference Endpoints](<https://devfeed.tech/articles/blazingly-fast-whisper-transcriptions-with-inference-endpoints-7192.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/fast-whisper-endpoints>)

Author: Morgan Funtowicz; Freddy Boulton; Steven Zheng; Vaibhav Srivastav; Erik Kaunismäki; Michelle Habonneau

Published: 2025-05-13T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

Hugging Face introduces an optimized Whisper inference endpoint powered by vLLM. It targets newer NVIDIA GPUs and combines PyTorch compilation, CUDA graphs, and float8 KV-cache quantization to improve transcription inference speed and memory efficiency.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## HuggingFace, IISc partner to supercharge model building on India's diverse languages

DevFeed: [HuggingFace, IISc partner to supercharge model building on India's diverse languages](<https://devfeed.tech/articles/huggingface-iisc-partner-to-supercharge-model-building-on-india-s-diverse-languages-7273.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/iisc-huggingface-collab>)

Author: Prasanta Kumar Ghosh; Nihar Desai; Sanka; Sujith Pulikodan

Published: 2025-02-27T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [audio](<https://devfeed.tech/tags/audio.md>), [community](<https://devfeed.tech/tags/community.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [google](<https://devfeed.tech/tags/google.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [india](<https://devfeed.tech/tags/india.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-collab](<https://devfeed.tech/tags/open-source-collab.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Hugging Face and IISc/ARTPARK are partnering to improve access to and usability of Project Vaani, an open-source multimodal dataset representing India's linguistic diversity. The dataset includes speech and transcribed text from languages and dialects across the country's districts, supporting speech recognition, language modeling, segmentation, and other AI applications.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Voice recognition

DevFeed: [Voice recognition](<https://devfeed.tech/articles/voice-recognition-36609.md>)

Original publisher: [Read original article](<http://www.imperialviolet.org/2023/07/29/voice-recognition.html>)

Author: Adam Langley

Published: 2023-07-29T00:00:00Z

Content type: opinion

Language: en

Sources: [ImperialViolet](<https://devfeed.tech/sources/imperialviolet.md>)

Topics: [Whisper](<https://devfeed.tech/topics/whisper.md>), [iOS](<https://devfeed.tech/topics/ios.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [caveat](<https://devfeed.tech/tags/caveat.md>), [ios](<https://devfeed.tech/tags/ios.md>), [llms](<https://devfeed.tech/tags/llms.md>), [performance](<https://devfeed.tech/tags/performance.md>), [script](<https://devfeed.tech/tags/script.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The author compares iOS 17 voice recognition with earlier results and finds that, although performance improved over iOS 16, it still produces too many errors with technical terms. The article then reports that Whisper solved the practical voice-recognition problem, while noting that it can append fabricated sentences that are removable.

### Source excerpt

Update: Evan let me know that Whisper solved the voice recognition problem. He has a wrapper that records from a microphone and prints the transcription here. Whisper is very impressive and the only caveat is that it sometimes inserts whole fabricated sentences at the end. The words always sort of make sense in context, but there were no sounds that could possibly have caused it. It's always at the very end in my experience, and it's no problem to remove it so, with that noted, you should ignore everything below because Whisper is a better answer. Last week's blog post was rather long, and had a greater than normal number of typos. (Thanks to people who pointed them out. I think I've fixed all the ones that were reported.) This was because I saw in reviews that iOS 17's voice recognition was supposed to be much improved, and I figured that I'd give it a try. I've always found iOS's recognition to be superior to Google Docs and I have an old iPad Pro that's good for betas. iOS's performance remains good and, yes, I think it's better than iOS 16. But it's still hardly at the level of "magic", especially when using technical terms. Here's a paragraph taken directly from the raw output of last week's post (I've highlighted errors with italics): It is integrated into the W3C credential management specification and so it is called via navigator . credentials . create and navigator .credentials. get. This document is about understanding the deeper structures that underpin web orphan rather than being a guy as to its details. So we will leave a great many details to the numerous guides to Web Oran that already exist on the web and instead focus on how structures from UF were carried over into Web orphan and updated. While it's nice that many of the words are there, with that density of errors doing all the corrections means that it's not clearly better than typing things out. However, the world is all aflutter about LLMs these days. Can they help? I wrote a script to chunk

[Next page](<https://devfeed.tech/tags/transcription.md?cursor=WyIyMDIzLTA3LTI5VDAwOjAwOjAwKzAwOjAwIiwgImE1Mjc3N2M0LTdiM2QtNDQ0MS04MDY4LWRmZTk4OTFlZTMzZSJd>)