# speech-to-speech

Published articles for speech-to-speech.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

DevFeed: [Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking](<https://devfeed.tech/articles/google-gemini-3-8-live-40886.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/selectel/news/1082998/>)

Author: techno\_mot (Selectel)

Published: 2026-09-16T13:40:43Z

Content type: news

Language: ru

Sources: [Tagir Valeev](<https://devfeed.tech/sources/tagir-valeev.md>)

Topics: [Google](<https://devfeed.tech/topics/google.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [llm](<https://devfeed.tech/tags/llm.md>), [realtime](<https://devfeed.tech/tags/realtime.md>), [selectel](<https://devfeed.tech/tags/selectel.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [synthid](<https://devfeed.tech/tags/synthid.md>), [tag-efc6fa45f1fe](<https://devfeed.tech/tags/tag-efc6fa45f1fe.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google introduced two native speech-to-speech models: Gemini 3.8 Live, focused on scale and cost, and Gemini 3.8 Live Extended Thinking, designed for multi-step tasks with reasoning during dialogue. The article discusses direct audio processing, asynchronous function calls, visual context, multilingual conversations, SynthID watermarking, pricing, and availability through the Gemini API and Google AI Studio.

### Source excerpt

15 сентября Google представила две нативные speech-to-speech модели, Gemini 3.8 Live и Gemini 3.8 Live Extended Thinking. Первая заточена под масштаб и цену, вторая -- под многошаговые задачи с рассуждением прямо в диалоге. Интереснее баллов то, что голосовые агенты наконец получили признаки продакшен-продукта. Асинхронные вызовы функций, предсказуемая цена, интеграции с тем, на чем такие системы реально собирают. Хочу напомнить, как это делалось раньше. Голосовой бот -- конвейер из трех сервисов. Распознавание речи, языковая модель и синтез. Задержки складываются, а вместе с текстом может потеряться все остальное -- интонация, скорость речи, эмоция, шум на фоне. Модель получает расшифровку и не знает, что собеседник злится. Проблемным было и перебивание. Пока распознавание не закрыло фразу, система вообще не понимает, что ее прервали. Нативная speech-to-speech модель работает с аудио напрямую и снимает оба ограничения разом. Читать далее

## Gemini Live audio

DevFeed: [Gemini Live audio](<https://devfeed.tech/articles/gemini-live-audio-31180.md>)

Original publisher: [Read original article](<https://simonwillison.net/2026/Sep/15/gemini-live/>)

Author: Simon Willison

Published: 2026-09-15T22:47:07Z

Content type: tutorial

Language: en

Sources: [Simon Willison's Weblog](<https://devfeed.tech/sources/simon-willison-s-weblog.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [browser](<https://devfeed.tech/topics/browser.md>), [Playback](<https://devfeed.tech/topics/playback.md>), [implementation](<https://devfeed.tech/topics/implementation.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [gemini-196](<https://devfeed.tech/tags/gemini-196.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [generative-ai-1-982](<https://devfeed.tech/tags/generative-ai-1-982.md>), [google](<https://devfeed.tech/tags/google.md>), [google-416](<https://devfeed.tech/tags/google-416.md>), [llm-release](<https://devfeed.tech/tags/llm-release.md>), [llm-release-231](<https://devfeed.tech/tags/llm-release-231.md>), [llms](<https://devfeed.tech/tags/llms.md>), [llms-1-948](<https://devfeed.tech/tags/llms-1-948.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [playback](<https://devfeed.tech/tags/playback.md>), [release](<https://devfeed.tech/tags/release.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [speech-to-text](<https://devfeed.tech/tags/speech-to-text.md>), [speech-to-text-21](<https://devfeed.tech/tags/speech-to-text-21.md>), [tools](<https://devfeed.tech/tags/tools.md>), [tools-78](<https://devfeed.tech/tags/tools-78.md>), [ui](<https://devfeed.tech/tags/ui.md>), [voice](<https://devfeed.tech/tags/voice.md>), [websocket](<https://devfeed.tech/tags/websocket.md>), [websockets](<https://devfeed.tech/tags/websockets.md>), [websockets-21](<https://devfeed.tech/tags/websockets-21.md>)

### AI overview

The article describes a browser-based web UI for trying Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking speech-to-speech models. The implementation supports model and voice selection, an optional system prompt, voice conversations, and interruption while the model is speaking. It uses no libraries, connecting to a WebSocket endpoint and using the Web Audio API for capture and playback.

### Source excerpt

Tool: Gemini Live audio Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family. I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking. The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback. Here's the Gemini Live tutorial for getting started with that WebSockets API. Tags: google, tools, websockets, generative-ai, llms, gemini, llm-release, speech-to-text

## Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

DevFeed: [Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking](<https://devfeed.tech/articles/introducing-gemini-3-8-live-and-3-8-live-extended-thinking-26922.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/>)

Author: Tom Ouyang

Published: 2026-09-15T17:05:57Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Google](<https://devfeed.tech/topics/google.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [cost](<https://devfeed.tech/tags/cost.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [none](<https://devfeed.tech/tags/none.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two models designed for near-real-time voice interaction and reasoning. The release describes visual grounding, multilingual conversation, background tool and API execution, and deeper reasoning for complex workflows.

### Source excerpt

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet, built for natural conversation.

## Grok Voice Think Fast 2.0: Benchmarks, pricing, specs

DevFeed: [Grok Voice Think Fast 2.0: Benchmarks, pricing, specs](<https://devfeed.tech/articles/grok-voice-think-fast-2-0-benchmarks-pricing-specs-16520.md>)

Original publisher: [Read original article](<https://appwrite.io/blog/post/whats-new-in-grok-voice-think-fast-20>)

Author: Aishwari Pahwa

Published: 2026-07-31T00:00:00Z

Content type: release

Language: en

Sources: [Appwrite Blog](<https://devfeed.tech/sources/appwrite-blog.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [latency](<https://devfeed.tech/tags/latency.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

This developer article introduces SpaceXAI's Grok Voice Think Fast 2.0, a speech-to-speech model announced on July 29, 2026. It covers vendor-reported benchmark results, latency, pricing, migration, and backend requirements, including 82.9% speech-to-speech quality, 0.70 seconds to first audio, and $0.08 per minute.

### Source excerpt

Grok Voice Think Fast 2.0 hits 0.70s time to first audio, 82.9% speech-to-speech quality, and $0.08 per minute. See the benchmarks, pricing, and migration path.

## Grok Voice Think Fast 2.0 now available on AI Gateway

DevFeed: [Grok Voice Think Fast 2.0 now available on AI Gateway](<https://devfeed.tech/articles/grok-voice-think-fast-2-0-now-available-on-ai-gateway-977.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/grok-voice-think-fast-2-0-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-29T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [API](<https://devfeed.tech/topics/api.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [latency](<https://devfeed.tech/tags/latency.md>), [model](<https://devfeed.tech/tags/model.md>), [playground](<https://devfeed.tech/tags/playground.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [server](<https://devfeed.tech/tags/server.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tool](<https://devfeed.tech/tags/tool.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Grok Voice Think Fast 2.0 from xAI is now available through AI Gateway as a speech-to-speech model. It processes audio input and output while reasoning in parallel, using fewer reasoning tokens to reduce tool-call delays. The article also describes transcription performance in noisy and compressed audio conditions and access through the AI SDK realtime API.

### Source excerpt

Grok Voice Think Fast 2.0 from xAI is now available on AI Gateway. It is a speech-to-speech voice model that takes audio in and audio out, improving on the previous Grok Voice model in reasoning, transcription accuracy, and conversation. The model reasons in parallel with speech, so it can think through a query while talking without adding latency. It has also been trained to use fewer reasoning tokens than before, so tool calls fire sooner, often before the end of the agent's first sentence. Transcription holds up in real-world conditions, including background noise and telephony compression. Use xai/grok-voice-think-fast-2.0 through the AI SDK's realtime API. Mint a short-lived token on the server so your API key never reaches the client: See the realtime documentation to build a voice agent with Grok Voice Think Fast 2.0. Try out the model in the AI Gateway playground. Read more

## Fluid, natural voice translation with Gemini 3.5 Live Translate

DevFeed: [Fluid, natural voice translation with Gemini 3.5 Live Translate](<https://devfeed.tech/articles/fluid-natural-voice-translation-with-gemini-3-5-live-translate-6153.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/fluid-natural-voice-translation-with-gemini-35-live-translate/>)

Author: Anuda Weerasinghe

Published: 2026-06-09T15:16:25Z

Content type: release

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [developer](<https://devfeed.tech/tags/developer.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [google-ai](<https://devfeed.tech/tags/google-ai.md>), [none](<https://devfeed.tech/tags/none.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google announces Gemini 3.5 Live Translate, an audio model for continuous speech-to-speech translation that preserves vocal qualities across more than 70 languages. Developers can access it in public preview through the Gemini Live API and Google AI Studio.

### Source excerpt

Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.

## Reachy Mini goes fully local

DevFeed: [Reachy Mini goes fully local](<https://devfeed.tech/articles/reachy-mini-goes-fully-local-7342.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/local-reachy-mini-conversation>)

Author: Amir Mahla; Andres Marafioti

Published: 2026-05-27T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [reachy](<https://devfeed.tech/topics/reachy.md>), [gemma4](<https://devfeed.tech/topics/gemma4.md>), [llama.cpp](<https://devfeed.tech/topics/llama-cpp.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [blog](<https://devfeed.tech/tags/blog.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llm](<https://devfeed.tech/tags/llm.md>), [local](<https://devfeed.tech/tags/local.md>), [local-ai](<https://devfeed.tech/tags/local-ai.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reachy](<https://devfeed.tech/tags/reachy.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [server](<https://devfeed.tech/tags/server.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>)

### AI overview

A tutorial for running fully local conversations with a Reachy Mini robot. It describes a cascaded VAD, speech-to-text, LLM, and text-to-speech pipeline using llama.cpp with Gemma 4, Silero VAD, Parakeet-TDT STT, and Qwen3-TTS, connected through a Realtime API-compatible WebSocket. The setup avoids cloud services, API keys, and sending data off the local machine.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## A New Framework for Evaluating Voice Agents (EVA)

DevFeed: [A New Framework for Evaluating Voice Agents (EVA)](<https://devfeed.tech/articles/a-new-framework-for-evaluating-voice-agents-eva-7048.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ServiceNow-AI/eva>)

Author: Tara Bogavelli; Gabrielle Gauthier Melancon; Katrina Stankiewicz; Nifemi Bamgbose; Hoang Nguyen; Raghav Mehndiratta; Hari Subramani; Fanny Riols

Published: 2026-03-24T02:01:52Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Bot](<https://devfeed.tech/topics/bot.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [demo](<https://devfeed.tech/tags/demo.md>), [eval](<https://devfeed.tech/tags/eval.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

EVA is an end-to-end framework for evaluating conversational voice agents across both task accuracy and conversational experience. It scores complete multi-turn spoken conversations, includes an airline dataset of 50 scenarios, and reports benchmark results for cascade and audio-native systems. The article highlights a recurring tradeoff: stronger task completion can coincide with worse user experience.

### Source excerpt

Conversational voice agents present a distinct evaluation challenge: they must simultaneously satisfy two objectives -- accuracy (completing the user's task correctly and faithfully) and conversational experience (doing so naturally, concisely, and in a way appropriate for spoken interaction).

## Faking it on the phone: How to tell if a voice call is AI or not

DevFeed: [Faking it on the phone: How to tell if a voice call is AI or not](<https://devfeed.tech/articles/faking-it-on-the-phone-how-to-tell-if-a-voice-call-is-ai-or-not-8332.md>)

Original publisher: [Read original article](<https://www.welivesecurity.com/en/business-security/faking-it-phone-how-tell-voice-call-ai/>)

Author: Phil Muncaster

Published: 2026-02-23T10:00:00Z

Content type: article

Language: en

Sources: [WeLiveSecurity](<https://devfeed.tech/sources/welivesecurity.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Authentication](<https://devfeed.tech/topics/authentication.md>), [Finance](<https://devfeed.tech/topics/finance.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [authentication](<https://devfeed.tech/tags/authentication.md>), [business](<https://devfeed.tech/tags/business.md>), [fraud](<https://devfeed.tech/tags/fraud.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [mfa](<https://devfeed.tech/tags/mfa.md>), [scams](<https://devfeed.tech/tags/scams.md>), [security](<https://devfeed.tech/tags/security.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

The article explains how generative AI enables convincing deepfake voice calls that can impersonate executives or suppliers and facilitate financial fraud, password resets, and MFA reset requests. It outlines how attackers obtain voice samples, select targets, and use scripted or near-real-time speech-to-speech tools.

### Source excerpt

Can you believe your ears? Increasingly, the answer is no. Here's what's at stake for your business, and how to beat the deepfakers.

## ServiceNow powers actionable enterprise AI with OpenAI

DevFeed: [ServiceNow powers actionable enterprise AI with OpenAI](<https://devfeed.tech/articles/servicenow-powers-actionable-enterprise-ai-with-openai-6649.md>)

Original publisher: [Read original article](<https://openai.com/index/servicenow-powers-actionable-enterprise-ai-with-openai>)

Published: 2026-01-20T05:45:00Z

Content type: news

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [building](<https://devfeed.tech/tags/building.md>), [business](<https://devfeed.tech/tags/business.md>), [customer](<https://devfeed.tech/tags/customer.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [global-affairs](<https://devfeed.tech/tags/global-affairs.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [openai](<https://devfeed.tech/tags/openai.md>), [operations](<https://devfeed.tech/tags/operations.md>), [platform](<https://devfeed.tech/tags/platform.md>), [product](<https://devfeed.tech/tags/product.md>), [sales](<https://devfeed.tech/tags/sales.md>), [scale](<https://devfeed.tech/tags/scale.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [technology](<https://devfeed.tech/tags/technology.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

ServiceNow and OpenAI announce a multi-year agreement to expand access to OpenAI frontier models across ServiceNow enterprise workflows. The collaboration also brings multimodal, speech-to-speech, and native voice capabilities to help enterprises automate work at scale within secure infrastructure.

### Source excerpt

ServiceNow expands access to OpenAI frontier models to power AI-driven enterprise workflows, summarization, search, and voice across the ServiceNow Platform.

## Improved Gemini audio models for powerful voice experiences

DevFeed: [Improved Gemini audio models for powerful voice experiences](<https://devfeed.tech/articles/improved-gemini-audio-models-for-powerful-voice-experiences-6189.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/improved-gemini-audio-models-for-powerful-voice-experiences/>)

Author: Bibo Xu

Published: 2025-12-12T17:50:50Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Google](<https://devfeed.tech/topics/google.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>)

Tags: [ai-studio](<https://devfeed.tech/tags/ai-studio.md>), [audio](<https://devfeed.tech/tags/audio.md>), [customer-service](<https://devfeed.tech/tags/customer-service.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [none](<https://devfeed.tech/tags/none.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [vertex-ai](<https://devfeed.tech/tags/vertex-ai.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google describes upgrades to Gemini 2.5 Native Audio for live voice agents, including better function calling, instruction following, and multi-turn conversations. The model is available across Google AI Studio and Vertex AI, with rollouts to Gemini Live and Search Live. The article also introduces beta live speech-to-speech translation in Google Translate.

### Source excerpt

An upgraded Gemini 2.5 Native Audio model across Google products and live speech translation in the Google Translate app.

## Real-time speech-to-speech translation

DevFeed: [Real-time speech-to-speech translation](<https://devfeed.tech/articles/real-time-speech-to-speech-translation-6852.md>)

Original publisher: [Read original article](<https://research.google/blog/real-time-speech-to-speech-translation/>)

Published: 2025-11-19T09:59:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [asr](<https://devfeed.tech/tags/asr.md>), [data](<https://devfeed.tech/tags/data.md>), [google](<https://devfeed.tech/tags/google.md>), [ml](<https://devfeed.tech/tags/ml.md>), [personalization](<https://devfeed.tech/tags/personalization.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Google researchers introduce an end-to-end speech-to-speech translation model that translates speech in real time while preserving the original speaker's voice, with a two-second delay. The system uses a streaming architecture, time-synchronized training data, and a scalable data-acquisition pipeline to support more languages and improve conversational naturalness.

### Source excerpt

Algorithms & Theory

## Evaluating Audio Reasoning with Big Bench Audio

DevFeed: [Evaluating Audio Reasoning with Big Bench Audio](<https://devfeed.tech/articles/evaluating-audio-reasoning-with-big-bench-audio-7128.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/big-bench-audio-release>)

Author: Micah Hill-Smith; George Cameron

Published: 2024-12-20T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [arena](<https://devfeed.tech/tags/arena.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [testing](<https://devfeed.tech/tags/testing.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

Big Bench Audio is a 1,000-question dataset for evaluating audio language-model reasoning across speech and text configurations. Initial results report a speech reasoning gap for GPT-4o between text-only and speech-to-speech performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Deploying Speech-to-Speech on Hugging Face

DevFeed: [Deploying Speech-to-Speech on Hugging Face](<https://devfeed.tech/articles/deploying-speech-to-speech-on-hugging-face-7462.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/s2s_endpoint>)

Author: Andres Marafioti; Derek Thomas; Diego Maniloff; Eustache Le Bihan

Published: 2024-10-22T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [docker](<https://devfeed.tech/tags/docker.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

A step-by-step guide to packaging and deploying Hugging Face's Speech-to-Speech pipeline on Inference Endpoints using a custom Docker image.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.