# Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

DevFeed: [Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking](<https://devfeed.tech/articles/google-gemini-3-8-live-40886.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/selectel/news/1082998/>)

Author: techno\_mot (Selectel)

Published: 2026-09-16T13:40:43Z

Content type: news

Language: ru

Sources: [Tagir Valeev](<https://devfeed.tech/sources/tagir-valeev.md>)

Topics: [Google](<https://devfeed.tech/topics/google.md>), [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [llm](<https://devfeed.tech/tags/llm.md>), [realtime](<https://devfeed.tech/tags/realtime.md>), [selectel](<https://devfeed.tech/tags/selectel.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [synthid](<https://devfeed.tech/tags/synthid.md>), [tag-efc6fa45f1fe](<https://devfeed.tech/tags/tag-efc6fa45f1fe.md>), [voice](<https://devfeed.tech/tags/voice.md>)

## AI overview

Google introduced two native speech-to-speech models: Gemini 3.8 Live, focused on scale and cost, and Gemini 3.8 Live Extended Thinking, designed for multi-step tasks with reasoning during dialogue. The article discusses direct audio processing, asynchronous function calls, visual context, multilingual conversations, SynthID watermarking, pricing, and availability through the Gemini API and Google AI Studio.

## Source excerpt

15 сентября Google представила две нативные speech-to-speech модели, Gemini 3.8 Live и Gemini 3.8 Live Extended Thinking. Первая заточена под масштаб и цену, вторая -- под многошаговые задачи с рассуждением прямо в диалоге. Интереснее баллов то, что голосовые агенты наконец получили признаки продакшен-продукта. Асинхронные вызовы функций, предсказуемая цена, интеграции с тем, на чем такие системы реально собирают. Хочу напомнить, как это делалось раньше. Голосовой бот -- конвейер из трех сервисов. Распознавание речи, языковая модель и синтез. Задержки складываются, а вместе с текстом может потеряться все остальное -- интонация, скорость речи, эмоция, шум на фоне. Модель получает расшифровку и не знает, что собеседник злится. Проблемным было и перебивание. Пока распознавание не закрыло фразу, система вообще не понимает, что ее прервали. Нативная speech-to-speech модель работает с аудио напрямую и снимает оба ограничения разом. Читать далее