# text-to-speech

Published articles for text-to-speech.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Build more natural voice experiences with GPT-Live-1 in the API

DevFeed: [Build more natural voice experiences with GPT-Live-1 in the API](<https://devfeed.tech/articles/build-more-natural-voice-experiences-with-gpt-live-1-in-the-api-6496.md>)

Original publisher: [Read original article](<https://openai.com/index/introducing-gpt-live-1-in-the-api>)

Published: 2026-09-10T00:00:00Z

Content type: release

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [product](<https://devfeed.tech/tags/product.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [release](<https://devfeed.tech/tags/release.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [tool](<https://devfeed.tech/tags/tool.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

GPT-Live-1 is launching in the API for full-duplex voice applications. The release emphasizes interruption handling, customizable conversational behavior, delegated reasoning and tool calls, long-session reliability, and telephony support.

### Source excerpt

GPT-Live-1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.

## Personalize your product's text-to-speech voice for any language: Fine-tuning with Kubeflow Trainer on Red Hat OpenShift AI

DevFeed: [Personalize your product's text-to-speech voice for any language: Fine-tuning with Kubeflow Trainer on Red Hat OpenShift AI](<https://devfeed.tech/articles/personalize-your-product-s-text-to-speech-voice-for-any-language-fine-tuning-with-kubeflow-trainer-on-red-hat-openshift-ai-12350.md>)

Original publisher: [Read original article](<https://developers.redhat.com/articles/2026/09/09/text-to-speech-for-any-language-fine-tuning-with-kubeflow-trainer-on-red-hat-openshift-ai>)

Author: Dmytro Hryshchenko, Abhijeet Dhumal

Published: 2026-09-09T03:32:28Z

Content type: tutorial

Language: en

Sources: [Red Hat](<https://devfeed.tech/sources/red-hat.md>), [Red Hat Developer](<https://devfeed.tech/sources/red-hat-developer.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [lora](<https://devfeed.tech/topics/lora.md>), [data](<https://devfeed.tech/topics/data.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [lora](<https://devfeed.tech/tags/lora.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [training](<https://devfeed.tech/tags/training.md>), [voice](<https://devfeed.tech/tags/voice.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

This tutorial explains how to fine-tune the open source Orpheus-3B text-to-speech model for Turkish using Red Hat OpenShift AI and Kubeflow Trainer. It describes packaging distributed training in a TrainJob, scaling across nodes and GPUs, and using LoRA to keep memory usage below 16 GB. The reported result reduces speech errors by more than 90% compared with the base model.

### Source excerpt

Can't Read, Won't Buy. That is the title CSA Research gave its survey of 8,709 consumers across 29 countries, and the numbers justify it: 76% prefer to buy in their own language, and 40% will never buy in another. The same rule governs what your product says out loud. The post Personalize your product's text-to-speech voice for any language: Fine-tuning with Kubeflow Trainer on Red Hat OpenShift AI appeared first on Red Hat Developer.

## 5 ways to use Gemini text-to-speech (TTS) in your apps with Firebase AI Logic

DevFeed: [5 ways to use Gemini text-to-speech (TTS) in your apps with Firebase AI Logic](<https://devfeed.tech/articles/5-ways-to-use-gemini-text-to-speech-tts-in-your-apps-with-firebase-ai-logic-16675.md>)

Original publisher: [Read original article](<https://firebase.blog/posts/2026/09/ai-logic-text-to-speech>)

Author: Ankita Saxena

Published: 2026-09-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Firebase Blog](<https://devfeed.tech/sources/firebase-blog.md>)

Topics: [Firebase](<https://devfeed.tech/topics/firebase.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [digital accessibility](<https://devfeed.tech/topics/digital-accessibility.md>), [Mobile](<https://devfeed.tech/topics/mobile.md>), [Web](<https://devfeed.tech/topics/web.md>)

Tags: [accessibility](<https://devfeed.tech/tags/accessibility.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-logic](<https://devfeed.tech/tags/ai-logic.md>), [firebase](<https://devfeed.tech/tags/firebase.md>), [firebase-ai-logic](<https://devfeed.tech/tags/firebase-ai-logic.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

The article presents five use cases for Gemini text-to-speech through Firebase AI Logic, including language-learning conversation practice, hands-free content reading, and accessibility-focused audio content. It also describes examples involving Finnish language exercises and recipe instructions.

### Source excerpt

News, tutorials, and updates from the Firebase team.

## Fish Audio models now available on Vercel AI Gateway for free

DevFeed: [Fish Audio models now available on Vercel AI Gateway for free](<https://devfeed.tech/articles/fish-audio-models-now-available-on-vercel-ai-gateway-for-free-934.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/fish-audio-models-now-available-on-ai-gateway-for-free>)

Author: Jerilyn Zheng

Published: 2026-08-19T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [free](<https://devfeed.tech/tags/free.md>), [launch](<https://devfeed.tech/tags/launch.md>), [models](<https://devfeed.tech/tags/models.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Fish Audio's text-to-speech and transcription models are available on Vercel AI Gateway. The models are free for 30 days, with AI SDK 7 support for speech generation and transcription, including timestamped segments and word-level timing.

### Source excerpt

Fish Audio's audio models are now available on AI Gateway. To celebrate the launch, every Fish Audio model is free on AI Gateway for the next 30 days, through September 18. Capability Regular Through September 18 Text-to-speech $15.00 per million characters Free Speech-to-text $0.36 per hour of audio Free Four models from Fish Audio are available, including their latest text-to-speech model: fish-audio/s2.1-pro (text-to-speech): Built for low-latency streaming; clones a voice from a reference recording. fish-audio/transcribe-1 (transcription): Returns the text along with the duration of the audio and timestamped segments, down to individual words. fish-audio/s2-pro (text-to-speech): Covers around eighty languages and takes inline tags, plain-language directions written into the text itself, so you can change how a single word or phrase is delivered instead of setting one style for the whole request. fish-audio/s1 (text-to-speech): Reads text that can carry markers for emotion, tone, and sound effects. How to use models during the offer period Using the standard model name (i.e., fish-audio/s2.1-pro) is free, but will automatically begin billing when the offer period ends. To ensure you aren't billed after the free period, add the -free suffix to the standard name, and the model will stop serving when the offer ends (i.e., fish-audio/s2.1-pro-free). Speech and transcription ship in the current AI SDK 7 release. Text-to-speech Generate spoken audio from text with generateSpeech and write the result: Speech-to-text Transcribe recordings into text with transcribe. The audio can be a buffer, a base64 string, or a URL: Each segment carries the text and its start and end time in seconds, down to individual words. Playground You can also try the Fish Audio models without writing any code. Open the models list, click into a model, and send text or audio to hear or read the result in your browser. For a full overview of how to utilize audio models, refer to the speech quickst

## Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS

DevFeed: [Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS](<https://devfeed.tech/articles/build-low-latency-multilingual-voice-agents-open-weights-full-deployment-control-with-nvidia-magpie-tts-7386.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/magpie-tts-multilingual-voice-agents>)

Author: Maryam Motamedi; Mikyas Desta; Jason Li; Jason Roche

Published: 2026-08-10T16:25:36Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [NVIDIA NIM](<https://devfeed.tech/topics/nvidia-nim.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [automation](<https://devfeed.tech/tags/automation.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [data](<https://devfeed.tech/tags/data.md>), [developers](<https://devfeed.tech/tags/developers.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-nim](<https://devfeed.tech/tags/nvidia-nim.md>), [open](<https://devfeed.tech/tags/open.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

This developer article presents NVIDIA Magpie Multilingual TTS as an open-weights, 364M-parameter text-to-speech model for building low-latency multilingual voice applications. It explains how a self-managed cascaded ASR, TTS, and LLM architecture can provide deployment control, tuning, privacy, and predictable performance across 12 languages.

### Source excerpt

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS Every voice interaction has a latency budget. By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retrieving context, and generating a response. Text-to-speech (TTS) is the final step -- and the one users notice most. If speech generation is slow, the whole experience feels slow.

## How we built a realtime system for responsive voice AI in six months

DevFeed: [How we built a realtime system for responsive voice AI in six months](<https://devfeed.tech/articles/how-we-built-a-realtime-system-for-responsive-voice-ai-in-six-months-6358.md>)

Original publisher: [Read original article](<https://openai.com/index/continuous-voice-interaction-with-gpt-live>)

Published: 2026-08-03T07:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [audio](<https://devfeed.tech/tags/audio.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [model](<https://devfeed.tech/tags/model.md>), [speech](<https://devfeed.tech/tags/speech.md>), [systems](<https://devfeed.tech/tags/systems.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>)

### AI overview

OpenAI describes GPT-Live, a full-duplex voice system designed for continuous, natural conversation. Its low-latency architecture streams audio in both directions, separates asynchronous delegation from the core voice path, and coordinates stateful inference, dynamic context management, and protocol-level optimization.

### Source excerpt

GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.

## Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

DevFeed: [Hugging Face and Cerebras bring Gemma 4 to real-time voice AI](<https://devfeed.tech/articles/hugging-face-and-cerebras-bring-gemma-4-to-real-time-voice-ai-7137.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/cerebras-gemma4-voice-ai>)

Author: Amir Mahla; Andres Marafioti; Leandro von Werra; Saurabh Vyas

Published: 2026-07-01T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [gemma4](<https://devfeed.tech/topics/gemma4.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-providers](<https://devfeed.tech/tags/inference-providers.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-collab](<https://devfeed.tech/tags/open-source-collab.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [systems](<https://devfeed.tech/tags/systems.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>)

### AI overview

Hugging Face and Cerebras present a real-time speech-to-speech pipeline that combines Cerebras inference, Google DeepMind's Gemma 4 31B language model, and Qwen text-to-speech. The modular, open architecture is designed for low latency, predictable performance, and adaptable voice experiences across assistants, robots, products, and research projects.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Build realtime voice agents on AI Gateway

DevFeed: [Build realtime voice agents on AI Gateway](<https://devfeed.tech/articles/build-realtime-voice-agents-on-ai-gateway-772.md>)

Original publisher: [Read original article](<https://vercel.com/blog/realtime-voice-agents-on-ai-gateway>)

Author: Kevin Dawkins

Published: 2026-06-29T07:00:00Z

Content type: tutorial

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [audio](<https://devfeed.tech/tags/audio.md>), [observability](<https://devfeed.tech/tags/observability.md>), [openai](<https://devfeed.tech/tags/openai.md>), [routing](<https://devfeed.tech/tags/routing.md>), [speech](<https://devfeed.tech/tags/speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [tools](<https://devfeed.tech/tags/tools.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Vercel AI Gateway adds beta support for realtime voice, text-to-speech, and speech-to-text in AI SDK 7. The article explains browser-based realtime sessions, token-based authentication, interruption handling, tool calls, and direct audio input/output models.

### Source excerpt

AI Gateway now supports audio/voice. You can add realtime voice, text to speech, and speech to text with the same calls you already use for text, image, and video, routed through AI Gateway alongside every other modality. Audio launches with models from OpenAI and xAI. Each call gets the same provider routing, observability, spend controls, and bring-your-own-key support you already use for your other models. These capabilities are in beta and available in AI SDK 7. Capability How it works Use it for Realtime voice Live audio in and out, for streaming, low-latency session Two-way voice agents and live conversation Text to speech Text in, audio file out, single request Voiceovers, spoken responses, audio versions of written content Speech to text Recorded audio in, text out, single request Transcribing voice notes, call recordings Getting started Realtime, speech, and transcription model are supported on AI SDK 7. Realtime voice agents Realtime turns your app into something a user can hold a conversation with. When they speak, the model responds right away. Because it replies in the moment instead of waiting for a full turn, users can interrupt and talk over it the way they would with a person. It fits voice assistants, customer support agents, hands-free tools, and anywhere a user would rather talk than type. What sets it apart from chaining models together is that a single realtime model hears audio and produces audio directly, instead of running a speech-to-text, then language model, then text-to-speech pipeline. In the browser, the useRealtime hook manages the WebSocket connection, microphone capture, and audio playback. The connection is authenticated with your AI Gateway credential, so you mint a short-lived token on the server and hand the browser only that token. Your API key never reaches the client. Add a route that mints the token: Then connect from a client component: The hook captures the microphone, streams the audio to the model through AI Gateway, and

## xAI Grok audio models now available on Vercel AI Gateway

DevFeed: [xAI Grok audio models now available on Vercel AI Gateway](<https://devfeed.tech/articles/xai-grok-audio-models-now-available-on-vercel-ai-gateway-1208.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/xai-grok-audio-models-now-available-on-vercel-ai-gateway>)

Author: Carlton Aikins

Published: 2026-06-29T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [API](<https://devfeed.tech/topics/api.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [React](<https://devfeed.tech/topics/react.md>), [browser](<https://devfeed.tech/topics/browser.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [browser](<https://devfeed.tech/tags/browser.md>), [models](<https://devfeed.tech/tags/models.md>), [observability](<https://devfeed.tech/tags/observability.md>), [react](<https://devfeed.tech/tags/react.md>), [release](<https://devfeed.tech/tags/release.md>), [responses](<https://devfeed.tech/tags/responses.md>), [routing](<https://devfeed.tech/tags/routing.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [server](<https://devfeed.tech/tags/server.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [vercel](<https://devfeed.tech/tags/vercel.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

xAI Grok audio models are now available through Vercel AI Gateway and the AI SDK 7 release. The integration supports realtime voice, text-to-speech, and speech-to-text, with routing, observability, and spend controls.

### Source excerpt

xAI's audio models are now live on AI Gateway. Realtime voice, text to speech, and speech to text are all available through the AI SDK with the same routing, observability, and spend controls as your other models. These capabilities are available on the AI SDK 7 release. Available models Capability Models Realtime voice xai/grok-voice-think-fast-1.0 Text to speech xai/grok-tts Speech to text xai/grok-stt Realtime A voice agent has two pieces: a server route that mints a short-lived token, so your API key never reaches the client, and a browser component that connects with it. Add the token route: this example sets model to xai/grok-voice-think-fast-1.0: Then connect from the browser. The useRealtimehook from @ai-sdk/react fetches that route and manages the WebSocket connection, microphone capture, and audio playback: Text to speech Generate spoken audio from text with generateSpeech. Pass a voice and an output format, then write the result to a file with xai/grok-tts: Speech to text Transcribe recordings into text with transcribe. This example uses xai/grok-stt: Playground You can also try the xAI audio models directly in the AI Gateway playground. Open the models list and click into any of the models to use them directly in the browser. The xai/grok-voice-think-fast-1.0 playground here allows you to talk to the agent and see responses instantly: More information Realtime quickstart Speech quickstart See all xAI models Read more

## Gemini 3.1 Flash TTS: the next generation of expressive AI speech

DevFeed: [Gemini 3.1 Flash TTS: the next generation of expressive AI speech](<https://devfeed.tech/articles/gemini-3-1-flash-tts-the-next-generation-of-expressive-ai-speech-6159.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/gemini-3-1-flash-tts-the-next-generation-of-expressive-ai-speech/>)

Author: Vilobh Meshram

Published: 2026-04-15T16:03:19Z

Content type: news

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>), [Google](<https://devfeed.tech/topics/google.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [building](<https://devfeed.tech/tags/building.md>), [developer](<https://devfeed.tech/tags/developer.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [generation](<https://devfeed.tech/tags/generation.md>), [google](<https://devfeed.tech/tags/google.md>), [none](<https://devfeed.tech/tags/none.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [updates](<https://devfeed.tech/tags/updates.md>)

### AI overview

Google introduces Gemini 3.1 Flash TTS, a text-to-speech model designed for more controllable, expressive, and natural AI speech. It offers native multi-speaker dialogue, support for more than 70 languages, audio tags for directing vocal style and delivery, and availability through the Gemini API, Google AI Studio, Vertex AI, and Google Vids.

### Source excerpt

Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

## Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro

DevFeed: [Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro](<https://devfeed.tech/articles/local-cpu-friendly-high-quality-tts-text-to-speech-with-kokoro-27459.md>)

Original publisher: [Read original article](<https://ariya.io/2026/03/local-cpu-friendly-high-quality-tts-text-to-speech-with-kokoro/>)

Published: 2026-04-01T04:42:35Z

Content type: tutorial

Language: en

Sources: [Ariya Hidayat](<https://devfeed.tech/sources/ariya-hidayat.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [container](<https://devfeed.tech/topics/container.md>), [Docker](<https://devfeed.tech/topics/docker.md>), [API](<https://devfeed.tech/topics/api.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [podman](<https://devfeed.tech/topics/podman.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [container](<https://devfeed.tech/tags/container.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [docker](<https://devfeed.tech/tags/docker.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [openai](<https://devfeed.tech/tags/openai.md>), [podman](<https://devfeed.tech/tags/podman.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [python](<https://devfeed.tech/tags/python.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

A tutorial on running Kokoro, an 82M-parameter text-to-speech model, locally with speech synthesis performed entirely on the CPU. It explains how to launch Kokoro-FastAPI in Docker or Podman, use its web UI or OpenAI-compatible speech API, run JavaScript and Python examples, select voices, and measure generation speed across CPUs.

### Source excerpt

Just a few years ago, realistic local speech generation seemed unimaginable. Today, its quality is exceptional and, crucially, it delivers these results without compromising privacy.

## Mistral AI Releases Voxtral TTS, a Multilingual Text-to-Speech Model

DevFeed: [Mistral AI Releases Voxtral TTS, a Multilingual Text-to-Speech Model](<https://devfeed.tech/articles/speaking-of-voxtral-7136.md>)

Original publisher: [Read original article](<https://mistral.ai/news/voxtral-tts/>)

Published: 2026-03-23T16:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [User experience (UX)](<https://devfeed.tech/topics/ux.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [api](<https://devfeed.tech/tags/api.md>), [audio](<https://devfeed.tech/tags/audio.md>), [customization](<https://devfeed.tech/tags/customization.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [quality](<https://devfeed.tech/tags/quality.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

Mistral AI has released Voxtral TTS, a 4B-parameter text-to-speech model designed to generate realistic, emotionally expressive speech in nine languages. It emphasizes low latency, contextual understanding, speaker modeling, and zero-shot cross-lingual voice adaptation for enterprise voice workflows and AI agents. The model is available through an API and Mistral Studio.

### Source excerpt

Voxtral TTS: A frontier, open-weights text-to-speech model that's fast, instantly adaptable, and produces lifelike speech for voice agents.

## WAXAL: A large-scale open resource for African language speech technology

DevFeed: [WAXAL: A large-scale open resource for African language speech technology](<https://devfeed.tech/articles/waxal-a-large-scale-open-resource-for-african-language-speech-technology-6927.md>)

Original publisher: [Read original article](<https://research.google/blog/waxal-a-large-scale-open-resource-for-african-language-speech-technology/>)

Published: 2026-03-06T20:06:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [data](<https://devfeed.tech/topics/data.md>), [Google](<https://devfeed.tech/topics/google.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [africa](<https://devfeed.tech/tags/africa.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [google](<https://devfeed.tech/tags/google.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [research](<https://devfeed.tech/tags/research.md>), [resources](<https://devfeed.tech/tags/resources.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google Research introduces WAXAL, an open-access speech dataset covering 27 Sub-Saharan African languages. The release includes approximately 1,846 hours of transcribed ASR data and more than 565 hours of high-fidelity TTS recordings under a CC-BY-4.0 license.

### Source excerpt

Natural Language Processing

## Voice Cloning with Consent

DevFeed: [Voice Cloning with Consent](<https://devfeed.tech/articles/voice-cloning-with-consent-7562.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/voice-consent-gate>)

Author: Margaret Mitchell; Lucie-Aimée Kaffee

Published: 2025-10-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [ethics](<https://devfeed.tech/tags/ethics.md>), [guide](<https://devfeed.tech/tags/guide.md>), [speech](<https://devfeed.tech/tags/speech.md>), [systems](<https://devfeed.tech/tags/systems.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-cloning](<https://devfeed.tech/tags/voice-cloning.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

This article presents a voice consent gate for voice cloning. The system requires a speaker to generate and speak a consent phrase, uses automatic speech recognition to verify it, and only then permits a text-to-speech voice-cloning model to run. The design makes consent a traceable and auditable prerequisite for AI action while addressing both the risks and beneficial uses of voice cloning.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## FastRTC: The Real-Time Communication Library for Python

DevFeed: [FastRTC: The Real-Time Communication Library for Python](<https://devfeed.tech/articles/fastrtc-the-real-time-communication-library-for-python-7194.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/fastrtc>)

Author: Freddy Boulton; Abubakar Abid

Published: 2025-02-25T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [FastAPI](<https://devfeed.tech/topics/fastapi.md>), [Python](<https://devfeed.tech/topics/python.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [WebRTC](<https://devfeed.tech/topics/webrtc.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [gradio](<https://devfeed.tech/topics/gradio.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [apis](<https://devfeed.tech/tags/apis.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [gradio](<https://devfeed.tech/tags/gradio.md>), [llm](<https://devfeed.tech/tags/llm.md>), [python](<https://devfeed.tech/tags/python.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [video](<https://devfeed.tech/tags/video.md>), [voice](<https://devfeed.tech/tags/voice.md>), [webrtc](<https://devfeed.tech/tags/webrtc.md>), [websocket](<https://devfeed.tech/tags/websocket.md>)

### AI overview

FastRTC is introduced as a Python library for building real-time audio and video AI applications. The article demonstrates core features including voice detection and turn taking, a WebRTC-enabled Gradio interface, phone access, WebRTC and WebSocket support, FastAPI integration, and speech-to-text and text-to-speech utilities.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Evaluating Audio Reasoning with Big Bench Audio

DevFeed: [Evaluating Audio Reasoning with Big Bench Audio](<https://devfeed.tech/articles/evaluating-audio-reasoning-with-big-bench-audio-7128.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/big-bench-audio-release>)

Author: Micah Hill-Smith; George Cameron

Published: 2024-12-20T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [arena](<https://devfeed.tech/tags/arena.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [testing](<https://devfeed.tech/tags/testing.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

Big Bench Audio is a 1,000-question dataset for evaluating audio language-model reasoning across speech and text configurations. Initial results report a speech reasoning gap for GPT-4o between text-only and speech-to-speech performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Build Your Own AI Voice Assistant

DevFeed: [Build Your Own AI Voice Assistant](<https://devfeed.tech/articles/build-your-own-ai-voice-assistant-5756.md>)

Original publisher: [Read original article](<https://neon.com/blog/pulse-voice-assistant>)

Author: Rishi Raj Jain

Published: 2024-12-19T02:28:26Z

Content type: tutorial

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Val Town](<https://devfeed.tech/topics/val-town.md>), [API](<https://devfeed.tech/topics/api.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Front end](<https://devfeed.tech/topics/frontend.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-voice-assistant](<https://devfeed.tech/tags/ai-voice-assistant.md>), [api](<https://devfeed.tech/tags/api.md>), [developers](<https://devfeed.tech/tags/developers.md>), [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [product](<https://devfeed.tech/tags/product.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [tools](<https://devfeed.tech/tags/tools.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

A developer guide to building Pulse, an AI voice assistant, using ElevenLabs for real-time voice responses, Neon for persistent data storage, and Next.js for the user interface, API requests, and real-time updates.

### Source excerpt

It's very likely that you, like many developers today, are exploring how to build AI apps using the hottest tools--and one of them is definitely ElevenLabs. That's why we've put together a guide that teaches you how to create an app like this--an AI voice assistant we've named Puls...

## Deploying Speech-to-Speech on Hugging Face

DevFeed: [Deploying Speech-to-Speech on Hugging Face](<https://devfeed.tech/articles/deploying-speech-to-speech-on-hugging-face-7462.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/s2s_endpoint>)

Author: Andres Marafioti; Derek Thomas; Diego Maniloff; Eustache Le Bihan

Published: 2024-10-22T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [docker](<https://devfeed.tech/tags/docker.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

A step-by-step guide to packaging and deploying Hugging Face's Speech-to-Speech pipeline on Inference Endpoints using a custom Docker image.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Implementing Multiplatform Kotlin Mobile

DevFeed: [Implementing Multiplatform Kotlin Mobile](<https://devfeed.tech/articles/implementing-multiplatform-kotlin-mobile-39207.md>)

Original publisher: [Read original article](<https://kt.academy/article/ak-kmp-kmm>)

Published: 2023-09-25T00:15:00Z

Content type: tutorial

Language: en

Sources: [Kt. Academy](<https://devfeed.tech/sources/kt-academy.md>)

Topics: [Kotlin Multiplatform](<https://devfeed.tech/topics/kotlin-multiplatform.md>), [multiplatform](<https://devfeed.tech/topics/multiplatform.md>), [Mobile](<https://devfeed.tech/topics/mobile.md>), [Android](<https://devfeed.tech/topics/android.md>), [iOS](<https://devfeed.tech/topics/ios.md>), [Development](<https://devfeed.tech/topics/development.md>), [business logic](<https://devfeed.tech/topics/business-logic.md>)

Tags: [android](<https://devfeed.tech/tags/android.md>), [business-logic](<https://devfeed.tech/tags/business-logic.md>), [coroutine](<https://devfeed.tech/tags/coroutine.md>), [dependencies](<https://devfeed.tech/tags/dependencies.md>), [ios](<https://devfeed.tech/tags/ios.md>), [kotlin](<https://devfeed.tech/tags/kotlin.md>), [logic](<https://devfeed.tech/tags/logic.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [multiplatform](<https://devfeed.tech/tags/multiplatform.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [workshop-learning-programming](<https://devfeed.tech/tags/workshop-learning-programming.md>)

### AI overview

This tutorial explains how Kotlin Multiplatform Mobile can share business logic between Android and iOS applications. It covers shared modules, generated platform libraries, ViewModel expect/actual classes, coroutine scopes, and interfaces for platform-specific dependencies such as text-to-speech.

### Source excerpt

How in Kotlin we can implement Android and iOS projects with shared logic.

## Everyday Accessibility

DevFeed: [Everyday Accessibility](<https://devfeed.tech/articles/everyday-accessibility-9382.md>)

Original publisher: [Read original article](<https://a11yproject.com/posts/everyday-accessibility/>)

Author: Michele A Williams; PhD

Published: 2021-06-15T00:00:00Z

Content type: article

Language: en

Sources: [The A11Y Project](<https://devfeed.tech/sources/the-a11y-project.md>)

Topics: [Accessibility](<https://devfeed.tech/topics/accessibility.md>), [alt text](<https://devfeed.tech/topics/alt-text.md>)

Tags: [accessibility](<https://devfeed.tech/tags/accessibility.md>), [alt-text](<https://devfeed.tech/tags/alt-text.md>), [images](<https://devfeed.tech/tags/images.md>), [readability](<https://devfeed.tech/tags/readability.md>), [social-media](<https://devfeed.tech/tags/social-media.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [writing](<https://devfeed.tech/tags/writing.md>)

### AI overview

This practical accessibility article explains how to make social media, digital, and multimedia content more inclusive. It recommends capitalizing words in hashtags, providing written text alongside images of text, adding concise alt text to images, limiting disruptive emoji use, and using built-in document styling features so assistive technologies can interpret content correctly.

### Source excerpt

Even with the best of intentions, you may be making decisions that exclude disabled users. Below are practical, intentional steps you can start today to create and share more accessible social media, digital, and multimedia content. Watch this post as a video Social Media Posts Social media is central to the world's conversations, so ensure you're not leaving anyone out. Use capital letters in hashtags Punctuation helps text-to-speech tools like screen readers speak more accurately, and helps with readability overall. This includes in hashtags, where capital letters help distinguish words. Key Point: Make it a habit to #UseCapitalLetters rather than #hardtoreadlowercase. Don't share just images of text Images of text (such as your favorite quote or saying) can be difficult for people with visual impairments to read and don't allow people with reading difficulties to leverage text-to-speech read-aloud features. When posting these kinds of images, also write out the text in your post. Key Point: Typing out the text in an image makes the content more accessible. Use "alt text" features to describe images Pictures are significant in social media, but also inaccessible to visually impaired readers unless they have a text-based description. When uploading, find the "image description" or "alternative/alt text" feature in the software to write a description. For sample instructions, review Alt Text for Social Media. When adding descriptions, don't write every detail, just enough to ensure everyone can understand your post and share in the image's meaning. For examples, read Scott Vinkle's Considerations when writing alt text. Key Point: Features like Twitter's "Add description" allow adding "alt text" for visually impaired readers. Limit your emojis Emojis are great for emphasis and playfulness, but adding too many sounds cluttered with text-to-speech software. Also, when placed unconventionally (such as throughout a sentence) this is distracting and hard to read for those

## openHAB 2.4 Release

DevFeed: [openHAB 2.4 Release](<https://devfeed.tech/articles/openhab-2-4-release-16718.md>)

Original publisher: [Read original article](<https://openhab.org/blog/2018-12-17-openhab-2-4-release.html>)

Published: 2018-12-16T23:57:45Z

Content type: release

Language: en

Sources: [openHAB](<https://devfeed.tech/sources/openhab.md>)

Topics: [releases](<https://devfeed.tech/topics/releases.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Raspberry Pi](<https://devfeed.tech/topics/raspberry-pi.md>)

Tags: [configuration](<https://devfeed.tech/tags/configuration.md>), [feature](<https://devfeed.tech/tags/feature.md>), [linux](<https://devfeed.tech/tags/linux.md>), [raspberry-pi](<https://devfeed.tech/tags/raspberry-pi.md>), [release](<https://devfeed.tech/tags/release.md>), [release-notes](<https://devfeed.tech/tags/release-notes.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [ui](<https://devfeed.tech/tags/ui.md>)

### AI overview

The openHAB 2.4 release introduces 34 new add-ons and updates the core runtime and existing add-ons. It highlights Profiles for reducing configuration complexity, new text-to-speech options including Google Cloud TTS and Pico TTS, and other improvements.

### Source excerpt

As for the past few years, the openHAB maintainers have decided to do a new release just in time for the holiday season, which is the busiest time of the year in our community. So what could be more welcome than a brand new stable release?

## Developer Contest: SoundCloud + Acapela Group Mashup

DevFeed: [Developer Contest: SoundCloud + Acapela Group Mashup](<https://devfeed.tech/articles/developer-contest-soundcloud-acapela-group-mashup-2016.md>)

Original publisher: [Read original article](<https://developers.soundcloud.com/blog//contest-soundcloud-and-acapela-group>)

Published: 2012-04-04T00:00:00Z

Content type: article

Language: en

Sources: [SoundCloud Backstage Blog](<https://devfeed.tech/sources/soundcloud-backstage-blog.md>)

Topics: [SDKs](<https://devfeed.tech/topics/sdks.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [coding-community](<https://devfeed.tech/topics/coding-community.md>), [API](<https://devfeed.tech/topics/api.md>), [App](<https://devfeed.tech/topics/app.md>), [Website](<https://devfeed.tech/topics/website.md>)

Tags: [announcements](<https://devfeed.tech/tags/announcements.md>), [api](<https://devfeed.tech/tags/api.md>), [app](<https://devfeed.tech/tags/app.md>), [audio](<https://devfeed.tech/tags/audio.md>), [community](<https://devfeed.tech/tags/community.md>), [creative](<https://devfeed.tech/tags/creative.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-community](<https://devfeed.tech/tags/developer-community.md>), [free](<https://devfeed.tech/tags/free.md>), [partner](<https://devfeed.tech/tags/partner.md>), [sdks](<https://devfeed.tech/tags/sdks.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voices](<https://devfeed.tech/tags/voices.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

SoundCloud and Acapela Group announce a developer contest inviting participants to combine SoundCloud's audio platform with Acapela's text-to-speech services and SDKs. Developers can create creative and accessible apps, with submissions due May 15th and prizes offered to the top three entries.

### Source excerpt

SoundCloud is teaming up with Acapela Group for our first Developer Contest. Acapela Group offers amazing text to speech solutions and have a variety of SDKs so you can write apps that create sound files using one of their voices (including hip-hop and country!). Take a listen to some sample voices. We're calling on our developer community to mashup SoundCloud and Acapela. Show us what you can do using a text to speech service with the best audio content platform on the web. Maybe you want to have your email read (or sung!) to you layered on some instrumental or drum tracks. We think there are loads of options for making interesting and accessible apps using these two services.