# How Together AI built the world's fastest speech-to-text stack

DevFeed: [How Together AI built the world's fastest speech-to-text stack](<https://devfeed.tech/articles/how-together-ai-built-the-world-s-fastest-speech-to-text-stack-80187.md>)

Original publisher: [Read original article](<https://www.together.ai/blog/how-together-ai-built-the-worlds-fastest-speech-to-text-stack>)

Author: Sebastien Beurnier

Published: 2026-05-29T00:00:00Z

Content type: article

Language: en

Sources: [Together.ai](<https://devfeed.tech/sources/together-ai.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Integration testing](<https://devfeed.tech/topics/integration-testing.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [Service fabric](<https://devfeed.tech/topics/service-fabric.md>), [On-device AI](<https://devfeed.tech/topics/on-device-ai.md>)

Tags: [batch](<https://devfeed.tech/tags/batch.md>), [compile](<https://devfeed.tech/tags/compile.md>), [concurrency](<https://devfeed.tech/tags/concurrency.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [cuda-graphs](<https://devfeed.tech/tags/cuda-graphs.md>), [domain-sockets](<https://devfeed.tech/tags/domain-sockets.md>), [epoll](<https://devfeed.tech/tags/epoll.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-text](<https://devfeed.tech/tags/speech-to-text.md>), [systems](<https://devfeed.tech/tags/systems.md>)

## AI overview

Together describes how it optimized production speech-to-text serving across the full system, including profile-tuned TensorRT encoder execution, GPU-side decoder control flow, lower-copy audio processing, evented streaming I/O, and Python garbage collection. The article says these changes improved throughput and latency, including a 2-3x faster decoder and the removal of periodic p95 latency spikes.

## Source excerpt

Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.