# Thinking Machines proposes interaction models for continuous, real-time multimodal interaction

DevFeed: [Thinking Machines proposes interaction models for continuous, real-time multimodal interaction](<https://devfeed.tech/articles/why-interaction-models-are-the-next-ai-frontier-18138.md>)

Original publisher: [Read original article](<https://hungrymindsdev.substack.com/p/why-interaction-models-are-the-next>)

Author: Alexandre Zajac

Published: 2026-07-06T15:31:47Z

Content type: opinion

Language: en

Sources: [Hungry Minds](<https://devfeed.tech/sources/hungry-minds.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [sglang](<https://devfeed.tech/topics/sglang.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [model](<https://devfeed.tech/tags/model.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

## AI overview

The article describes Thinking Machines' interaction-model approach, which replaces discrete conversational turns with time-aligned 200-millisecond micro-turns so input and output can occur concurrently. It also outlines lightweight audio and video encoders, coordination between a fast interaction model and a slower reasoning model, and streaming inference work in SGLang. The claimed applications include live translation, real-time commentary, mid-sentence corrections, and video-triggered responses.

## Source excerpt

PLUS: Reddit's anti-spam internals 👨💻, p99 0ms autocomplete ⚡, YouTube leaks creators' videos 🚨