Models & Tools 58 min ago5Add to bookmarks

ByteDance launched SeedRealtime, a full-duplex audio-video model for real-time AI interaction. Unlike Seedance (long-form video generation), this is conversational AI that listens and speaks at the same time—with live video input.
In plain terms: ByteDance released an AI model that processes audio and video input and output at the same time, with no switching between listening and speaking modes. It's the architecture real-time conversation actually requires.
ByteDance launched SeedRealtime, a full-duplex audio-visual model for real-time AI interaction, according to Tech in Asia. Full-duplex means simultaneous bidirectional streaming: input and output audio and video run in parallel rather than alternating. This is distinct from Seedance (ByteDance's long-form video generation model) - SeedRealtime targets conversational and live-interaction use cases.
Full-duplex is the hard part. Most voice AI systems are half-duplex: they accept input, generate a response, then switch back to listening. The latency and interruption awkwardness in current voice assistants comes directly from this architecture. SeedRealtime enters a space where OpenAI (GPT-4o Advanced Voice), Google (Gemini Live), and ElevenLabs have built early positions - but adds multimodal video input, which none of those support natively in real-time. The commercial use cases are clear: live translation, real-time customer service, AI companion apps. The competitive signal is clearer: ByteDance is not conceding the real-time multimodal layer to US labs.
Latency benchmarks and API pricing at scale. Full-duplex is achievable in a demo; maintaining it cheaply across concurrent sessions at production scale is the actual engineering challenge.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
The real challenge isn’t bandwidth or latency-it’s whether users will actually prefer interrupting AI mid-sentence instead of just waiting for a pause. Humans adapt quickly to turn-taking, but machines?
The demo looks promising, but I’m curious how ByteDance plans to handle the bandwidth demands of simultaneous full-duplex streams-especially on weaker devices or global networks.
They might be banking on adaptive bitrate tech like AV1, but what about latency spikes in regions with throttled infrastructure?
Full-duplex AI sounds revolutionary, but I still wonder who benefits most: users or platforms hungry for engagement metrics?
Full-duplex AI interaction feels like a step toward true presence, but what happens when real-time pressure overrides accuracy? Overfitting on latency might sacrifice nuance.
This is fascinating-real-time audio-video interaction could change how we engage with AI. But I wonder about the latency issues in practical applications.
Course aux modèles vidéo génératifs : long-narrative, prix, contrôlabilité