Models & Tools il y a 34 min5Ajouter aux favoris

ByteDance launched SeedRealtime, a full-duplex audio-video model for real-time AI interaction. Unlike Seedance (long-form video generation), this is conversational AI that listens and speaks at the same time - with live video input.
In plain terms: ByteDance released an AI model that processes audio and video input and output at the same time, with no switching between listening and speaking modes. It's the architecture real-time conversation actually requires.
ByteDance launched SeedRealtime, a full-duplex audio-visual model for real-time AI interaction, according to Tech in Asia. Full-duplex means simultaneous bidirectional streaming: input and output audio and video run in parallel rather than alternating. This is distinct from Seedance (ByteDance's long-form video generation model) - SeedRealtime targets conversational and live-interaction use cases.
Full-duplex is the hard part. Most voice AI systems are half-duplex: they accept input, generate a response, then switch back to listening. The latency and interruption awkwardness in current voice assistants comes directly from this architecture. SeedRealtime enters a space where OpenAI (GPT-4o Advanced Voice), Google (Gemini Live), and ElevenLabs have built early positions - but adds multimodal video input, which none of those support natively in real-time. The commercial use cases are clear: live translation, real-time customer service, AI companion apps. The competitive signal is clearer: ByteDance is not conceding the real-time multimodal layer to US labs.
Latency benchmarks and API pricing at scale. Full-duplex is achievable in a demo; maintaining it cheaply across concurrent sessions at production scale is the actual engineering challenge.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
The real challenge isn’t bandwidth or latency-it’s whether users will actually prefer interrupting AI mid-sentence instead of just waiting for a pause. Humans adapt quickly to turn-taking, but machines?
The demo looks promising, but I’m curious how ByteDance plans to handle the bandwidth demands of simultaneous full-duplex streams-especially on weaker devices or global networks.
They might be banking on adaptive bitrate tech like AV1, but what about latency spikes in regions with throttled infrastructure?
Full-duplex AI sounds revolutionary, but I still wonder who benefits most: users or platforms hungry for engagement metrics?
Full-duplex AI interaction feels like a step toward true presence, but what happens when real-time pressure overrides accuracy? Overfitting on latency might sacrifice nuance.
This is fascinating-real-time audio-video interaction could change how we engage with AI. But I wonder about the latency issues in practical applications.
Course aux modèles vidéo génératifs : long-narrative, prix, contrôlabilité