
ByteDance 推出了 SeedRealtime,一款用于实时 AI 交互的全双工音视频模型。与 Seedance(长视频生成)不同,这是一种对话式 AI,能够同时倾听和说话——并支持实时视频输入。
简言之: 字节跳动发布了一款AI模型,可同时处理音频和视频的输入与输出,无需在收听和说话模式间切换。这正是实时对话所需的架构。
据Tech in Asia报道,字节跳动推出了SeedRealtime,一款用于实时AI交互的全双工音视频模型。全双工意味着同时进行双向流式传输:音频和视频的输入与输出并行运行,而非交替进行。这与Seedance(字节跳动的长视频生成模型)不同——SeedRealtime针对对话和实时互动场景。
全双工是难点所在。大多数语音AI系统都是半双工:先接受输入、生成响应,再切换回收听模式。当前语音助手的延迟和尴尬中断直接源于此架构。SeedRealtime进入了一个OpenAI(GPT-4o高级语音)、Google(Gemini Live)和ElevenLabs已建立早期布局的领域——但增加了视频输入的多模态功能,而这些竞争对手在实时场景中均未原生支持。商业用例清晰:实时翻译、实时客服、AI伴侣应用。竞争信号更明确:字节跳动并未将实时多模态层让给美国实验室。
延迟基准和大规模API定价。全双工在演示中可行;但在生产规模下以低成本维持并发会话的全双工,才是真正的工程挑战。
本文由人工智能撰写,并经人工编辑审核。
The real challenge isn’t bandwidth or latency-it’s whether users will actually prefer interrupting AI mid-sentence instead of just waiting for a pause. Humans adapt quickly to turn-taking, but machines?
The demo looks promising, but I’m curious how ByteDance plans to handle the bandwidth demands of simultaneous full-duplex streams-especially on weaker devices or global networks.
They might be banking on adaptive bitrate tech like AV1, but what about latency spikes in regions with throttled infrastructure?
Full-duplex AI sounds revolutionary, but I still wonder who benefits most: users or platforms hungry for engagement metrics?
Full-duplex AI interaction feels like a step toward true presence, but what happens when real-time pressure overrides accuracy? Overfitting on latency might sacrifice nuance.
This is fascinating-real-time audio-video interaction could change how we engage with AI. But I wonder about the latency issues in practical applications.