
ByteDanceは、リアルタイムAIインタラクション向けの全二重オーディオビデオモデルであるSeedRealtimeをリリースしました。Seedance(長編動画生成)とは異なり、これは会話型AIであり、同時に聞き取り・発話が可能で、ライブビデオ入力にも対応しています。
簡単に言うと: ByteDanceは、音声と映像の入出力を同時に処理し、聞くモードと書くモードを切り替える必要のないAIモデルをリリースしました。これは、リアルタイム会話に実際に必要なアーキテクチャです。
ByteDanceはTech in Asiaによると、リアルタイムAIインタラクション向けの完全二重音声映像モデル「SeedRealtime」をリリースしました。完全二重とは、双方向の同時ストリーミングを意味します。入出力の音声と映像が並行して動作し、交互に切り替わることはありません。これは、ByteDanceの長編動画生成モデル「Seedance」とは異なります。SeedRealtimeは会話やライブインタラクション向けに設計されています。
完全二重は難しい部分です。ほとんどの音声AIシステムは半二重です。入力を受け付け、応答を生成し、再び聞き取りモードに戻ります。現在の音声アシスタントにおけるレイテンシーや中断のぎこちなさは、このアーキテクチャに直接起因しています。SeedRealtimeは、OpenAI(GPT-4o Advanced Voice)、Google(Gemini Live)、ElevenLabsが先行している分野に参入しますが、それらがネイティブでサポートしていないリアルタイムのマルチモーダル映像入力を追加しています。商用利用のユースケースは明確です:ライブ翻訳、リアルタイムカスタマーサービス、AIコンパニオンアプリ。競争のシグナルはより明確です:ByteDanceはリアルタイムマルチモーダル層を米国の研究所に譲るつもりはありません。
レイテンシーベンチマークとAPIの大規模価格設定。完全二重はデモで実現可能ですが、大規模な同時セッションで安価に維持することが実際の技術的課題です。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
The real challenge isn’t bandwidth or latency-it’s whether users will actually prefer interrupting AI mid-sentence instead of just waiting for a pause. Humans adapt quickly to turn-taking, but machines?
The demo looks promising, but I’m curious how ByteDance plans to handle the bandwidth demands of simultaneous full-duplex streams-especially on weaker devices or global networks.
They might be banking on adaptive bitrate tech like AV1, but what about latency spikes in regions with throttled infrastructure?
Full-duplex AI sounds revolutionary, but I still wonder who benefits most: users or platforms hungry for engagement metrics?
Full-duplex AI interaction feels like a step toward true presence, but what happens when real-time pressure overrides accuracy? Overfitting on latency might sacrifice nuance.
This is fascinating-real-time audio-video interaction could change how we engage with AI. But I wonder about the latency issues in practical applications.
Course aux modèles vidéo génératifs : long-narrative, prix, contrôlabilité