ByteDanceのSeedRealtimeは、全二重の音声と映像を同時に実行します

継続中のトピック : Course aux modèles vidéo génératifs : long-narrative, prix, contrôlabilité· パート 2/2

モデルとツール 57 min ago5ブックマークに追加

ByteDanceのSeedRealtimeは、全二重の音声と映像を同時に実行します
イラスト : Léa Fontaine

ByteDanceは、リアルタイムAIインタラクション向けの全二重オーディオビデオモデルであるSeedRealtimeをリリースしました。Seedance(長編動画生成)とは異なり、これは会話型AIであり、同時に聞き取り・発話が可能で、ライブビデオ入力にも対応しています。

簡単に言うと: ByteDanceは、音声と映像の入出力を同時に処理し、聞くモードと書くモードを切り替える必要のないAIモデルをリリースしました。これは、リアルタイム会話に実際に必要なアーキテクチャです。

事実

ByteDanceはTech in Asiaによると、リアルタイムAIインタラクション向けの完全二重音声映像モデル「SeedRealtime」をリリースしました。完全二重とは、双方向の同時ストリーミングを意味します。入出力の音声と映像が並行して動作し、交互に切り替わることはありません。これは、ByteDanceの長編動画生成モデル「Seedance」とは異なります。SeedRealtimeは会話やライブインタラクション向けに設計されています。

私たちの見解

完全二重は難しい部分です。ほとんどの音声AIシステムは半二重です。入力を受け付け、応答を生成し、再び聞き取りモードに戻ります。現在の音声アシスタントにおけるレイテンシーや中断のぎこちなさは、このアーキテクチャに直接起因しています。SeedRealtimeは、OpenAI(GPT-4o Advanced Voice)、Google(Gemini Live)、ElevenLabsが先行している分野に参入しますが、それらがネイティブでサポートしていないリアルタイムのマルチモーダル映像入力を追加しています。商用利用のユースケースは明確です:ライブ翻訳、リアルタイムカスタマーサービス、AIコンパニオンアプリ。競争のシグナルはより明確です:ByteDanceはリアルタイムマルチモーダル層を米国の研究所に譲るつもりはありません。

見るべきポイント

レイテンシーベンチマークとAPIの大規模価格設定。完全二重はデモで実現可能ですが、大規模な同時セッションで安価に維持することが実際の技術的課題です。

本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。

編集部について
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
この記事は役に立ちましたか?

5 人がこの記事を評価しました

いいね
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
シェア:
コメント (5)

ログインして議論に参加しましょう。

TechSavvy47 05 Aug 2026 · 13:05

The real challenge isn’t bandwidth or latency-it’s whether users will actually prefer interrupting AI mid-sentence instead of just waiting for a pause. Humans adapt quickly to turn-taking, but machines?

Dr. Emily 05 Aug 2026 · 12:42

The demo looks promising, but I’m curious how ByteDance plans to handle the bandwidth demands of simultaneous full-duplex streams-especially on weaker devices or global networks.

J.P.R. 05 Aug 2026 · 15:05

They might be banking on adaptive bitrate tech like AV1, but what about latency spikes in regions with throttled infrastructure?

Critique42 05 Aug 2026 · 12:25

Full-duplex AI sounds revolutionary, but I still wonder who benefits most: users or platforms hungry for engagement metrics?

J.P.R. 2 05 Aug 2026 · 12:20

Full-duplex AI interaction feels like a step toward true presence, but what happens when real-time pressure overrides accuracy? Overfitting on latency might sacrifice nuance.

curio_usa 05 Aug 2026 · 12:18

This is fascinating-real-time audio-video interaction could change how we engage with AI. But I wonder about the latency issues in practical applications.

トピックの経過

Course aux modèles vidéo génératifs : long-narrative, prix, contrôlabilité

  1. 1Seedance 2.5:ByteDanceがビデオ生成を30秒のシングルショットに拡張02/08/2026
  2. 2ByteDanceのSeedRealtimeは、全二重の音声と映像を同時に実行します05/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
テーマ
探索
インフォメーション