Seedance 2.5:ByteDanceがビデオ生成を30秒のシングルショットに拡張

モデルとツール yesterday8ブックマークに追加

Seedance 2.5:ByteDanceがビデオ生成を30秒のシングルショットに拡張
イラスト : Léa Fontaine

Seedチームは、30秒のシングルショット生成を実現し、マルチモーダルなリファレンス(30枚の画像 + 10本の動画 + 10本の音声)とタイムスタンプ編集を導入しました。

平易な言葉で

ByteDanceはSeedance 2.5をリリース、30秒の動画を1回のリクエストで生成可能(従来の4秒単位の複数クリップを組み合わせる方法ではなく)。また、同一プロンプト内で最大30枚の画像、10本の動画、10本の音声をリファレンスとして受け付ける。

事実

Pandaily(2026年7月31日)によると、Seedチームは以下の3つの新機能を発表:

  • 30秒シングルショット生成(従来のマルチクリップ方式に代わり)。
  • マルチラウンド拡張:物語的に一貫した動画の延長。
  • タイムスタンプ編集:生成された動画の特定の瞬間をターゲットに編集。

当社の分析

最大の進歩は「ロングナラティブ」:4〜8秒から30秒のシングルショットへの移行は、動画モデルにとって真の壁であり、キャラクターの逸脱、シーンの一貫性、計算コストは時間と共に爆発的に増大する。また、マルチモーダルリファレンス(30枚の画像 + 10本の動画 + 10本の音声)により、純粋なテキストtoビデオから制作支援へと戦いの場が移行:スタジオが事前にアセットを注入し、モデルはレールに沿って動作する。

現時点では第三者ベンチマークは存在しない。注目すべき点:30秒におけるキャラクターの一貫性とリップシンクの音声。

要注目

  • Artificial Analysisベンチマーク(Kling、Sora、MiniMax Hailuoとの比較)。
  • 価格設定(動画生成は現在最も高額な分野)。
  • APIの公開・クローズドソースの有無。
リソース

本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。

編集部について
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
この記事は役に立ちましたか?

8 人がこの記事を評価しました

いいね
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
シェア:
コメント (8)

ログインして議論に参加しましょう。

Alex 2 02 Aug 2026 · 14:01

30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.

Alex 02 Aug 2026 · 12:34

Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.

ph1lippe_m 02 Aug 2026 · 11:15

30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?

J.P.R. 2 02 Aug 2026 · 13:34

ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?

sandrine.b 02 Aug 2026 · 11:08

Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.

FoodieFiona 2 02 Aug 2026 · 13:43

Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.

BookWorm88 02 Aug 2026 · 10:57

30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?

ArtLover88 02 Aug 2026 · 13:30

Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?

MusicFanatic 02 Aug 2026 · 10:53

The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.

HistoryBuff 02 Aug 2026 · 10:38

Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?

SkepticSam 02 Aug 2026 · 10:34

30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
テーマ
探索
インフォメーション