
Seedチームは、30秒のシングルショット生成を実現し、マルチモーダルなリファレンス(30枚の画像 + 10本の動画 + 10本の音声)とタイムスタンプ編集を導入しました。
ByteDanceはSeedance 2.5をリリース、30秒の動画を1回のリクエストで生成可能(従来の4秒単位の複数クリップを組み合わせる方法ではなく)。また、同一プロンプト内で最大30枚の画像、10本の動画、10本の音声をリファレンスとして受け付ける。
Pandaily(2026年7月31日)によると、Seedチームは以下の3つの新機能を発表:
最大の進歩は「ロングナラティブ」:4〜8秒から30秒のシングルショットへの移行は、動画モデルにとって真の壁であり、キャラクターの逸脱、シーンの一貫性、計算コストは時間と共に爆発的に増大する。また、マルチモーダルリファレンス(30枚の画像 + 10本の動画 + 10本の音声)により、純粋なテキストtoビデオから制作支援へと戦いの場が移行:スタジオが事前にアセットを注入し、モデルはレールに沿って動作する。
現時点では第三者ベンチマークは存在しない。注目すべき点:30秒におけるキャラクターの一貫性とリップシンクの音声。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.