
Seed团队将单次生成时间推进至30秒,支持多模态参考(30张图像 + 10段视频 + 10段音频)及时间戳编辑功能。
字节跳动发布 Seedance 2.5,一种视频生成模型,能够在单次请求中生成30秒视频——而非通过拼接多个4秒片段。它还支持在同一提示中使用最多30张图像、10段视频和10段音频作为参考。
Pandaily(2026年7月31日)报道了Seed团队的三项新功能:
最大的突破在于“长叙事”:从4-8秒提升至30秒单次生成是视频模型的真正瓶颈——角色偏移、场景一致性及计算成本随时长激增。此外,多模态参考输入(30张图+10段视频+10段音频)将竞争焦点从纯文本转向视频生成:工作室可预注入素材,模型则保持在可控范围内。
目前尚无第三方基准测试。我们关注的重点:30秒内角色一致性及音频唇同步。
本文由人工智能撰写,并经人工编辑审核。
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.