Models & Tools 14 min ago7Add to bookmarks

The Seed team pushes single-shot generation to 30 seconds, with multi-modal references (30 images + 10 videos + 10 audios) and timestamp editing.
ByteDance releases Seedance 2.5, a video model capable of generating 30 seconds in a single request—not by stitching multiple 4-second clips. It also accepts up to 30 images, 10 videos, and 10 audio files as references in a single prompt.
Pandaily (July 31, 2026) reports three new features from the Seed team:
The big leap is in long-form narrative: going from 4–8 seconds to 30 seconds in a single shot is the real wall for video models—the drift of characters, scene consistency, and compute costs explode with duration. The multi-modal reference input (30 images + 10 videos + 10 audio files) also shifts the battle from pure text-to-video to assisted production: the studio pre-injects its assets, and the model stays on track.
No third-party benchmarks yet. What we’re watching: character consistency at 30 seconds and audio lipsync.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.