Models & Tools yesterday8Add to bookmarks

The Seed team pushes single-shot generation to 30 seconds, with multi-modal references (30 images + 10 videos + 10 audios) and timestamp editing.
ByteDance releases Seedance 2.5, a video model capable of generating 30 seconds in a single request—not by stitching multiple 4-second clips. It also accepts up to 30 images, 10 videos, and 10 audio files as references in a single prompt.
Pandaily (July 31, 2026) reports three new features from the Seed team:
The big leap is in long-form narrative: going from 4–8 seconds to 30 seconds in a single shot is the real wall for video models—the drift of characters, scene consistency, and compute costs explode with duration. The multi-modal reference input (30 images + 10 videos + 10 audio files) also shifts the battle from pure text-to-video to assisted production: the studio pre-injects its assets, and the model stays on track.
No third-party benchmarks yet. What we’re watching: character consistency at 30 seconds and audio lipsync.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.