Models & Tools il y a 4 min7Ajouter aux favoris

Le Seed team pousse la génération à 30 s single-shot, avec références multi-modales (30 images + 10 vidéos + 10 audios) et édition timestamp.
ByteDance sort Seedance 2.5, un modèle vidéo capable de générer 30 secondes en une seule requête - pas en agrégeant plusieurs clips de 4 s. Il accepte aussi jusqu'à 30 images, 10 vidéos et 10 audios comme références dans un même prompt.
Pandaily (31 juillet 2026) rapporte trois briques nouvelles du Seed team :
Le pas long, c'est le long-narrative : passer de 4-8 s à 30 s single-shot est le vrai mur des modèles vidéo - la dérive des personnages, la cohérence de scène et le coût compute explosent avec la durée. Le multi-modal reference (30 img + 10 vid + 10 audio en input) déplace en outre la lutte du pur text-to-video vers la production assistée : le studio pré-injecte ses assets, le modèle reste dans les rails.
Zéro benchmark tiers à ce stade. Ce qu'on regarde : la cohérence de personnage à 30 s et le lipsync audio.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.