Modelle & Werkzeuge 3 min ago7Zu Lesezeichen hinzufügen

Der Seed-Team treibt die Generation auf 30-Sekunden-Single-Shots voran, mit multimodalen Referenzen (30 Bilder + 10 Videos + 10 Audios) und Zeitstempel-Bearbeitung.
ByteDance stellt Seedance 2.5 vor, ein Video-Modell, das in einer einzigen Anfrage 30 Sekunden generieren kann – nicht durch das Zusammenfügen mehrerer 4-Sekunden-Clips. Es akzeptiert zudem bis zu 30 Bilder, 10 Videos und 10 Audios als Referenzen in einer einzigen Eingabeaufforderung.
Pandaily (31. Juli 2026) berichtet über drei neue Funktionen des Seed-Teams:
Der große Schritt ist das long-narrative: Von 4–8 Sekunden auf 30 Sekunden in einem einzigen Durchgang zu kommen, ist die eigentliche Hürde für Video-Modelle – die Abweichung der Charaktere, die Szenen-Kohärenz und die Rechenkosten explodieren mit der Dauer. Die multimodale Referenzierung (30 Bilder + 10 Videos + 10 Audios als Eingabe) verschiebt zudem den Fokus vom reinen Text-zu-Video hin zur assistierten Produktion: Das Studio injiziert seine Assets vorab, das Modell bleibt auf Kurs.
Noch keine unabhängigen Benchmarks in diesem Stadium. Was wir beobachten: die Charakter-Kohärenz über 30 Sekunden und das Lipsync mit Audio.
Artikel von künstlicher Intelligenz erstellt, unter menschlicher redaktioneller Kontrolle geprüft.
Melden Sie sich an, um an der Diskussion teilzunehmen.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.