Modelos y Herramientas yesterday8Añadir a favoritos

El equipo Seed lleva la generación a 30 segundos en un solo disparo, con referencias multimodales (30 imágenes + 10 videos + 10 audios) y edición con marca de tiempo.
ByteDance lanza Seedance 2.5, un modelo de vídeo capaz de generar 30 segundos en una sola solicitud —no agregando varios clips de 4 s—. También acepta hasta 30 imágenes, 10 vídeos y 10 audios como referencias en una misma entrada.
Pandaily (31 de julio de 2026) reporta tres novedades del equipo Seed:
El salto largo es el narrativa larga: pasar de 4-8 s a 30 s en un solo paso es el verdadero muro de los modelos de vídeo —la deriva de los personajes, la coherencia de escena y el coste computacional se disparan con la duración—. La referencia multimodal (30 img + 10 vid + 10 audio en entrada) desplaza además la batalla del puro text-to-video hacia la producción asistida: el estudio preinyecta sus assets y el modelo se mantiene en los carriles.
Sin benchmarks de terceros por ahora. Lo que observamos: la coherencia de personajes a 30 s y el sincronismo labial con el audio.
Artículo producido por inteligencia artificial, revisado bajo control editorial humano.
Inicia sesión para unirte a la conversación.
30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.