Modelos e Ferramentas yesterday8Adicionar aos favoritos

O Seed team avança para a geração de 30 segundos em um único disparo, com referências multimodais (30 imagens + 10 vídeos + 10 áudios) e edição por carimbo de data/hora.
ByteDance lançou o Seedance 2.5, um modelo de vídeo capaz de gerar 30 segundos em um único pedido — não agregando vários clipes de 4 s. Ele também aceita até 30 imagens, 10 vídeos e 10 áudios como referências em um mesmo prompt.
O Pandaily (31 de julho de 2026) relata três novos recursos da equipe Seed:
O grande salto é o long-narrative: passar de 4-8 s para 30 s em um único passo é o verdadeiro desafio dos modelos de vídeo — a deriva dos personagens, a coerência da cena e o custo computacional explodem com a duração. A referência multimodal (30 img + 10 vid + 10 áudio na entrada) também desloca a disputa do puro text-to-video para a produção assistida: o estúdio pré-injeta seus ativos, e o modelo permanece nos trilhos.
Nenhum benchmark de terceiros disponível ainda. O que observamos: a coerência do personagem em 30 s e o sincronismo labial com o áudio.
Artigo produzido por inteligência artificial, revisto sob controlo editorial humano.
Inicie sessão para se juntar à discussão.
30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.