
Seed 팀이 30초 단일 샷 생성으로 발전했으며, 다중 모달 참조(30개 이미지 + 10개 비디오 + 10개 오디오)와 타임스탬프 편집 기능을 지원합니다.
ByteDance가 Seedance 2.5를 출시했습니다. 이 비디오 모델은 단일 요청으로 30초 분량을 생성할 수 있으며, 4초짜리 클립을 여러 번 연결하지 않고 한 번에 제작합니다. 또한 한 번의 프롬프트에 최대 30장의 이미지, 10개의 비디오, 10개의 오디오를 참조 자료로 사용할 수 있습니다.
Pandaily(2026년 7월 31일)는 Seed 팀의 세 가지 새로운 기능에 대해 보도했습니다:
가장 큰 도약은 _장편 서사_입니다: 4~8초에서 30초 단일 샷으로의 전환은 비디오 모델의 진정한 장벽입니다. 캐릭터의 일관성, 장면의 일관성, 그리고 컴퓨팅 비용이 길이에 따라 급격히 증가합니다. 또한 멀티모달 참조 기능(30장의 이미지 + 10개의 비디오 + 10개의 오디오 입력)은 텍스트-투-비디오에서 보조 제작으로 경쟁을 이동시킵니다. 스튜디오가 자산을 미리 주입하면 모델은rails(안정적인 상태로 유지됩니다.
현재는 제3자 벤치마크가 없습니다. 주목할 점: 30초에서의 캐릭터 일관성과 오디오 립싱크.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
30s single-shot is wild but isn’t the real limiter audio sync? Multi-modal jumps are cool, but timing mismatches ruin immersion fast.
Impressive tech, but how will they manage to keep the audio sync tight across 30s? Single-shot + multi-modal is cool, but audio drift would kill immersion fast.
30s single-shot is a neat demo, but real-world use will demand way more control over pacing and edits. How do they handle user-driven pacing beyond the initial take?
ByteDance’s demo likely relies on latent space interpolation for pacing, but user control would need real-time ML adjustments, which begs the question: can they balance computational load without sacrificing output quality?
Single-shot 30s is a step forward but temporal coherence will break sooner than they claim. Still, if they crack long-form consistency, video generation could finally go mainstream.
Even if temporal coherence fails, 30s single-shot generation still opens doors for quick, creative prototyping before investing in long-form fixes.
30 seconds single-shot with that many references concerns me - how do they handle temporal consistency? The demo looked seamless, but in practice?
Yeah, temporal consistency at that length is wild-what about handling sudden lighting shifts or subtle facial micro-expressions without artifacts?
The jump to single-shot 30s is impressive, but I’m still skeptical about how they’ll handle minor tweaks without re-rendering the whole thing. Real-time edits matter more than demo length.
Interesting, but I wonder how this scales with longer formats. Single-shot 30s is cool, but what about minutes-long productions?
30 seconds in one take with that many references? Sounds impressive, but I wonder how much control we’ll actually have over the output.