ByteDance Launches Seedance 2.5 Video AI Model: 30-Second Generation with Long Narrative Power
ByteDance has domestically released Seedance 2.5, a video generation model capable of producing 30-second clips with coherent storytelling, multi-round extension, and multimodal reference handling, aiming at film, advertising, education, and industrial applications.
ByteDance today officially rolled out its latest video generation model, Seedance 2.5, across Chinese platforms. The model breaks the 15-second generation ceiling by delivering high-quality 30-second video clips in a single run, with support for seamless multi-round extensions.
The model is now available on Jimeng AI and Doubao Professional Edition, and API access will soon be provided through Volcano Ark. Built on the unified multimodal audio-visual joint generation architecture introduced in Seedance 2.0, this upgrade focuses on long narrative coherence, richer multimodal referencing, and advanced editing capabilities.
From Snapshots to Stories
Seedance 2.5 is designed to shift video AI from “generating a clip” to “completing a creation.” It can orchestrate multiple logically connected shots within 30 seconds, allowing stories to unfold with setup, development, turn, and conclusion.
A demo released by ByteDance shows a singer’s one-shot performance journey: the camera pushes through heavy red curtains into a warm backstage makeup room; a young female singer adjusts her earpiece and is prompted by staff; she walks through the corridor, interacts with dancers, takes the microphone; then steps onto the stage as the shot pulls back to reveal a stadium panorama with lights, banners, and a roaring crowd.
Multi-round extension inherits the same character likeness, scene, visual style, voice, and sound effects. Another sample depicts a boy running through a train car with a soccer ball, the metro stopping, the boy darting out, and a man finally catching up—all in continuous action across two 30-second segments.
Multimodal Reference at Scale
Users can feed the model up to 30 images, 10 video clips, and 10 audio files as reference material in a single input. The system comprehends composition, setting, style, character traits, and props across these sources, then applies them according to text instructions.
In group scenes, Seedance 2.5 simultaneously reconstructs multiple characters and their voices while preserving stable subject features. A 30-second concert showcase — shot in 16:9 cinematic realistic style — blends a scene reference, a pianist, cellist, violinist, vocalist, orchestra, choir, and audience into one smooth narrative.
ByteDance is targeting a broad range of industries: film and television, advertising, education, industrial manufacturing, embodied intelligence, and autonomous driving. With its extended duration and narrative logic, Seedance 2.5 signals a step toward AI tools that act as creative collaborators rather than mere clip generators.
Further details, demos, and technical documentation are available on the official Seedance 2.5 project page.