
ByteDance's flagship. 15s video with native audio in a single generation.
ByteDance's flagship video model. Picture and sound are generated together in one pass, so lip movements, effects, and music land in perfect sync. Up to 15 seconds with native audio, and up to 9 reference images to keep characters and scenes consistent shot to shot. Comes in four rungs: Mini for the lowest cost, Fast for quicker turnaround, Standard for the full flagship, and Cinematic, the premium high-bitrate version with roughly 3x the detail retention and output up to true 4K.

“Armoured hero landing on a rain-soaked rooftop, shockwave cracking concrete, cape billowing, lightning illuminating a neon cityscape behind, dramatic low angle, volumetric rain, Unreal Engine 5 cinematic”

“High-speed car chase along a sun-drenched Miami coastal highway, matte black supercar drifting sideways through an intersection, tyre smoke, palm trees blurring, helicopter tracking shot, golden hour lens flare”

“Anime warrior standing on a floating crystal platform above clouds, glowing energy sword raised, a colossal dragon emerging from the storm below, cel-shaded rendering, volumetric god rays, epic fantasy”

“Cyberpunk samurai walking through a neon-drenched market street in futuristic Tokyo, holographic ads floating overhead, rain cascading off a translucent umbrella, katana on back, atmospheric fog, chrome reflections”

“Open-world gameplay shot: figure on a motorcycle cresting a hill overlooking a vast coastal city at sunset, ocean to the horizon, winding highway below, photorealistic, cinematic colour grading”

“Colossal mech robot emerging from stormy ocean waves, searchlights cutting through spray and fog, fighter jets banking away, lightning illuminating armour plating, dramatic low angle, IMAX scale, teal and orange”
Type a detailed prompt describing the video you want, or upload a reference image as a starting frame.
Pick your resolution and duration. See the credit cost before you generate.
Your video is ready in 1-3 minutes. Download, iterate, or extend the sequence.
Jump into the Studio and start generating. Plans from $11/month.
Seedance 2.0 is ByteDance's flagship video generation model. It uses a dual-branch Multi-Modal Diffusion Transformer: one branch generates video frames, the other generates audio waveforms, connected by a cross-attention bridge that synchronises them at every step. The result is video and audio created together in a single pass, not audio bolted on after the fact.
The multi-modal reference system is the standout capability. Feed Seedance 2.0 up to 9 reference images for character appearance and scene consistency. Use @1, @2 etc. in your prompt to direct specific references: '@1 walks through the market while @2 watches from a balcony.' The model decouples content from motion, letting you combine a character from one reference with camera movement from another. This is directing, not prompting.
Output reaches 15 seconds at 24fps with native dual-channel stereo audio. Flow Matching replaces traditional Gaussian diffusion with a more direct mathematical path from noise to output, delivering 30% faster generation than Seedance 1.5 while improving quality. 480p and 720p options are available for faster iteration at lower credit cost. The family runs in four rungs: Mini, the budget rung, at 480p/720p for the lowest cost; Fast for quicker turnaround; Standard for the full 1080p flagship; and Cinematic, the premium high-bitrate version. Cinematic is the same model delivered in a cinema-grade encode carrying roughly three times the data of a standard delivery, so skin texture, fabric, foliage, and fast motion hold their detail, with output up to true 4K. The workflow that earns its keep: iterate on the cheaper rungs until the shot is right, then re-run the winning prompt and references on Cinematic for the final take.
Professional video generation. Plans from $11/month.