Quick answer
Prompt Seedance 2.5 like a production brief, not a sentence. Keep the 2.0 grammar (subject, environment, action, camera, lighting, mood), then add what 2.5 needs: split the clip into 3 to 5 timed beats with one job each, give every reference an explicit job in the prompt text, put the reference you most need preserved first, and describe the audio instead of just enabling it. Use 8 to 12 purposeful references, not 50.
Seedance 2.5 landed on 31 July 2026, and the prompting advice that spread with it mostly restates what worked on 2.0. That advice is not wrong, it is just incomplete: the model now holds 30 seconds instead of 12, accepts 50 references instead of 12, and generates its own audio in the same pass. Each of those changes what a good prompt looks like.
This is the complete version. It covers the grammar, the beat structure, the full reference stack, how to keep a real person's face intact, the parameters worth setting, and the failure modes that cost people afternoons. Where a section needs more room than a handbook allows, it links to a deeper guide.
One thing up front about where the numbers come from, because this field is thick with vendors quoting each other. Model specs below are sourced and dated. Techniques marked as ours come from running a production video pipeline: thousands of paid generations, a failure classifier, and a lot of money spent proving things do not work. That pipeline currently runs Seedance 2.0, not 2.5. The prompt grammar carries across, and we say so rather than dressing 2.0 results up as 2.5 results.
The formula that did not change
Everything from the Seedance 2.0 prompt guide still applies. If you have not internalised it, start there. The core is unchanged:
Subject + Environment + Action + Camera + Lighting + Style/Mood
Vague subjects still produce generic output. "A woman" drifts; "a woman with a short black bob in a red wool coat" holds. Precise camera verbs still outperform adjectives. One strong lighting keyword still beats three competing ones. And the first 20 to 30 words still carry the most weight.
What 2.5 adds sits on top of that, and it is mostly about scale.
What actually changed, with the specs corrected
- Launched
- 31 July 2026. Jimeng and Doubao first, then Dreamina and third-party APIs.
- Clip length
- 4 to 30 seconds in one pass, or duration set to auto. Beta extension toward ~3 minutes.
- Resolution
- 480p and 720p, with 1080p arriving around 17 August 2026. 4K is an upscale applied after generation.Widely mis-stated: that it "scales to native 4K". It does not. An upscale is not native capture, and the difference shows on fine texture and text. (Disputed)
- Aspect ratios
- 16:9, 21:9 and 9:16.
- References
- 50 total: 30 images, 10 video clips, 10 audio clips. Clips run about 1.8 to 30.2s each, capped around 30.2s per modality.
- Reference syntax
- @Image1, @Video1, @Audio1, numbered by upload order rather than by type.
- Audio
- Co-generated with the picture in one pass. generate_audio defaults to true.
- Prompt adherence
- Roughly 20% better than 2.0, by ByteDance's own measure. (Unconfirmed)
- Consistency ceiling
- Reliable to about 8 people or products in one clip; it degrades past that.
- Weights
- Closed. API and first-party apps only.
Compiled from ByteDance launch materials and provider API documentation (fal, KIE, BytePlus, EvoLink). Model specs move fast; check before budgeting a production.
Amber rows are ones where published sources disagree. We’ve said which reading we trust and why rather than picking one silently.
The resolution row is the one worth pausing on, because a lot of comparison pages have it backwards and it will wreck a delivery plan. Seedance 2.5 is not a 4K model today. If native resolution is what you need this month, MiniMax H3 ships native 2K and Seedance does not.
Structure the clip as beats, not a description
This is the single highest-leverage change, and it is where most 30-second generations fall apart.
A 12-second prompt can be one continuous description. Write 30 seconds the same way and the model has too much room: it wanders, rushes the middle, or holds one moment for far too long. The fix is to say what happens when.
- 0s–6sbeat 1
Establish
A woman in a red raincoat waits alone at a rain-soaked bus stop at night, sodium streetlights overhead. Static wide shot. She checks her watch. - 6s–14sbeat 2
Arrival
A vintage car pulls into frame, headlights flaring through the rain. Slow push in as she leans down to the passenger window. - 14s–24sbeat 3
Interior turn
Inside the car. Warm dashboard glow. She laughs at something the driver says, the tension going out of her shoulders. Rack focus from her face to the rearview mirror. - 24s–30sbeat 4
Exit
Exterior. The car pulls away down the empty street, taillights receding. Crane up to the skyline.
Three rules make beats work:
- One job per beat. One location change, one emotional turn, or one camera move. A beat asked to do three things produces mush.
- Make transitions explicit. "Inside the car" tells the model to cut. Without it you get a smeared morph between locations, which is the single most recognisable AI-video artifact.
- Carry continuity across the gap. When a subject passes behind something, say what is unchanged when they reappear: same face, same coat, same bag. Otherwise the model treats the reappearance as a fresh invention.
Beats are not the same as a montage. Narrative scenes want clarity and few cuts; montages want coverage and variety. Say which you are making.
The full 30-second guide covers cause-and-effect sequencing, occlusion continuity, and chaining clips past 30 seconds with return_last_frame.
Spend the reference budget deliberately
Fifty slots is a ceiling, not a target. Every reference is an instruction, and instructions that overlap get averaged. Averaging is where identity drift comes from.
Think in layers, and fill only the ones your shot actually needs:
- Identity×3
- The face you cannot afford to lose. First in upload order, always.
- World×3
- Location plate, key prop, wardrobe.
- Style×1
- One colour-grade frame or film still that defines the look.
- Motion×1
- A clip whose choreography or camera move you want followed.
- Sound×1
- The track to cut to, or a voice reference.
Most strong generations land between 8 and 12 references. The 50-slot ceiling exists for productions with many characters, products and locations. It is not a score to max out.
Two rules govern the whole stack:
- Order is priority. References are numbered by upload order, not by type, and the model weights early ones more heavily.
@Image1should be whatever you most need preserved. For real-person work that is always the face. - Name each reference's job in the prompt text. The model will not reliably infer that image 3 is the location and image 4 is the jacket. Say so, and say it narrowly: "@Image1 defines the woman's face and green jacket only." Naming what to use works better than listing what to avoid.
A reference that is uploaded but never mentioned in the prompt text may simply not bind. We have lost whole generations to this: the render comes back looking like a text-to-video result, and it reads as a prompt-adherence failure when it is actually a wiring failure. If your output ignored the reference completely, suspect the plumbing before you rewrite the prompt.
The full reference guide covers all three modalities in depth, including what video and audio references are genuinely good at versus what people assume they do.
Build the reference image, do not just attach one
Most guides stop at "attach a reference." The bigger lever is what you attach.
There are two constructions worth knowing, and they fail in different directions:
- A character sheet is one image showing your subject from several angles and poses. It gives the model a rounded sense of a person and leaves it free to stage the scene. Best when you care about who far more than where.
- A storyboard grid is one image of several panels showing specific moments. It pins staging, environment and on-screen elements. Best when the scene has to look a particular way: a stadium, a game UI, a specific set.
The tradeoff nobody mentions: a grid sequences by design, so it can introduce hard cuts between panels when you wanted continuous motion. And any generated reference introduces drift, because the generation step re-renders your subject rather than compositing them, so faces and hair change before the video model ever sees them.
Character sheet or storyboard grid walks through building both, and when each one wins.
Keeping a real person's face
If you are putting a real person in a scene, this is the whole game, and it has one counter-intuitive rule.
The standard advice, including our own older writing, says a photo captured mid-action beats a static one. For likeness that is backwards. Testing this repeatedly on our own pipeline, the pattern was consistent: a source photo of someone walking drifts badly, while a photo of someone in a held pose or gesture holds remarkably well.
The reason is that the model is not really refusing to hold still, it is completing the action the frame implies. A walking photo is an unfinished step, so the model finishes it. A stance or a gesture has nothing to finish.
So for likeness work, photo selection outranks prompt wording. Everything else stacks on top: identity reference first, stated narrowly, never contradicted in prose, with three to five visible anchors repeated in the same words every time you generate.
How to keep a real person's face consistent is the full version.
Prompt the audio, do not just enable it
Audio is co-generated with the picture, so the words you use for sound genuinely shape the edit. generate_audio is on by default, and leaving it at that produces generic ambience.
- Diegetic detail beats a switch. "Rain on the car roof, wipers squeaking, muffled jazz from the radio" beats "generate audio."
- Write dialogue with delivery. The line and how it is said: she says, almost laughing: "you're late." For two-handers, say who speaks and whose mouth stays closed.
- Use music as structure. With an
@Audio1track, say what syncs: cuts land on the beat, final pose on the last hit. - Keep sound next to the event that makes it. Sound described beside its trigger lands in time; sound listed in a block at the end drifts.
Lip-sync is meaningfully tighter than 2.0, so dialogue scenes that were a coin flip before are worth attempting now.
Camera vocabulary is still the biggest quality lever
Precise camera language remains the highest-leverage skill in Seedance prompting, and 2.5 responds to more of it. Rack focus, crane moves and whip pans join the reliable push-ins and steadicam follows. Our camera movements library has copy-paste phrasing for 54 moves.
Two habits worth keeping:
Do not use the bare word "fast." It remains the single keyword most likely to degrade motion quality into blur and warping. Describe the speed concretely instead: "whip pan," "explosive burst," "snaps into position."
Say where the subject sits in frame. "Dynamic camerawork" is ignored. "Keep the red player in the left third of frame" is followed.
The parameters worth setting
| Parameter | What it does |
|---|---|
duration | 4 to 30s, or auto. Per-second billing means length is a cost decision, not a default. |
generate_audio | On by default. Turn it off only if you are scoring in post. |
camera_fixed | True for locked-off shots. Beats writing "static camera" in prose. |
seed | Fix it to change one variable at a time between generations. |
return_last_frame | Hands back the final frame as a clean image — the raw material for chaining a follow-up clip. |
Resolution is a workflow decision as much as a quality one: iterate at the cheapest tier, finalise higher. Because 2.5 also prices by reference count, a heavy reference stack makes drafts more expensive, which is another argument for 9 references instead of 40.
If you want somewhere to run these, Kie.ai carries Seedance and MiniMax H3 on the same API, which is useful here for a reason beyond price: when a generation goes wrong, being able to throw the identical prompt at a second model is the fastest way to tell a model problem from a provider problem.
A complete prompt, dissected
Every clause below is doing declared work. Nothing is decoration.
The man from @Image1 (identity: face and hairline only), wearing the suit from @Image2.
Identity first in upload order, scoped narrowly so the model does not also copy @Image1's background.
He crosses the marble lobby from @Image4.
Location pinned to a plate rather than described, which saves words and holds better.
Grade follows @Image6: moody amber tungsten, cinematic film grain, 24fps.
One style reference and one lighting keyword. Competing lighting words cancel out.
0-8s: steadicam follow from behind as he crosses the floor, staff turning to watch.
Beat one: one job, one camera move.
8-18s: he pushes through brass doors into the night. Whip pan to reveal a waiting crowd of photographers, flashbulbs strobing.
The explicit door-to-street transition is what prevents a morph.
18-30s: slow push in on his face, half-smile, rack focus to the marquee behind him.
Beat three resolves. Three beats over 30s, not six.
Cut to @Audio1, beats landing on the flashes. Diegetic sound: crowd murmur, camera shutters, distant traffic.
Music given a sync job, ambience described rather than left to the model.
Avoid: morphing, distortion, sudden scene jumps, unnatural facial movement, extra limbs, blurred text, watermarks.
A short, targeted avoid list. Long generic ones dilute and can backfire.
Copy gives you the prompt with the annotations stripped.
For photorealistic live-action, add 真人实拍 ("real person live-action footage") near the end of the prompt. It pushes the model away from illustration, which matters most when your reference image is itself a generated or illustrated frame. Placement is not cosmetic: leading with it tends to produce Chinese-language dialogue. Leave it out entirely for anime, 3D, pixel-art or other deliberately non-photoreal styles.
When it goes wrong
Most troubleshooting advice amounts to "change one thing," which is correct and useless if you do not know which thing. Symptoms map to causes fairly reliably:
The output ignored my reference photo completely: different subject, different setting.Plumbing
What’s actually happening
Change this one thing
It keeps failing with a generic "Internal Error, please try again later."Moderation
What’s actually happening
Change this one thing
My subject will not hold still, no matter how forcefully I write "frozen" or "does not move".Model limit
What’s actually happening
Change this one thing
The full fault finder covers the rest, including competing camera instructions, why long negative lists backfire, and an honest list of things Seedance will not do regardless of how you ask.
The mistakes worth naming
- Writing 30 seconds as one unbroken paragraph. Beats or mush.
- Uploading references without assigning jobs. Unlabelled inputs get averaged.
- Contradicting your references in prose. If
@Image2is a navy suit, do not write "black tuxedo." The model splits the difference and loses both. - Maxing duration by default. Per-second billing makes an unnecessary 30 seconds cost double.
- Treating audio as a checkbox. Described sound produces sound design; an enabled switch produces generic ambience.
- Rewriting everything when 90% worked. You had one failure and five variables. Change the one instruction connected to the one failure.
Where to go next
- All 50 references, in depth — images, video and audio, and what each is actually good for
- Character sheet or storyboard grid — building the reference, not just attaching it
- Writing 30 seconds that does not wander — beats, continuity, chaining past the limit
- Keeping a real person's face consistent — the likeness problem, properly
- When Seedance ignores your prompt — the fault finder
- Seedance 2.5 vs MiniMax H3 — same launch day, opposite bets
- Seedance 2.0 fundamentals — the grammar underneath all of it
We write these because we run this pipeline daily and most of what is published about it is guesswork. If you would rather see the output than write the prompts, Starrd is the same techniques applied a few thousand times: pre-engineered scene templates where the beat structure, reference ordering and camera grammar are already solved. Useful as a reference for what the model can do, even if you never use it.