Skip to main content
seedance 2.5seedancepromptsprompt guideai videobytedance

Seedance 2.5 Prompt Guide: The Complete Handbook (2026)

How to actually prompt Seedance 2.5 — the 30-second beat structure, all 50 reference slots, keeping a real face consistent, and the failure modes nobody documents. Written from a production pipeline, with the specs corrected.

Brian Bautista · Co-Founder & Creative Director|July 16, 202616 min read

Quick answer

Prompt Seedance 2.5 like a production brief, not a sentence. Keep the 2.0 grammar (subject, environment, action, camera, lighting, mood), then add what 2.5 needs: split the clip into 3 to 5 timed beats with one job each, give every reference an explicit job in the prompt text, put the reference you most need preserved first, and describe the audio instead of just enabling it. Use 8 to 12 purposeful references, not 50.

Seedance 2.5 landed on 31 July 2026, and the prompting advice that spread with it mostly restates what worked on 2.0. That advice is not wrong, it is just incomplete: the model now holds 30 seconds instead of 12, accepts 50 references instead of 12, and generates its own audio in the same pass. Each of those changes what a good prompt looks like.

This is the complete version. It covers the grammar, the beat structure, the full reference stack, how to keep a real person's face intact, the parameters worth setting, and the failure modes that cost people afternoons. Where a section needs more room than a handbook allows, it links to a deeper guide.

One thing up front about where the numbers come from, because this field is thick with vendors quoting each other. Model specs below are sourced and dated. Techniques marked as ours come from running a production video pipeline: thousands of paid generations, a failure classifier, and a lot of money spent proving things do not work. That pipeline currently runs Seedance 2.0, not 2.5. The prompt grammar carries across, and we say so rather than dressing 2.0 results up as 2.5 results.

The formula that did not change

Everything from the Seedance 2.0 prompt guide still applies. If you have not internalised it, start there. The core is unchanged:

Subject + Environment + Action + Camera + Lighting + Style/Mood

Vague subjects still produce generic output. "A woman" drifts; "a woman with a short black bob in a red wool coat" holds. Precise camera verbs still outperform adjectives. One strong lighting keyword still beats three competing ones. And the first 20 to 30 words still carry the most weight.

What 2.5 adds sits on top of that, and it is mostly about scale.

What actually changed, with the specs corrected

Seedance 2.5as of 17 August 2026
Launched
31 July 2026. Jimeng and Doubao first, then Dreamina and third-party APIs.
Clip length
4 to 30 seconds in one pass, or duration set to auto. Beta extension toward ~3 minutes.
Resolution
480p and 720p, with 1080p arriving around 17 August 2026. 4K is an upscale applied after generation.Widely mis-stated: that it "scales to native 4K". It does not. An upscale is not native capture, and the difference shows on fine texture and text. (Disputed)
Aspect ratios
16:9, 21:9 and 9:16.
References
50 total: 30 images, 10 video clips, 10 audio clips. Clips run about 1.8 to 30.2s each, capped around 30.2s per modality.
Reference syntax
@Image1, @Video1, @Audio1, numbered by upload order rather than by type.
Audio
Co-generated with the picture in one pass. generate_audio defaults to true.
Prompt adherence
Roughly 20% better than 2.0, by ByteDance's own measure. (Unconfirmed)
Consistency ceiling
Reliable to about 8 people or products in one clip; it degrades past that.
Weights
Closed. API and first-party apps only.

Compiled from ByteDance launch materials and provider API documentation (fal, KIE, BytePlus, EvoLink). Model specs move fast; check before budgeting a production.

Amber rows are ones where published sources disagree. We’ve said which reading we trust and why rather than picking one silently.

The resolution row is the one worth pausing on, because a lot of comparison pages have it backwards and it will wreck a delivery plan. Seedance 2.5 is not a 4K model today. If native resolution is what you need this month, MiniMax H3 ships native 2K and Seedance does not.

Structure the clip as beats, not a description

This is the single highest-leverage change, and it is where most 30-second generations fall apart.

A 12-second prompt can be one continuous description. Write 30 seconds the same way and the model has too much room: it wanders, rushes the middle, or holds one moment for far too long. The fix is to say what happens when.

A 30-second clip, carved into four beats4 beats · 30s
  1. 0s6sbeat 1

    Establish

    A woman in a red raincoat waits alone at a rain-soaked bus stop at night, sodium streetlights overhead. Static wide shot. She checks her watch.
  2. 6s14sbeat 2

    Arrival

    A vintage car pulls into frame, headlights flaring through the rain. Slow push in as she leans down to the passenger window.
  3. 14s24sbeat 3

    Interior turn

    Inside the car. Warm dashboard glow. She laughs at something the driver says, the tension going out of her shoulders. Rack focus from her face to the rearview mirror.
  4. 24s30sbeat 4

    Exit

    Exterior. The car pulls away down the empty street, taillights receding. Crane up to the skyline.

Three rules make beats work:

  • One job per beat. One location change, one emotional turn, or one camera move. A beat asked to do three things produces mush.
  • Make transitions explicit. "Inside the car" tells the model to cut. Without it you get a smeared morph between locations, which is the single most recognisable AI-video artifact.
  • Carry continuity across the gap. When a subject passes behind something, say what is unchanged when they reappear: same face, same coat, same bag. Otherwise the model treats the reappearance as a fresh invention.

Beats are not the same as a montage. Narrative scenes want clarity and few cuts; montages want coverage and variety. Say which you are making.

The full 30-second guide covers cause-and-effect sequencing, occlusion continuity, and chaining clips past 30 seconds with return_last_frame.

Spend the reference budget deliberately

Fifty slots is a ceiling, not a target. Every reference is an instruction, and instructions that overlap get averaged. Averaging is where identity drift comes from.

Think in layers, and fill only the ones your shot actually needs:

A typical strong allocation, 9 of 50 slots9 of 50 slots spent
Images7/30
Video clips1/10
Audio clips1/10
Identity×3
The face you cannot afford to lose. First in upload order, always.
World×3
Location plate, key prop, wardrobe.
Style×1
One colour-grade frame or film still that defines the look.
Motion×1
A clip whose choreography or camera move you want followed.
Sound×1
The track to cut to, or a voice reference.

Most strong generations land between 8 and 12 references. The 50-slot ceiling exists for productions with many characters, products and locations. It is not a score to max out.

Two rules govern the whole stack:

  1. Order is priority. References are numbered by upload order, not by type, and the model weights early ones more heavily. @Image1 should be whatever you most need preserved. For real-person work that is always the face.
  2. Name each reference's job in the prompt text. The model will not reliably infer that image 3 is the location and image 4 is the jacket. Say so, and say it narrowly: "@Image1 defines the woman's face and green jacket only." Naming what to use works better than listing what to avoid.
Warning

A reference that is uploaded but never mentioned in the prompt text may simply not bind. We have lost whole generations to this: the render comes back looking like a text-to-video result, and it reads as a prompt-adherence failure when it is actually a wiring failure. If your output ignored the reference completely, suspect the plumbing before you rewrite the prompt.

The full reference guide covers all three modalities in depth, including what video and audio references are genuinely good at versus what people assume they do.

Build the reference image, do not just attach one

Most guides stop at "attach a reference." The bigger lever is what you attach.

There are two constructions worth knowing, and they fail in different directions:

  • A character sheet is one image showing your subject from several angles and poses. It gives the model a rounded sense of a person and leaves it free to stage the scene. Best when you care about who far more than where.
  • A storyboard grid is one image of several panels showing specific moments. It pins staging, environment and on-screen elements. Best when the scene has to look a particular way: a stadium, a game UI, a specific set.

The tradeoff nobody mentions: a grid sequences by design, so it can introduce hard cuts between panels when you wanted continuous motion. And any generated reference introduces drift, because the generation step re-renders your subject rather than compositing them, so faces and hair change before the video model ever sees them.

Character sheet or storyboard grid walks through building both, and when each one wins.

Keeping a real person's face

If you are putting a real person in a scene, this is the whole game, and it has one counter-intuitive rule.

The standard advice, including our own older writing, says a photo captured mid-action beats a static one. For likeness that is backwards. Testing this repeatedly on our own pipeline, the pattern was consistent: a source photo of someone walking drifts badly, while a photo of someone in a held pose or gesture holds remarkably well.

The reason is that the model is not really refusing to hold still, it is completing the action the frame implies. A walking photo is an unfinished step, so the model finishes it. A stance or a gesture has nothing to finish.

So for likeness work, photo selection outranks prompt wording. Everything else stacks on top: identity reference first, stated narrowly, never contradicted in prose, with three to five visible anchors repeated in the same words every time you generate.

How to keep a real person's face consistent is the full version.

Prompt the audio, do not just enable it

Audio is co-generated with the picture, so the words you use for sound genuinely shape the edit. generate_audio is on by default, and leaving it at that produces generic ambience.

  • Diegetic detail beats a switch. "Rain on the car roof, wipers squeaking, muffled jazz from the radio" beats "generate audio."
  • Write dialogue with delivery. The line and how it is said: she says, almost laughing: "you're late." For two-handers, say who speaks and whose mouth stays closed.
  • Use music as structure. With an @Audio1 track, say what syncs: cuts land on the beat, final pose on the last hit.
  • Keep sound next to the event that makes it. Sound described beside its trigger lands in time; sound listed in a block at the end drifts.

Lip-sync is meaningfully tighter than 2.0, so dialogue scenes that were a coin flip before are worth attempting now.

Camera vocabulary is still the biggest quality lever

Precise camera language remains the highest-leverage skill in Seedance prompting, and 2.5 responds to more of it. Rack focus, crane moves and whip pans join the reliable push-ins and steadicam follows. Our camera movements library has copy-paste phrasing for 54 moves.

Two habits worth keeping:

Do not use the bare word "fast." It remains the single keyword most likely to degrade motion quality into blur and warping. Describe the speed concretely instead: "whip pan," "explosive burst," "snaps into position."

Say where the subject sits in frame. "Dynamic camerawork" is ignored. "Keep the red player in the left third of frame" is followed.

The parameters worth setting

ParameterWhat it does
duration4 to 30s, or auto. Per-second billing means length is a cost decision, not a default.
generate_audioOn by default. Turn it off only if you are scoring in post.
camera_fixedTrue for locked-off shots. Beats writing "static camera" in prose.
seedFix it to change one variable at a time between generations.
return_last_frameHands back the final frame as a clean image — the raw material for chaining a follow-up clip.

Resolution is a workflow decision as much as a quality one: iterate at the cheapest tier, finalise higher. Because 2.5 also prices by reference count, a heavy reference stack makes drafts more expensive, which is another argument for 9 references instead of 40.

If you want somewhere to run these, Kie.ai carries Seedance and MiniMax H3 on the same API, which is useful here for a reason beyond price: when a generation goes wrong, being able to throw the identical prompt at a second model is the fastest way to tell a model problem from a provider problem.

A complete prompt, dissected

Every clause below is doing declared work. Nothing is decoration.

30-second hotel exit, annotated
Identity

The man from @Image1 (identity: face and hairline only), wearing the suit from @Image2.

Identity first in upload order, scoped narrowly so the model does not also copy @Image1's background.

World

He crosses the marble lobby from @Image4.

Location pinned to a plate rather than described, which saves words and holds better.

Style

Grade follows @Image6: moody amber tungsten, cinematic film grain, 24fps.

One style reference and one lighting keyword. Competing lighting words cancel out.

Timing

0-8s: steadicam follow from behind as he crosses the floor, staff turning to watch.

Beat one: one job, one camera move.

Timing

8-18s: he pushes through brass doors into the night. Whip pan to reveal a waiting crowd of photographers, flashbulbs strobing.

The explicit door-to-street transition is what prevents a morph.

Timing

18-30s: slow push in on his face, half-smile, rack focus to the marquee behind him.

Beat three resolves. Three beats over 30s, not six.

Sound

Cut to @Audio1, beats landing on the flashes. Diegetic sound: crowd murmur, camera shutters, distant traffic.

Music given a sync job, ambience described rather than left to the model.

Camera

Avoid: morphing, distortion, sudden scene jumps, unnatural facial movement, extra limbs, blurred text, watermarks.

A short, targeted avoid list. Long generic ones dilute and can backfire.

Copy gives you the prompt with the annotations stripped.

Pro Tip

For photorealistic live-action, add 真人实拍 ("real person live-action footage") near the end of the prompt. It pushes the model away from illustration, which matters most when your reference image is itself a generated or illustrated frame. Placement is not cosmetic: leading with it tends to produce Chinese-language dialogue. Leave it out entirely for anime, 3D, pixel-art or other deliberately non-photoreal styles.

When it goes wrong

Most troubleshooting advice amounts to "change one thing," which is correct and useless if you do not know which thing. Symptoms map to causes fairly reliably:

The three that cost the most time3 symptoms
The output ignored my reference photo completely: different subject, different setting.Plumbing

What’s actually happening

Almost never the prompt. Either the reference parameter name was wrong, or the reference was uploaded but never named in the prompt text. Some APIs silently ignore unknown parameters instead of erroring, which quietly turns an image-to-video call into text-to-video.

Change this one thing

Check the wiring before touching the words. Confirm the reference actually attached, and that the prompt names it with the correct tag and capitalisation.
It keeps failing with a generic "Internal Error, please try again later."Moderation

What’s actually happening

That message is sometimes a content refusal wearing an outage costume. It reads as transient, so every retry goes back to the same model that just refused, and they all fail identically. Photos of children and certain likeness edits are common triggers.

Change this one thing

If three retries fail identically while other jobs succeed, treat it as a refusal rather than an outage. Change the input photo or route to a different model instead of retrying.
My subject will not hold still, no matter how forcefully I write "frozen" or "does not move".Model limit

What’s actually happening

This is a property of generative video, not a prompt bug. We spent four renders and roughly $4.86 proving it, then reproduced the same failure on an unrelated model family. These models animate people; that is what they are for. Restating it in the prompt does not change the outcome.

Change this one thing

Stop prompting for it. Start from a held-pose photo rather than a walking one, or use a depth-parallax renderer, which freezes by construction because nothing is being generated.

The full fault finder covers the rest, including competing camera instructions, why long negative lists backfire, and an honest list of things Seedance will not do regardless of how you ask.

The mistakes worth naming

  • Writing 30 seconds as one unbroken paragraph. Beats or mush.
  • Uploading references without assigning jobs. Unlabelled inputs get averaged.
  • Contradicting your references in prose. If @Image2 is a navy suit, do not write "black tuxedo." The model splits the difference and loses both.
  • Maxing duration by default. Per-second billing makes an unnecessary 30 seconds cost double.
  • Treating audio as a checkbox. Described sound produces sound design; an enabled switch produces generic ambience.
  • Rewriting everything when 90% worked. You had one failure and five variables. Change the one instruction connected to the one failure.

Where to go next


We write these because we run this pipeline daily and most of what is published about it is guesswork. If you would rather see the output than write the prompts, Starrd is the same techniques applied a few thousand times: pre-engineered scene templates where the beat structure, reference ordering and camera grammar are already solved. Useful as a reference for what the model can do, even if you never use it.

Frequently Asked Questions

Do Seedance 2.0 prompts still work on 2.5?

Yes, unchanged. Seedance 2.5 adds structure on top of the 2.0 grammar rather than replacing it. The same layered shot description, precise camera verbs and single strong lighting keyword all still apply. What 2.5 adds is scale: longer clips that need beat structure, a much larger reference budget that needs deliberate allocation, and native audio that responds to being described rather than merely switched on.

How many references should I actually use?

Between 8 and 12 for almost everything. The 50-slot ceiling exists for productions with many characters, products and locations, not as a target. Every reference is an instruction, and references that overlap or contradict each other get averaged, which is where identity drift comes from. Use as few references as pin down what you cannot afford to lose, and give each one an explicit job in the prompt text.

Does Seedance 2.5 do 4K?

Not natively, despite a lot of pages saying otherwise. As of August 2026 it generates at 480p and 720p, with 1080p arriving around 17 August. 4K is available as an upscale applied after generation, which is a different thing from native 4K and looks different. Budget and plan for 720p or 1080p as your real output resolution.

How long can a Seedance 2.5 clip be?

4 to 30 seconds in a single pass, and the API also accepts duration set to auto. There is a beta long-video path extending toward roughly three minutes. Because billing is per second, a 30-second generation costs about double the 15-second version of the same idea, so ask for the length the story actually needs rather than defaulting to the maximum.

Why does my character's face keep changing?

Usually one of three things: the identity reference is not first in upload order, the reference has no explicit job stated in the prompt text, or your prose contradicts the reference. A fourth cause catches people out on real-person work: if your source photo shows someone mid-stride, the model tends to complete that motion, and faces drift most while a subject is moving. A held pose or gesture holds far better than a walking shot.

What does 真人实拍 do in a Seedance prompt?

It means 'real person live-action footage' and it pushes the model toward photorealistic rendering instead of drifting into illustration, which matters most when your reference input is itself illustrated. Placement matters: put it near the end of the prompt. Leading with it tends to make the model generate Chinese-language dialogue. Leave it out entirely for anime, 3D, pixel-art or other deliberately non-photoreal styles.

Is Seedance 2.5 better than MiniMax H3 (Hailuo 03)?

They launched the same day and made opposite bets. Seedance 2.5 goes long, with 30-second single-pass clips and 50 reference slots. MiniMax H3 goes sharp and open, with native 2K at 24fps, 15-second clips and open weights. If you need duration and heavy reference control, Seedance. If you need resolution today, or you need to run the model yourself, H3.

About the author

Brian Bautista · Co-Founder & Creative Director

Brian is co-founder and creative director at Starrd, working as a creative technologist and data scientist. He tracks viral AI-video trends, designs Starrd's scene templates, and writes the deep-dive model comparisons and prompting breakdowns.

Part of The Seedance Guide

More guides in this series

Related Articles

Ready to create your own video?

Pick a template, upload your photos, and generate a cinematic AI video in minutes.

Browse Templates