Quick answer
Split the clip into three to five timed beats and give each one job: a single location change, emotional turn or camera move. Name every transition explicitly, because an unstated cut becomes a smeared morph. State what stays unchanged when a subject reappears from behind something. Write reactions a beat after their triggers rather than alongside them. And do not default to 30 seconds, because per-second billing makes an unnecessary 30-second clip cost double the 15-second version of the same idea.
Seedance 2.5 holds 30 seconds in a single pass. That is a genuine capability jump, and it breaks a habit that worked fine at 12 seconds.
At 12 seconds you can write one continuous description and get a coherent shot. Write 30 seconds the same way and the model has too much room. It wanders, compresses the middle, or holds a single moment far longer than you meant. Nothing is wrong with the prompt except that it never said how to spend the time.
This is the structural half of the Seedance prompt handbook.
Beats, not description
The fix is to say what happens when. Three to five beats, each with one job.
- 0s–6sbeat 1
Establish
A woman in a red raincoat waits alone at a rain-soaked bus stop at night, sodium streetlights overhead. Static wide shot. She checks her watch. - 6s–14sbeat 2
Arrival
A vintage car pulls into frame, headlights flaring through the rain. Slow push in as she leans down to the passenger window. - 14s–24sbeat 3
Interior turn
Inside the car. Warm dashboard glow. She laughs at something the driver says, the tension going out of her shoulders. Rack focus from her face to the rearview mirror. - 24s–30sbeat 4
Exit
Exterior. The car pulls away down the empty street, taillights receding. Crane up to the skyline.
One job per beat. One location change, one emotional turn, or one camera move. A beat carrying three of those produces mush, because the model has to average them into the same span of time.
Three to five, not eight. Under three and the model allocates duration itself, which on longer clips means a rushed middle. Over five and no beat gets enough screen time to land, so the clip reads as a list of events rather than a scene. Three to four meaningful beats per 15 seconds is a good working rhythm.
Front-load anyway. The first 20 to 30 words still carry the most weight, exactly as they did on 2.0. Beat one should establish the things you cannot afford to have wrong.
Name every transition
An unstated cut becomes a morph, and a morph between locations is the single most recognisable AI-video artifact there is.
"Inside the car" tells the model to cut. Without it, you get a smeared interpolation from the street into the car interior, with the woman's coat becoming upholstery on the way. Two words at the head of the beat prevent it.
The same applies to time. If beat three happens later that night, say so. The model will otherwise assume continuity and light it accordingly.
Carry continuity across the gap
Whenever a subject passes behind something, the model treats their reappearance as a fresh generation unless told otherwise. This is where a character quietly changes coat colour halfway through a clip.
State the occlusion and the invariants:
14-18s: she passes behind the stone column and is hidden for about two seconds. 18-22s: the same woman emerges from the far side of the column: same face, same green coat, same black boots, same red suitcase. Her walking pace is unchanged.
Naming what stays the same is more effective than naming what must not change. The first is a description the model can render; the second is an instruction it has to interpret.
Let effects follow their causes
Longer clips give you room for cause and effect, and the way to get it is to write the reaction as a separate event slightly after the trigger, rather than describing both at once.
"The customer looks surprised as the glass falls" tends to produce a surprised face and a falling glass happening simultaneously, which reads as uncanny. Written in sequence, with the contact first and the reactions staggered, it reads as physics:
8-10s: the glass slips from the edge of the counter and strikes the tile, shattering outward. 10-11s: the nearest customer flinches, shoulders rising. 11-13s: two people at the far table turn toward the sound a beat later.
That staggering, the far table turning a beat later, is what makes a crowd look like a crowd rather than a chorus line.
Describe physical motion in phases
For anything where the physics has to read correctly, break the movement into contact phases instead of naming the action. Not "he does a kickflip" but the approach, the specific contact, the force transfer, and the settle:
the rear foot snaps the tail against the concrete, the board rises, the front foot slides forward, then the knees absorb the landing
Same for liquids, cloth and smoke: name the direction of force and the final settling state rather than asking for realistic physics, which is not an instruction the model can act on.
Narrative and montage want opposite things
These pull in different directions, so say which you are making.
Narrative wants clarity and few cuts. Beats are moments in one continuous situation, transitions are deliberate, and the camera behaves consistently.
Montage wants coverage and variety: faster cuts, varied angles, inserts, and no obligation to maintain a single continuous space. A montage prompt written with narrative discipline feels sluggish, and a narrative prompt written with montage energy feels incoherent.
Dialogue needs its own timing
Dialogue does not fit into a beat implicitly; give it a slot.
0-3s: she looks straight into the phone camera and says, plainly: I bought this for the commute. 3-8s: she turns the bag to show the side pocket. She does not speak while turning it. 8-12s: she looks back up and adds, half-smiling: it fits a laptop, barely.
Two rules: say who speaks and whose mouth stays closed in a multi-character shot, and give non-speaking action its own span so the model does not stretch the line across it. Lip-sync on 2.5 is meaningfully tighter than on 2.0, so dialogue that was a coin flip before is worth attempting now.
Going past 30 seconds
Two routes.
Chaining is the controllable one. Ask the API to return the last frame of a clip (return_last_frame), then hand that frame to the next generation as @Image1 and frame the prompt as a continuation rather than a new scene:
Use @Image1 as the exact first frame and continue forward from that moment. The same man, same suit, same lobby, same amber tungsten grade. 0-8s: he reaches the revolving door and pushes through into the street.
Say "continue forward from that moment" explicitly. Without it, the model tends to treat the frame as a style reference and restart the action.
Beta long-video extends toward roughly three minutes in one pass. It is less controllable per segment, and worth reaching for only when the shot genuinely cannot be cut.
Do not default to 30 seconds
Billing is per second, and on 2.5 it also scales with reference count. An unnecessary 30-second generation costs about double the 15-second version of the same idea, and you will iterate more than once.
The workflow that saves the most money: draft short and low-resolution until the structure is right, then extend and finalise the version that works. If you need somewhere to do that cheaply, Kie.ai runs Seedance per-second, which makes throwaway drafts genuinely cheap.
Common mistakes
- One unbroken paragraph for 30 seconds. The default failure. Beats or mush.
- Eight beats. Each too short to land; the clip becomes a list.
- Unstated transitions. Morphs instead of cuts.
- Simultaneous cause and effect. Reads uncanny. Stagger it.
- Silent occlusions. The coat changes colour behind the pillar.
- Maxing duration by reflex. Double the cost for no story reason.
- Mixing montage energy into a narrative scene. Decide which one you are making.
Where to go next
- The prompt handbook for the full grammar
- All 50 reference slots for what the beats are being populated with
- Camera movements library for the moves each beat can name
- When Seedance ignores your prompt if the pacing fix did not take
Every scene template on Starrd is a beat-structured prompt underneath, refined over a lot of paid generations. Worth a look as a set of worked examples, even if you write your own.