Quick answer
Put the identity reference first in upload order and state its job narrowly, never contradict it in prose, and repeat three to five visible anchors in identical wording every time you generate. Then the lever most people miss: choose a source photo of a held pose or gesture rather than someone mid-stride. Video models complete the action a frame implies, so a walking photo keeps moving and drifts, while a stance has nothing to finish and holds.
Identity drift is the hardest problem in AI video and the one with the most bad advice attached. The usual list (attach references, describe your character, be consistent) is correct and nowhere near sufficient, because it treats this as a prompting problem.
The single biggest lever is not in the prompt. It is which photo you start from.
A note on provenance, because this field is full of confident guessing: the technique list below is standard and widely corroborated. The photo-selection finding at the centre of this piece is ours, from our own pipeline, and it came out of a failed project rather than a successful one. That pipeline runs Seedance 2.0, not 2.5. The behaviour is a property of how generative video models treat a still frame rather than of a version number, but the evidence is 2.0 evidence and we would rather say so.
The finding that inverts the usual advice
Standard guidance, including our own older writing, says a photo captured mid-action makes a better reference than a static one. It feels right: more energy, more information, more to work with.
For likeness, it is backwards.
We spent a stretch of last month trying to build a time-freeze shot, where the subject holds still and the camera orbits them. Six renders failed identically: the subject walked, every time. We rewrote the prompt six ways, restated "does not move" five times in one attempt, switched reference construction twice, and then reproduced the exact same failure on a completely unrelated model family. Nothing moved the outcome.
Then one render worked. The prompt was not meaningfully different. The photo was: a subject standing and pointing, rather than mid-stride. The pose held for essentially the whole clip, face steady, while the camera orbited with real parallax.
Going back through the failures, all six had started from someone walking.
The explanation is simple once you see it. The model was never refusing to hold still. It was completing the action the frame implies. A walking photo is an unfinished step, so the model finishes the step. A stance, a gesture, a held expression: there is nothing to finish, so nothing gets invented.
That reframes photo selection entirely. You are not choosing the most flattering or the most dynamic picture. You are choosing a state rather than a transition, because every transition is an instruction to move, and faces drift most while a subject moves.
The practical version: before you upload, ask "is this person doing something, or partway through something?" Standing, sitting, pointing, arms folded, laughing, holding an object: all states, all hold well. Walking, running, jumping, turning, reaching: all transitions, all invite the model to finish the motion.
Choosing the source photo
In rough order of impact:
- A held pose, not a transition. Covered above. This outranks everything else on this list.
- The face unobstructed and reasonably large in frame. Sunglasses, heavy shadow across one side, a hand near the jaw, or a face occupying 5% of the frame all cost you likeness.
- Even, unremarkable lighting. Strong coloured light bakes into the model's idea of the person. Flat daylight is boring and reliable.
- Neutral or natural expression. An extreme expression gets treated as a facial feature and shows up in frames where it makes no sense.
- Recent and consistent. If you pass several photos, they should plausibly be the same person on the same day. Different weights, hairstyles or ages across references get averaged into someone who is neither.
Then the reference stack
With the right photo chosen, the standard techniques do their work. They are widely documented and they are genuinely necessary; they are just not sufficient on their own.
Identity goes first. References are numbered by upload order, and the model weights early ones more heavily. @Image1 should be the face. Not the location, not the wardrobe, not the mood board.
State the job narrowly. Not "@Image1 is a reference" but "@Image1 defines the woman's face and hairline only." Naming what to use works better than listing what to avoid, and scoping it stops the model also importing the photo's background, wardrobe or lighting.
Never contradict a reference in prose. If @Image2 is a navy suit, do not write "black tuxedo." The model splits the difference and you lose both. This is the most common self-inflicted drift there is.
Repeat visible anchors verbatim. Pick three to five concrete, visible traits and use the same words every single time: "short black bob, silver hoop earrings, rust-red leather jacket." Not "rust-red" in one generation and "burnt orange" in the next unless you want variation. Consistency of wording produces consistency of output to a degree that feels superstitious until you test it.
Three references, three angles. Front, three-quarter, profile beats six near-identical selfies. The model needs something to draw on when the camera swings around; redundant references just get averaged.
Respect the ceiling. Around eight people or products in one clip, consistency starts slipping. Past that, restage the shot rather than adding references.
The generated-reference trap
If your workflow builds an intermediate image (a character sheet, a storyboard grid, a stylised frame) from the user's photo before generating video, understand what that step costs.
The image model re-renders the person. It does not composite the original pixels; it generates a new picture of someone who resembles them. Faces and hair shift there, before the video model has seen anything at all. So the video is not drifting from your subject; it is faithfully rendering an intermediate that already drifted.
Two mitigations:
- Send the original photo as a second reference alongside the generated one. One extra slot, and the highest-value slot you will spend. The video model receives both and pulls likeness partway back.
- Skip the intermediate where the shot allows it. The raw photograph straight to the video model gives the best likeness available. You give up staging control and aspect normalisation in exchange.
Character sheet or storyboard grid covers the construction side of this in full.
The Chinese phrasing that actually helps
Seedance is a ByteDance model, and a few Chinese phrases are followed more reliably than their English equivalents. The one worth knowing for likeness work:
真人实拍 — real-person live-action footage. It pushes the model toward photorealistic rendering rather than drifting toward illustration, which matters most when your reference input is itself generated or stylised.
Two rules attach to it, and both are easy to get wrong:
- Put it near the end of the prompt. Leading with it tends to make the model generate Chinese-language dialogue, which is a memorable way to ruin a take.
- Leave it out for non-photoreal styles. Anime, 3D, pixel-art and chibi templates are deliberately not live-action, and this phrase fights them.
A short pre-flight check
Before you spend the generation:
- Source photo is a held pose, not mid-stride
- Face unobstructed, evenly lit, reasonably large in frame
- Identity reference is
@Image1, first in upload order - Its job is stated narrowly in the prompt text
- Nothing in the prose contradicts it
- Three to five visible anchors, worded identically to last time
- If you built an intermediate image, the original photo is attached too
- Fewer than about eight identities in the shot
-
真人实拍near the end, if and only if the style is live-action
When it still drifts
If you have done all of the above and the likeness is still wrong, the cause is usually not identity at all. An output that ignores the reference completely (different person, different setting) is almost always a wiring problem rather than a prompting one, and an output that fails identically three times in a row may be a refusal wearing an outage costume. The fault finder covers both.
And if the subject simply will not hold still no matter what you write: that one is not fixable from the prompt. It is a property of what these models are for. Start from a held-pose photo, or use a renderer that freezes by construction.
Every generation on Starrd is a real person's photo going into a cinematic scene, so this is the problem we spend the most time on. The templates there have the reference ordering, the anchor wording and the construction already solved per scene, which is a reasonable place to see the techniques above working before you build your own.
Related: the full prompt handbook, all 50 reference slots.