Skip to main content
seedanceseedance 2.5character consistencyidentity driftface consistencyai videolikeness

How to Keep a Real Person's Face Consistent in AI Video

Identity drift is the hardest problem in AI video, and the biggest lever is not the prompt. What actually holds a real likeness across a Seedance generation, including the photo-selection rule that inverts the usual advice.

Brian Bautista · Co-Founder & Creative Director|August 17, 20268 min read

Quick answer

Put the identity reference first in upload order and state its job narrowly, never contradict it in prose, and repeat three to five visible anchors in identical wording every time you generate. Then the lever most people miss: choose a source photo of a held pose or gesture rather than someone mid-stride. Video models complete the action a frame implies, so a walking photo keeps moving and drifts, while a stance has nothing to finish and holds.

Identity drift is the hardest problem in AI video and the one with the most bad advice attached. The usual list (attach references, describe your character, be consistent) is correct and nowhere near sufficient, because it treats this as a prompting problem.

The single biggest lever is not in the prompt. It is which photo you start from.

A note on provenance, because this field is full of confident guessing: the technique list below is standard and widely corroborated. The photo-selection finding at the centre of this piece is ours, from our own pipeline, and it came out of a failed project rather than a successful one. That pipeline runs Seedance 2.0, not 2.5. The behaviour is a property of how generative video models treat a still frame rather than of a version number, but the evidence is 2.0 evidence and we would rather say so.

The finding that inverts the usual advice

Standard guidance, including our own older writing, says a photo captured mid-action makes a better reference than a static one. It feels right: more energy, more information, more to work with.

For likeness, it is backwards.

We spent a stretch of last month trying to build a time-freeze shot, where the subject holds still and the camera orbits them. Six renders failed identically: the subject walked, every time. We rewrote the prompt six ways, restated "does not move" five times in one attempt, switched reference construction twice, and then reproduced the exact same failure on a completely unrelated model family. Nothing moved the outcome.

Then one render worked. The prompt was not meaningfully different. The photo was: a subject standing and pointing, rather than mid-stride. The pose held for essentially the whole clip, face steady, while the camera orbited with real parallax.

Going back through the failures, all six had started from someone walking.

The explanation is simple once you see it. The model was never refusing to hold still. It was completing the action the frame implies. A walking photo is an unfinished step, so the model finishes the step. A stance, a gesture, a held expression: there is nothing to finish, so nothing gets invented.

That reframes photo selection entirely. You are not choosing the most flattering or the most dynamic picture. You are choosing a state rather than a transition, because every transition is an instruction to move, and faces drift most while a subject moves.

Pro Tip

The practical version: before you upload, ask "is this person doing something, or partway through something?" Standing, sitting, pointing, arms folded, laughing, holding an object: all states, all hold well. Walking, running, jumping, turning, reaching: all transitions, all invite the model to finish the motion.

Choosing the source photo

In rough order of impact:

  1. A held pose, not a transition. Covered above. This outranks everything else on this list.
  2. The face unobstructed and reasonably large in frame. Sunglasses, heavy shadow across one side, a hand near the jaw, or a face occupying 5% of the frame all cost you likeness.
  3. Even, unremarkable lighting. Strong coloured light bakes into the model's idea of the person. Flat daylight is boring and reliable.
  4. Neutral or natural expression. An extreme expression gets treated as a facial feature and shows up in frames where it makes no sense.
  5. Recent and consistent. If you pass several photos, they should plausibly be the same person on the same day. Different weights, hairstyles or ages across references get averaged into someone who is neither.

Then the reference stack

With the right photo chosen, the standard techniques do their work. They are widely documented and they are genuinely necessary; they are just not sufficient on their own.

Identity goes first. References are numbered by upload order, and the model weights early ones more heavily. @Image1 should be the face. Not the location, not the wardrobe, not the mood board.

State the job narrowly. Not "@Image1 is a reference" but "@Image1 defines the woman's face and hairline only." Naming what to use works better than listing what to avoid, and scoping it stops the model also importing the photo's background, wardrobe or lighting.

Never contradict a reference in prose. If @Image2 is a navy suit, do not write "black tuxedo." The model splits the difference and you lose both. This is the most common self-inflicted drift there is.

Repeat visible anchors verbatim. Pick three to five concrete, visible traits and use the same words every single time: "short black bob, silver hoop earrings, rust-red leather jacket." Not "rust-red" in one generation and "burnt orange" in the next unless you want variation. Consistency of wording produces consistency of output to a degree that feels superstitious until you test it.

Three references, three angles. Front, three-quarter, profile beats six near-identical selfies. The model needs something to draw on when the camera swings around; redundant references just get averaged.

Respect the ceiling. Around eight people or products in one clip, consistency starts slipping. Past that, restage the shot rather than adding references.

The generated-reference trap

If your workflow builds an intermediate image (a character sheet, a storyboard grid, a stylised frame) from the user's photo before generating video, understand what that step costs.

The image model re-renders the person. It does not composite the original pixels; it generates a new picture of someone who resembles them. Faces and hair shift there, before the video model has seen anything at all. So the video is not drifting from your subject; it is faithfully rendering an intermediate that already drifted.

Two mitigations:

  • Send the original photo as a second reference alongside the generated one. One extra slot, and the highest-value slot you will spend. The video model receives both and pulls likeness partway back.
  • Skip the intermediate where the shot allows it. The raw photograph straight to the video model gives the best likeness available. You give up staging control and aspect normalisation in exchange.

Character sheet or storyboard grid covers the construction side of this in full.

The Chinese phrasing that actually helps

Seedance is a ByteDance model, and a few Chinese phrases are followed more reliably than their English equivalents. The one worth knowing for likeness work:

真人实拍 — real-person live-action footage. It pushes the model toward photorealistic rendering rather than drifting toward illustration, which matters most when your reference input is itself generated or stylised.

Two rules attach to it, and both are easy to get wrong:

  • Put it near the end of the prompt. Leading with it tends to make the model generate Chinese-language dialogue, which is a memorable way to ruin a take.
  • Leave it out for non-photoreal styles. Anime, 3D, pixel-art and chibi templates are deliberately not live-action, and this phrase fights them.

A short pre-flight check

Before you spend the generation:

  • Source photo is a held pose, not mid-stride
  • Face unobstructed, evenly lit, reasonably large in frame
  • Identity reference is @Image1, first in upload order
  • Its job is stated narrowly in the prompt text
  • Nothing in the prose contradicts it
  • Three to five visible anchors, worded identically to last time
  • If you built an intermediate image, the original photo is attached too
  • Fewer than about eight identities in the shot
  • 真人实拍 near the end, if and only if the style is live-action

When it still drifts

If you have done all of the above and the likeness is still wrong, the cause is usually not identity at all. An output that ignores the reference completely (different person, different setting) is almost always a wiring problem rather than a prompting one, and an output that fails identically three times in a row may be a refusal wearing an outage costume. The fault finder covers both.

And if the subject simply will not hold still no matter what you write: that one is not fixable from the prompt. It is a property of what these models are for. Start from a held-pose photo, or use a renderer that freezes by construction.


Every generation on Starrd is a real person's photo going into a cinematic scene, so this is the problem we spend the most time on. The templates there have the reference ordering, the anchor wording and the construction already solved per scene, which is a reasonable place to see the techniques above working before you build your own.

Related: the full prompt handbook, all 50 reference slots.

Frequently Asked Questions

Why does my face change halfway through an AI video?

Usually one of four causes, in rough order of frequency: the identity reference is not first in upload order so the model weights it lower than something else, the reference has no job stated in the prompt so it gets averaged with the other inputs, your prose contradicts the reference, or your source photo shows someone in motion. The last one catches people out most because it has nothing to do with the prompt at all.

What kind of photo works best as an identity reference?

A held pose or a gesture, shot clearly, with the face unobstructed and evenly lit. Standing, sitting, pointing, arms crossed, mid-laugh: anything that is a state rather than a transition. Avoid photos taken mid-stride, mid-jump or mid-turn. Video models complete the action a frame implies, so a photo of an unfinished movement invites the model to finish it, and faces drift most while a subject is moving.

How many identity references should I use?

Three is a good default and more than six rarely helps. What matters is angle coverage rather than count: a front view, a three-quarter and a profile beat six near-identical selfies, because the model needs something to draw on when the camera moves around the subject. Redundant references get averaged, and averaging is itself a source of drift.

Does a character sheet fix identity drift?

It fixes angle coverage, which is a different problem. A character sheet stops the model inventing the far side of a face during a camera move. It does not stop drift caused by the sheet itself, because generating that sheet re-renders the person rather than compositing them, so the likeness has already shifted before the video model sees it. Send the original photo alongside the sheet.

How many people can Seedance keep consistent in one clip?

About eight people or products before consistency starts slipping noticeably. Below that, per-subject identity references work well. Above it you are fighting the model, and the usual fix is to restage the shot so fewer identities need to be held at once, rather than adding more references.

What are the Chinese keywords people use for character consistency?

Seedance is a ByteDance model, and some Chinese phrasing is followed more reliably than its English equivalent. The most useful is 真人实拍, meaning real-person live-action footage, which pushes the model toward photorealism instead of drifting into illustration. Placement matters: put it near the end of the prompt, because leading with it tends to make the model generate Chinese-language dialogue. Leave it out for deliberately non-photoreal styles.

About the author

Brian Bautista · Co-Founder & Creative Director

Brian is co-founder and creative director at Starrd, working as a creative technologist and data scientist. He tracks viral AI-video trends, designs Starrd's scene templates, and writes the deep-dive model comparisons and prompting breakdowns.

Related Articles

Ready to create your own video?

Pick a template, upload your photos, and generate a cinematic AI video in minutes.

Browse Templates