Quick answer
Use a character sheet when you care about who is in the shot: one image showing your subject from several angles and poses, which gives the video model a rounded sense of the person and leaves it free to stage the scene. Use a storyboard grid when the scene has to look a specific way: one image of several panels pinning environment, staging and on-screen elements. Grids buy control at the cost of hard cuts between panels, and any generated reference introduces identity drift because it re-renders your subject rather than compositing them.
Every Seedance guide tells you to attach a reference image and label its job. That advice is correct and it stops one step too early, because the bigger lever is what image you attach.
Two constructions do almost all the work. They fail in opposite directions, and knowing which failure you can live with is most of the decision.
This is one of the few areas where we have genuinely unusual data: our pipeline builds a reference image for every generation, and we have watched both constructions succeed and fail across a lot of paid renders. A caveat worth repeating: that pipeline runs Seedance 2.0, not 2.5. The two constructions behave the same way because the underlying behaviour is about how a video model reads a composite image, not about a version number, but the numbers below come from 2.0.
The two constructions
A character sheet is one image showing your subject from several angles and poses. The standard build is six views: three turnaround angles (front, three-quarter, profile) and three action poses. It answers "who is this person, from any angle" and says nothing about where they are.
A storyboard grid is one image divided into panels, each showing a specific moment of the shot. It answers "what does this scene look like" and pins environment, staging, framing and on-screen elements.
| Character sheet | Storyboard grid | |
|---|---|---|
| Protects | Identity, from any camera angle | Environment, staging, on-screen elements |
| Leaves open | Where the scene happens, how it is staged | Exactly what the subject looks like |
| Model behaviour | Free to stage; motion tends to be continuous | Sequences the panels; tends to cut between them |
| Best for | Transformations, performance, anything camera-orbiting | Stadiums, game UI, broadcast graphics, specific sets |
| Main failure | Scene drifts from what you imagined | Hard cuts, and worse likeness |
When each one wins
The question to ask is not "which is better." It is what can I not afford to get wrong in this shot.
If the answer is the person, use a character sheet. Self-insert work, transformations, anything where a viewer knows the face and will notice it slipping. The sheet gives the model enough angles that a camera orbit does not force it to invent the far side of a head, and because the sheet carries no scene information the model stages the shot itself, which it is good at.
If the answer is the place, use a grid. A stadium with the right scoreboard, a game menu with the right UI, a broadcast frame with a chyron in the right position. These are things the model will approximate badly from prose and reproduce well from a picture.
If both matter equally, do not try to make one image do both jobs. A character sheet plus a separate environment plate as a second reference beats a grid that tries to carry identity and staging at once. That is what the expanded 2.5 reference budget is genuinely for.
The three things nobody mentions
1. A grid sequences by design, so it cuts
This is the one that surprises people. Panels read as an ordered series of moments, so the model treats them as shots to move through and produces hard cuts between them. If you wanted one continuous take, a grid actively fights you.
We hit this hard while testing a camera-orbit shot: a six-panel grid of a single frozen instant, shot from six angles, produced not an orbit but six cuts. Rebuilding the same idea as a four-panel character sheet with an explicit instruction to interpolate rather than cut fixed the cutting completely.
If you must use a grid and you want continuous motion, say so in the prompt in those words: interpolate smoothly between these views, never cut. It is not a guarantee, but the grid's sequencing bias is strong enough that leaving it unaddressed almost always produces cuts.
2. Any generated reference has already lost some likeness
This is the big one for real-person work, and it is structural rather than fixable by prompting.
When you build a character sheet or a grid from someone's photo, the image model re-renders the person. It does not composite the original pixels; it generates a new picture of someone who looks like them. Faces and hair shift in that step, before the video model has seen anything. So the video is not drifting from your subject, it is faithfully rendering an intermediate that already drifted.
Two mitigations, in order of effectiveness:
- Pass the original photo as a second reference alongside the generated sheet. The video model gets both and pulls likeness partway back. This is cheap and it works.
- Skip the intermediate entirely where the shot allows it. Sending the real photograph straight to the video model gives the best likeness available, because nothing has re-rendered the face. The cost is that you lose the staging control the sheet was buying, and you lose any aspect-ratio normalisation.
3. Multi-view consistency holds the clothes and loses the limbs
Across a four to six panel build, current image models hold wardrobe, lighting and scene very well. What drifts is arm position and the stride phase of background people.
That matters more than it sounds, because the video model reads that drift as a motion cue. Two panels of the same person with slightly different arm positions look, to a video model, like the beginning and end of a movement. Sometimes that is a free gift. When you wanted stillness, it is the reason you did not get it.
Building a character sheet
Six views is the sweet spot: enough angles to survive a camera move, few enough that each panel stays large enough to hold facial detail. Under three views the model invents; past eight, each panel gets too small to be useful.
A character reference sheet of the subject, six views on a clean neutral background, arranged 3x2. Top row: front view, three-quarter view, profile view. Head and shoulders to mid-thigh, neutral expression, even soft lighting. Bottom row: three action poses appropriate to the scene, full body, same wardrobe and same lighting as the top row. Consistent identity across all six panels: same face, same hairline, same build, same clothing. No text, no panel borders, no watermarks.
Two rules that matter more than the wording:
- Keep lighting identical across panels. Different lighting per panel reads as different scenes, and the video model will try to reconcile them.
- Keep the expression neutral in the turnaround row. A smile in one view and not another is another false motion cue.
Building a storyboard grid
Fewer panels, larger each. Three to six is the useful range, laid out in reading order because that is the order the model will move through them.
A 2x2 storyboard grid, four panels in reading order, showing one continuous scene. Panel 1: wide establishing shot of the location, subject entering frame left. Panel 2: medium shot, subject centred, [the specific staging that matters]. Panel 3: close-up on the subject's face, same wardrobe and lighting. Panel 4: wide shot from the opposite angle, subject exiting frame right. Identical subject, wardrobe, lighting and location across all four panels. No text overlays, no panel numbers, no watermarks.
Say "no panel numbers" and "no text overlays" explicitly. Image models love adding them, and any text baked into a reference frame tends to survive into the video as garbled lettering.
Match the reference image's aspect ratio to your intended output. A 9:16 reference feeding a 16:9 generation invites recomposition, and the most common symptom is cropped heads. If your tooling exposes aspect on both the reference and the video, set both.
A decision you can make in ten seconds
- Is there anything in the scene the model will get wrong from a description? A specific stadium, a game UI, a broadcast graphic, a branded set. If yes, lean grid.
- Will a viewer know this face? If yes, lean character sheet, and pass the original photo as a second reference regardless of which you choose.
- Does the camera move around the subject? If yes, character sheet. Orbits need angles, and grids give cuts.
- Both 1 and 2 are true? Character sheet for the person, separate environment plate for the place. Two references, two jobs.
Common mistakes
- Building a grid when you needed a sheet, because a grid felt more thorough. More panels is not more control; it is more sequencing.
- Different lighting across panels. The most common self-inflicted wound in both constructions.
- Letting text into the reference. Panel numbers and captions survive into the video as garbled lettering.
- Forgetting the original photo. If you built a generated reference from a real person, send the photo too. It is one extra slot and it is the highest-value slot you will spend.
- Assuming the sheet fixes likeness on its own. It fixes angle coverage. Likeness is a separate fight, covered in keeping a real person's face consistent.
Where this fits
The reference image is one layer of a larger stack. Once you have built the right one, the full reference guide covers how to spend the other 49 slots, and the prompt handbook covers the grammar that goes around them. If your reference is being ignored entirely rather than drifting, that is usually a wiring problem, and the fault finder has it.
We build a reference image for every generation that runs through Starrd, which is why we have opinions about this. If you would rather not build one, the templates there have the construction already chosen per scene: character sheets where identity carries the shot, grids where the environment does.