Quick answer
To make a jab jab slip roll uppercut AI video, you need motion control (motion transfer), not a text prompt — a fixed shadowboxing clip drives every punch and slip, and your photo supplies the fighter. The photo has to show a full-body, upright subject so the model has a skeleton to drive. Kling 2.6 Motion Control handles it well. Starrd's Jab Jab Slip Roll Uppercut template does it in one tap from a single photo.
What You're Trying to Make
One photo of you, and out comes a clip of you shadowboxing like a pro fighter — double jab, slip, roll under the punch, snap the uppercut — in your own room, in your own clothes. Same combination, same rhythm, same swagger as the original. You just never had to learn it.
This guide covers where the clip comes from, why it needs motion control instead of a text prompt, the photo that works, and the one-tap way to make yours.
Fastest way — Jab Jab Slip Roll Uppercut on Starrd does the whole thing in one tap: upload one photo and it squares you up into a fighting stance, runs the motion pass, and returns a 20-second clip with the original audio. 100 credits (about $4.99), no prompt to write. Want the full method first? Read on. ↓
Where the Clip Comes From
The source is a home video of Conor McGregor shadowboxing in a dining room — polo shirt, jeans, no gloves, running a fast combination around the dinner table. Fight and MMA pages keep reposting it with the combination as the caption: jab, slip, uppercut, backhand, roll. It works as a trend because it's short, it's readable (you can count every punch), and it's impressive without being out of reach — it looks like something you could almost do.
The AI version keeps all of that and swaps the fighter. Your photo supplies the person; the clip supplies every move.
Why This Is Motion Control, Not a Text Prompt
You can't prompt your way to this one either.
Type "a man shadowboxing" into a text-to-video model and you'll get some punches — in a random order, at a random pace, different every render. The trend isn't "someone boxing." It's this exact combination, beat for beat.
That's what motion control (motion transfer) does. You give it two things:
- A driving clip — the shadowboxing video, which supplies the motion and the audio.
- A still image — you, which supplies the person, the clothes, and the room.
The model estimates a skeleton from the driving clip and maps that movement onto your photo. Same combo, same timing — different fighter.
The Catch — It Needs Your Whole Body
This is where most DIY attempts go wrong.
Motion control estimates a skeleton from your still image. A selfie or a waist-up photo has no legs to drive, so the footwork — the half of this combo that makes it look real — has nothing to land on. The model either invents legs badly or crops you awkwardly.
So before the motion pass, the photo gets transformed: same you, same outfit, same room, but standing full length in a relaxed fighting stance — feet apart, fists up near the chin, elbows clear of the body, head to shoes in frame. One clean, unobstructed silhouette. Then the combo has something to drive.
If your photo is a close-up against a plain wall, that's fine — the stance transform extends the same setting around you so your whole body fits. What it can't fix is something in front of you: a hand on your shoulder, a pet in your arms, or a table edge crossing your body will break the silhouette.
The Fastest Way — Use the Template on Starrd
The Jab Jab Slip Roll Uppercut template packages every step above into one upload — the stance transform, the motion pass, and the original audio.
- Pick one clear photo of yourself. Good light, nothing in front of you. Full body is best, but half-body works.
- Open the Jab Jab Slip Roll Uppercut template in the Starrd app or web library.
- Upload and tap generate. It squares you up into a fighting stance, runs the motion pass on Kling 2.6, and returns a 20-second clip.
Jab Jab Slip Roll Uppercut
Upload one photo and watch yourself shadowbox the full combo like a pro — jabs, slips, rolls and the uppercut. No editing, no prompt writing.
Or, Build It Yourself
Step 1 — Get the driving clip
You need the shadowboxing clip, trimmed to the combination. Its background never appears in the output, and burned-in captions don't transfer — only the motion and the audio come across. One rule: exactly one person in frame. Kling refuses a driving clip with two people in it.
Step 2 — Square up your photo
Run your photo through an image-to-image model with a stance brief:
Transform the subject from the reference photo into a full-body shot of the SAME subject standing TALL and FULLY UPRIGHT in a relaxed boxing stance — NOT sitting, crouching, or leaning. Feet shoulder-width apart, both fists raised loosely near the chin, elbows bent and clearly separated from the torso, hands free and unobstructed, looking straight ahead with calm confident focus. PRESERVE the exact identity (face, hair, skin, clothing) from the reference — only the pose changes. One clear unobstructed silhouette. KEEP THE EXACT SAME background, setting, and lighting as the original; if the photo is a close-up, extend that same setting naturally so the whole body stands inside it. Wide shot with the camera pulled back: the ENTIRE body is visible from the top of the head down to both feet and shoes. Vertical 9:16, photorealistic. No text, no logos, no watermarks, no gloves.
The line doing the most work is the wide-shot framing. Without it, image models love to crop at the hips — which looks fine as a photo and ruins the footwork in the video. "Elbows clearly separated from the torso" matters too: arms pressed against the body give the skeleton estimate nothing to separate.
Step 3 — Run the motion pass
Feed the upright still and the driving clip to Kling 2.6 Motion Control at 720p. Use 2.6, not 3.0 — 3.0 turns every failure into the same opaque error. And skip 480p: Kling's motion-control path rejects it.
The video prompt barely matters — the timing comes from the clip. A line of energy is enough:
The subject shadowboxes like a seasoned fighter, throwing sharp jabs, slipping and rolling under punches and snapping an uppercut, light on their feet with smooth confident footwork. Focused, cocky, explosive energy.
Step 4 — Post it
Vertical, sound on, around 20 seconds. Put the combo in the caption — people search the combination itself. Label it as AI.
Common Mistakes That Tank Your Video
- Prompting it instead of driving it. A described boxer isn't the trend. The exact combination is.
- A waist-up photo with no transform. No legs means no footwork, and the footwork sells it.
- Something crossing your body. A bag strap, a drink, another person's arm — anything in front of you breaks the silhouette.
- A driving clip with two people in it. Kling refuses it outright.
- Rendering at 480p. The motion-control path only takes 720p and up.
Related Reading
- How to Make a Kung Fu Cat AI Video — the same motion-control chassis, with a screaming karate pet.
- How to Make an AI UFC Fighter Video — walkouts, tap-outs and boxing matches with you as the fighter.
- How to Make the "If You Grab Me, I'm Gonna Bite You" AI Pet Video — another one-photo motion-control format.
- Viral AI Video Trends of 2026 — what's climbing right now.