Skip to main content
how totonight we're getting naughtygetting naughtynaughty dancenaughty dance trendtonight we're getting nattytonight we getting notigetting notidate night dancecouple dancecouples dance trendboyfriend girlfriend danceai dance videoai couple videoai video generatorviral tiktok trendtiktok dance 2026

How to Make the Tonight We're Getting Naughty AI Video (2026 Dance Trend)

Two photos, one couple's routine — and it happens in your own kitchen, not a studio set. The Tonight We're Getting Naughty trend, rebuilt from your photos so the room behind you is actually yours. No prompt, no green screen, no CapCut.

Starrd Team|September 20, 20268 min read

Quick answer

Upload two photos to Starrd's Tonight We're Getting Naughty template and it puts both of you side by side doing the couple's routine — in the real room from your photos, not a stock set. It costs 100 credits (about $5), takes two photos, and comes out vertical 9:16 at 9 seconds with the real sound underneath.

Two photos in, and out comes the two of you doing the routine in your own kitchen — same walls, same floor, same light, one locked vertical shot. Here's exactly what you get:

The Tonight We're Getting Naughty routine, from two photos — in the room from the uploaded photo, 9 seconds

Tonight We're Getting Naughty

Two photos, your own place. You and your person, running the routine — real room, real sound. No prompt to write.

Try It

What the Sound Is

A short, fast couple's dance that broke on TikTok in 2026 off the hook line tonight we're getting naughty. It travelled the way these routines usually do — a few seconds of choreography, easy enough to learn in one watch, hard enough that a pair has to actually rehearse it once.

The genuinely useful thing to know is that nobody agrees on how to spell it. It gets searched as naughty, as natty, and as noti, and TikTok carries separate discovery pages for each. That is unusual, and it is free reach: most people caption one spelling and quietly miss the other two.

Is This Still a Trend?

Yes, and it behaves differently from the music-video trends. There's no set to recognise and no artist cameo to reference — the whole appeal is that it's two people who actually know each other, in a real house, slightly out of sync on the hard beat.

That's also why it survived into the AI era instead of being flattened by it. A single-face swap gets you nowhere: swap one head into a couple's dance and the other dancer is still a stranger, which is exactly the thing the format is about.

The Fastest Way

Tonight We're Getting Naughty

Two photos. You and your person doing the routine, in your own room. 9 seconds, vertical, real sound.

Try It

Upload two photos, one person each. You get back a 9-second vertical clip with the real sound on it. No prompt, no CapCut, no clearing the furniture out of your living room.

Why It Uses Your Room, Not a Set

Most templates put you somewhere: a booth, a red carpet, a stadium. This one does the opposite — it reads the background of your photo and rebuilds that room to dance in.

It matters more than it sounds. A couple's dance filmed in an obviously fake studio reads as a filter. The same routine in a kitchen with the wrong number of mugs on the shelf reads as a video someone actually made. The format's credibility comes from the room being unremarkable and real.

Three rules it follows:

  • Your room, not a nicer version of it. It won't tidy up, restyle or upgrade what's behind you. A cluttered kitchen stays a cluttered kitchen.
  • Two photos, two different places? The first one wins. Both people end up in one room together — never a split frame, never a different room behind each person.
  • No usable background, and only then a fallback. A blank studio wall or a tight head crop has nothing to rebuild, so it uses a warm living room instead of stranding you in a grey void. A plain bedroom with a bed and a lamp counts as usable — it will use that.

So the photo you upload first decides the location. Choose accordingly.

Build It Yourself

If you'd rather assemble it by hand, here's the shape of it.

1. Build the still first, with the room already in it. Don't animate your raw selfies. Generate one still of both people standing in the room, then animate that:

A photorealistic VERTICAL 9:16 photograph of the TWO subjects from the reference photos dancing together in a real room.LEFT subject: the first subject from the reference photos, standing TALL and FULLY UPRIGHT, arms raised and clearly separated from the body at about shoulder height, mid-step in a dance. The face is awake and animated — eyes open and lit up, plainly enjoying themselves. Not a stiff, blank or deadpan face.RIGHT subject: the second subject from the reference photos, same posture and the same animated face.PRESERVE the exact identity of each subject — face, hair, skin tone and clothing — exactly as in their reference photo. Do NOT blend the two subjects together and do NOT swap their clothing.They stand side by side with a clear gap of empty floor between them, neither overlapping nor touching. BOTH are shown FULL BODY from head to feet, feet flat on the floor, filling most of the frame height.SETTING — USE THE REAL PLACE FROM THE REFERENCE PHOTOS. Look at what is BEHIND the subjects in their photos and rebuild that same real place: the same wall colour, flooring, furniture, windows, lighting and decor, with the light coming from the same direction. Extend it naturally to give them floor to dance on. Keep it recognisably THEIR place — do not tidy it, redecorate it or swap it for a nicer version. If the two photos show different places, use the place from the FIRST subject's photo for both.Photorealistic, sharp focus, natural colors. Vertical 9:16. No text, no captions, no logos, no watermarks, no other people.

2. Drive it with the routine itself. Take that still into a video model that accepts both an image reference and a video reference, and hand it a clip of the dance you want copied. Every step and every arm movement comes from the driving clip — not from your prompt.

3. Lock left to left and right to right. With two dancers you have to say which copies which, explicitly, or the model averages them into the same moves and you get a mirror instead of a duet.

4. Ask for the faces, by part, not by feeling. This is the one people miss. A copied body over a dead face is uncannier than no video at all. Say that each character's eyes, eyebrows, mouth and whole expression follow their own dancer, moment for moment — and name no particular expression. Write smiling and you'll get one frozen smile held for ten seconds, which is the same deadpan problem arrived at from the other side.

5. Do not describe the moves. The counterintuitive one. Writing bobbing on the beat or stepping side to side makes the output worse — the model satisfies your description with the smallest possible version of it and discards the real choreography.

Pro Tip

Same goes for framing. Adding "full body", "feet on the floor" or "static locked-off camera" to the video prompt measurably suppresses the movement. Put that language in the still prompt, where it belongs, and keep the video prompt about nothing but the performance.

6. Put the real sound on afterward. Render silent and lay the track over the top. Generated audio over a sound this recognisable is instantly obvious — and asking a model to synthesise audio over a copyrighted reference clip tends to fail moderation anyway.

Common Mistakes

Uploading the wrong photo first. The first photo's room is the room. If you want the kitchen, lead with the kitchen.

Two tight head crops. Nothing behind either face means nothing to rebuild, and you get the fallback set instead of your house.

Uploading one group photo for both slots. Two photos, one person in each. Hand it a group shot and the model has to guess who's who.

Describing the choreography. Covered above, and it's the single biggest quality killer on two-person templates.

Naming an emotion instead of the facial parts. "Make them look flirty" gets you one held expression. Point the eyes and mouth at the driving clip and let the routine supply the rest.

Letting the model generate the audio. Everyone knows this sound. A near-miss is worse than silence.

FAQ

What is the Tonight We're Getting Naughty trend? A two-person dance routine, usually filmed at home in one vertical take. The AI version rebuilds both people from two photos.

Is it naughty, natty or noti? All three are searched, and TikTok has pages for each. Put more than one spelling in your caption.

Why two photos? Two dancers, mapped independently — left to left, right to right. Single-face swaps can't do the second person.

Does it use my room? Yes. It rebuilds the background from your photo, and only falls back to a set room if there's nothing usable behind you.

What if the photos are from different rooms? The first photo's room wins, and both of you end up in it.

How much? 100 credits, about $5. Credits never expire.

Frequently Asked Questions

What is the Tonight We're Getting Naughty trend?

It's a two-person dance trend — a couple, siblings, roommates or friends running a short synced routine side by side, usually filmed at home on a phone in one locked vertical shot. The AI version takes two photos and rebuilds both people doing the routine, in the room from your own photo.

Is it spelled naughty, natty or noti?

All three, and that matters if you are writing captions. People search the sound as 'tonight we're getting naughty', 'tonight we're getting natty' and 'tonight we getting noti', and TikTok carries separate discovery pages for each spelling. If you want the video found, put more than one spelling in your caption or hashtags.

Why does it take two photos?

Because the routine has two dancers and each one is mapped independently — the person on the left copies the left dancer, the person on the right copies the right. That is the part single-face CapCut versions cannot do: they swap one face onto existing footage, so the second person is still a stranger.

Does it use my own room?

Yes, and that is the unusual part. Most templates drop you onto a fixed set. This one reads the background of your photo — the walls, the floor, the furniture, the light — and rebuilds that room to dance in. If your photos have no usable background, it falls back to a warm living room rather than leaving you in a void.

What if the two photos were taken in different places?

The first photo's room wins, and both people end up in it together. It never splits the frame or puts a different room behind each person. So upload the photo of the room you actually want first.

How much does it cost?

100 credits, about $5, same as every other video template. No subscription, and credits never expire. New accounts start with 20 welcome credits, which covers a photo pack or a couple of AI songs but not a full video.

What photos work best?

Two clear, well-lit, front-facing photos, one person per photo. Unlike most templates, it is worth giving this one some room in the frame — a shot with your actual kitchen, bedroom or living room visible behind you gives it something to rebuild. A tight head crop on a blank wall forces the fallback set.

Why does it only need 9 seconds?

Because the routine is 9 seconds. The clip is as long as the dance and stops there, rather than padding out to a round number and leaving the last beats — and the last of the music — to run out into silence.

Do I need to label it as AI?

Yes — TikTok, Instagram and YouTube all require AI-generated content to be disclosed. It costs you nothing here; the clip is obviously a bit, and it travels fine with the label on.

Part of AI Dance Videos

More guides in this series

Related Articles

Ready to create your own video?

Pick a template, upload your photos, and generate a cinematic AI video in minutes.

Browse Templates