Skip to main content
how towalk it like i talk itwalk it talk itcarpool karaokecar karaokesinging in the carrap in the cargroup videofriend group videosquad videofour friendsroad trip videoai video generatorai group videolandscape ai videoviral tiktok trend

How to Make the Walk It Like I Talk It AI Video (4-Person Carpool Trend)

Four photos, one car, four seats. The carpool singalong rebuilt so you and three of your people are the ones in the front and back — landscape 16:9, 15 seconds, cuts and all. The slot order is the seating plan, so read that part first.

Starrd Team|September 21, 20267 min read

Quick answer

Upload four photos to Starrd's Walk It Like I Talk It template and it puts all four of you in one car doing the carpool singalong. The photo order is the seating plan, running left to right: photo 1 is the front passenger, photos 2 and 3 are the back seat, photo 4 is the driver. It costs 100 credits (about $5) and comes out landscape 16:9 at 15 seconds with the track underneath.

Four photos in, and out comes you and three of your people in one car, trading the verse — four fixed seats, the camera cutting between them, landscape. Here's exactly what you get:

Walk It Like I Talk It, from four photos — four seats, 15 seconds, landscape

Walk It Like I Talk It

Four photos, one car. You and three of yours, running the song. No prompt to write.

Try It

Read This Part First: The Slot Order Is the Seating Plan

Most templates take one photo, or two. This one takes four, and unlike the others the order you upload them in decides where everyone sits. It runs left to right across the frame:

SlotSeatScreen time
Photo 1Front passenger, far leftMost — they lead the song
Photo 2Back seat, leftMiddle
Photo 3Back seat, rightMiddle
Photo 4Driver, far rightLeast, plus a close-up at the end

The upload slots are labelled in the app, so you don't have to hold this in your head. But it is worth knowing photo 1 is the hero — that seat is on screen for roughly two-thirds of the clip, gets the opening shot, and carries the verse. Put your main person there.

Pro Tip

If you only have three people, still fill all four slots — duplicate someone into the back seat rather than leaving a slot empty. And if one friend insists on being the driver, that's slot 4, not slot 2. It reads backwards, and there's a real reason for it below.

Why the Driver Is Slot 4

This is the kind of detail that normally stays buried, but it changes how you use the template, so it's worth a paragraph.

The obvious ordering is front row then back row — photo 1 front passenger, photo 2 driver, photos 3 and 4 in the back. That's what we built first, and on a cast of four people it looked like it worked.

It didn't. Tested against a second, deliberately different cast, the image model re-seated three of the four slots, because when it isn't certain who's who it falls back to placing reference photos left to right across the frame. Front-row-then-back-row fights that. Left-to-right agrees with it.

So the contract matches the model instead of arguing with it. The cost is that the driver — the far-right seat — is the last slot rather than the second.

Being straight about the limits: this improves the odds, it doesn't make them certain. Across our test generations slot 1 has come out right every single time — the hero seat is the stable one. The other three shuffle occasionally, and when they do it's almost always whoever is in slot 2 turning up at the wheel. It's a regenerate-and-move-on problem, not a broken template, but it's worth knowing before you promise a specific friend the driver's seat.

Where the Song Comes From

Walk It Talk It is the Migos single from Culture II (2018), featuring Drake, and the phrase in the hook — "walk it like I talk it" — is the part everyone actually says. The carpool singalong format it's staged in is the one you already know: a camera on the windscreen, four people in an SUV, nobody performing to anything but each other.

That format is the whole reason this template is landscape. A car interior with four people in it does not survive a vertical crop — you lose the back seat or you lose the driver. So this one renders 16:9, which makes it one of only two wide templates on Starrd.

The Fastest Way

Walk It Like I Talk It

Four photos, four seats, 15 seconds landscape. The camera cuts between all of you.

Try It

Four photos, 100 credits, no prompt. The template holds the car, the seats, the cuts and the track; your photos only decide who's in which seat.

Build It Yourself

If you'd rather wire it up by hand, the shape is a video-to-video job: a fixed clip supplies the camera and the performance, and a single composed still supplies the four faces.

1. Make the still first. One landscape image of all four people in a car, one per seat, cleanly separated. This is the step that decides whether the casting is right, and it's cheap to iterate on — generate the still alone and look at it before you pay for any video.

2. Then drive it with a clip. The clip supplies the camera, the cuts and the timing. Your still supplies who the people are.

Image 1 shows the FOUR characters in the car. Video 1 is the performance to copy.DIVISION OF SOURCES — follow this strictly:
From Image 1 take ONLY the four characters' identities and wardrobe, and the car cabin they are riding in.
From Video 1 take EVERYTHING ELSE: the framing, the camera position, WHERE AND WHEN THE CAMERA CUTS, which seats are on screen in each shot, and the moment-to-moment performance.
Where Image 1 and Video 1 disagree about anything other than who the four characters are and what they are wearing, Video 1 WINS.SEATS ARE FIXED. No character ever changes seat, and no character ever appears in a seat that is not theirs.CUTS: the camera cuts between angles at exactly the moments Video 1 cuts, to exactly the same angles. Do not smooth a cut into a drift, and do not invent a cut that Video 1 does not make.FACES: each character's eyes, eyebrows, mouth and jaw follow the face and mouth of the person in their seat in Video 1, word for word, for the whole clip.

Model choice: run this on Seedance 2.0 in its video-to-video mode. It animates several characters independently, which is the whole requirement here — engines that drive every subject in a still from one skeleton will make all four people move in unison, which looks exactly as wrong as it sounds.

Common Mistakes

Uploading a group photo into one slot. Each slot is one person. A photo with two faces in it makes the casting ambiguous and you'll get a blend.

Assuming the driver is slot 2. It's slot 4. See above.

Expecting every cut to land. The edit this was built from cuts fourteen times in fifteen seconds; the driving clip carries seven and a render comes back with seven or eight. All four camera setups are there every time — the fastest flurries are what get smoothed. That's the engine, not your photos.

Cropping it vertical afterwards. You'll cut two people out of their seats. If you need a vertical post, letterbox it rather than crop it.

Padding the audio to a round number. If you're building this yourself: cut the track longer than the video, not shorter. A track that ends first silently truncates the video, and one cut to exactly the same length can still leave you a beat of dead air.

FAQ

Which photo goes where? Left to right: photo 1 front passenger, photo 2 back-left, photo 3 back-right, photo 4 driver. Photo 1 gets the most screen time.

Can I use pets? Not yet — see the note in the FAQ block above. Camden Bop and Hotel Lobby both handle animals today.

Why landscape? Because four people in a car don't fit in a vertical frame without losing two of them.

How long does it take? A few minutes end to end, most of it the video step.

Frequently Asked Questions

Which photo goes in which seat?

The photo order is the seating plan, and it runs left to right across the frame. Photo 1 is the front passenger on the far left — they lead the song and get the most screen time. Photo 2 is the back seat on the left. Photo 3 is the back seat on the right. Photo 4 is the driver on the far right. The upload slots are labelled, so you do not have to memorise it, but the order does matter and swapping two photos swaps two people.

Why is the driver the last slot instead of the second?

Because left to right is the ordering the image model actually follows. The intuitive alternative — driver in slot 2 — was tested on both a human cast and an animal cast and re-seated people in three of four slots, because the model places reference photos left to right across the frame when it is unsure who is who. Matching the slot order to the seat order goes with that instead of fighting it. It is not perfect: slot 1, the front passenger, has landed correctly every time we have tested it, while the other three occasionally shuffle — most often whoever is in slot 2 ending up at the wheel. If that happens, generate it again.

How many photos does it take?

Four, and all four are required. This is the first Starrd template that asks for four. If you only have three people, put your best photo in slot 1 — that is the seat with the most screen time.

Is it vertical or landscape?

Landscape. This one comes out 16:9 at 864 by 496, because a car interior with four people in it does not survive a vertical crop — you would lose two of them. It is one of only two Starrd templates that renders wide. Post it to YouTube, X or a landscape Reel rather than a straight TikTok upload.

How much does it cost?

100 credits, about $5, the same as every other video template. No subscription, and credits never expire. New accounts start with 20 welcome credits, which is not enough for a full video but covers a photo pack or a couple of AI songs.

Does it work on pets?

Not yet. The template renders people reliably, but an all-animal cast currently has the original performers bleeding back into about a quarter of the frames, so we have not shipped it for pets. If you want a pet in a dance template today, Camden Bop and Hotel Lobby both handle animals properly.

What photos work best?

Four clear, well-lit, front-facing photos, one person each. Faces roughly head-and-shoulders is ideal. Avoid group shots in a single slot — each slot is one person, and handing it a photo with two faces in it makes the casting ambiguous.

Why is it 15 seconds?

Because that is the model ceiling and the section runs that long. The clip ends on the ad-lib rather than fading out mid-bar, and the audio is cut two seconds longer than the video so it never truncates early.

Are the camera cuts real?

Yes — the template follows a fixed driving clip that cuts between four setups: a close-up of the front passenger, a two-shot of the back seat, a wide of all four, and a close-up of the driver at the end. The edit it was built from cut fourteen times in fifteen seconds; the driving clip carries seven of those and a render reproduces seven or eight. The fastest flurries get smoothed out, which is the engine rather than anything to do with your photos.

Do I need to label it as AI?

Yes — TikTok, Instagram and YouTube all require AI-generated content to be disclosed. It costs you nothing; the clip reads as a bit anyway and travels fine with the label on.

Part of AI Dance Videos

More guides in this series

Related Articles

Ready to create your own video?

Pick a template, upload your photos, and generate a cinematic AI video in minutes.

Browse Templates