Skip to main content
gta 6 prompt geminigta 6 gemini promptnano banana gta 6nano banana pro gta promptgta 6 ai promptsreal to gtaput yourself in gta 6gemini gta 6 image

GTA 6 Prompts for Gemini & Nano Banana Pro (Real-to-GTA Look)

Copy-paste GTA 6 prompts written for Gemini's Nano Banana Pro — the photoreal 'real to GTA' trend. How to keep your face recognizable, why keyword stacks fail, and the prompts that actually generate.

Brian Bautista · Co-Founder & Creative Director|August 17, 20267 min read

Quick answer

Nano Banana Pro inside Gemini is the strongest model for the photoreal 'real to GTA' look — your photo re-rendered as an actual in-game screenshot rather than a drawing. Direct it in plain sentences, not keyword stacks or --flags, which it ignores. Upload one clear, well-lit selfie, explicitly tell it to keep your face recognizable and not to beautify you, name the Grand Theft Auto VI render style, lock Vice City in one clause, state the aspect ratio, and end with 'no HUD, no on-screen text, no logos.'

Nano Banana Pro is the "real to GTA" model

There are two GTA AI looks, and people constantly try to make one model do both.

The illustrated look is the original "GTA Me" loading-screen portrait — cel-shaded, bold outlines, painted cover-art finish. Midjourney still owns that one.

The photoreal look — the currently-viral "real to GTA" trend — turns your photo into what reads as an actual in-game screenshot from a modern console engine. Not a drawing. That's Nano Banana Pro, inside Gemini, and it is not close.

Two reasons it wins here:

  1. It holds a real face. GPT Image 2 and Midjourney both drift your likeness toward a generic attractive average. Nano Banana Pro holds bone structure, skin tone, and hairline far more faithfully — which is the entire point when the shot is you in Vice City.
  2. It renders materials, not illustration. Wet asphalt, subsurface skin, screen-space reflections, physically based surfaces. That's what sells "in-game screenshot" over "AI picture."

Where it loses: text inside the image. Wanted posters, HUD overlays, mission cards, menu screens — Nano Banana Pro smears the lettering. Use ChatGPT and GPT Image 2 for those.

The prompting rule people get wrong

Nano Banana Pro ignores keyword stacks.

Prompts that look like this do almost nothing:

GTA 6, vice city, neon, 8k, hyperrealistic, cinematic, unreal engine, masterpiece, trending on artstation, --ar 16:9 --v 6

That's Midjourney grammar. Nano Banana Pro doesn't parse --flags at all, and comma-separated tag soup gets averaged into mush. It responds to plain directing sentences — the way you'd brief a photographer.

Same shot, phrased for the model:

Re-render this person as a photorealistic character in Grand Theft Auto VI. Third-person game camera, standing on a wet neon-lit Vice City street at night. Keep their face clearly recognizable. 16:9, no HUD.

Pro Tip

If you're porting a prompt from Midjourney, strip every flag and rewrite the tag list as one or two sentences. That single edit fixes most "Nano Banana ignored my prompt" complaints.

Copy-paste GTA 6 prompts for Gemini

Photoreal in-game screenshot (the real-to-GTA look)

The one everyone is actually looking for
Re-render the person in this photo as a photorealistic in-game character from Grand Theft Auto VI. Third-person over-the-shoulder game camera, mid-shot. They stand on a wet neon-lit Vice City street at night — art-deco storefronts, palm trees, a parked convertible, puddle reflections, light rain in the air. Render it like a modern console game engine: physically based materials, subtle subsurface scattering on skin, screen-space reflections, soft volumetric haze around the neon, filmic magenta-and-teal colour grade, slight film grain. Keep my face, hairstyle and skin tone clearly recognizable — do not beautify or replace my features. Give me a cool, confident, moody main-character look: hard cold stare into the camera, mouth closed, absolutely not smiling. 16:9. No HUD, no on-screen text, no logos, no watermark.

Golden-hour version

Daylight, because everyone else's is at night
Re-render the person in this photo as a photorealistic Grand Theft Auto VI character, shot at golden hour instead of at night. Third-person game camera, low angle. They lean against an invented convertible muscle car on a palm-lined Vice City beachfront boulevard, ocean glittering behind, pastel art-deco motels down the strip, long warm shadows across the asphalt. Modern console engine render — physically based materials, realistic skin, warm saturated grade with blown highlights, heat haze, film grain. Keep my face, hairstyle and skin tone clearly recognizable, do not beautify. Confident neutral expression, mouth closed. 16:9. No HUD, no on-screen text, no logos.

Couple / two-shot

Two photos in, one Vice City cutscene out
Re-render the two people in these photos as photorealistic Grand Theft Auto VI characters standing together in a cutscene two-shot. Medium shot, waist up, both facing camera on a neon-lit Vice City rooftop at night, city skyline and palm trees behind, wet ground reflecting magenta and teal signage. Modern console game engine render — physically based materials, realistic skin, cinematic depth of field, filmic grade, film grain. Keep both faces, hairstyles and skin tones clearly recognizable — do not beautify or replace their features. Both look straight into camera, confident, mouths closed, not smiling. 16:9. No HUD, no on-screen text, no logos.

Your dog as a Vice City character

Absurdly reliable, absurdly shareable
Re-render the pet in this photo as a photorealistic character in Grand Theft Auto VI. It sits upright in the passenger seat of an invented convertible on a neon-lit Vice City boulevard at night, one paw on the door, city lights streaking past behind. Modern console game engine render — realistic fur with individual strand detail, physically based materials, magenta-and-teal neon grade, shallow depth of field, film grain. Keep the pet's exact breed, markings, fur colour and face clearly recognizable — do not change its coat pattern. 16:9. No HUD, no on-screen text, no logos.

Getting your face to survive

Likeness is the whole game with Nano Banana Pro, and most of it is decided before the prompt runs.

The upload matters more than the wording:

  • Front-facing, eyes visible, no sunglasses
  • Even light — hard shadow across half the face costs likeness immediately
  • Face large in frame; a full-body shot from ten feet away gives the model almost nothing to work with
  • One person per photo unless the prompt is explicitly a two-shot

The wording that matters:

Keep my face, hairstyle and skin tone clearly recognizable — do not beautify or replace my features.

Include that sentence verbatim. Without it, the model quietly "improves" you into someone else and the result stops being funny.

Warning

Don't ask for a smile. GTA leads never grin, and a smiling render instantly reads as an AI photo rather than a game character. "Mouth closed, not smiling" is doing real work in every prompt above.

Which Google model for which job

You wantUse
Photoreal "real to GTA" stills, your faceNano Banana Pro (Gemini)
Text inside the image — posters, HUD, menusGPT Image 2 via ChatGPT
Illustrated cel-shaded loading-screen artMidjourney V8.1
A single cinematic 4K shot with audioVeo 3.1
A multi-shot trailer up to 15 secondsKling 3.0

Gemini has no path to a full GTA trailer. Veo 3.1 caps around 8 seconds per shot, which is one beat, not a sequence — and stitching beats while keeping one face consistent across every cut is where prompt-by-prompt workflows fall apart.

GTA 6 Trailer

Upload one selfie and get a cinematic 15-second GTA 6 trailer starring you — neon Vice City, the crew, the heist, the getaway. Your face stays consistent across every cut. No prompts, no model-picking.

Try It

Keep going


Fan-made parody prompts in the style of Grand Theft Auto VI. Not affiliated with or endorsed by Rockstar Games. None of these prompts request official assets, logos, or wordmarks.

Frequently Asked Questions

Can Gemini make GTA 6 images?

Yes. Gemini generates images through Nano Banana Pro, and it is currently the best model available for the photoreal 'real to GTA' look — your photo turned into what reads as an actual in-game screenshot rather than an illustration. It also holds an uploaded face more faithfully than GPT Image 2 or Midjourney, which is why the viral real-to-GTA trend runs on it.

How do I prompt Nano Banana Pro for GTA 6?

Write plain directing sentences, not keyword stacks. Nano Banana Pro ignores comma-separated tag lists and Midjourney-style --flags; it responds to instructions phrased the way you would brief a photographer. Name the Grand Theft Auto VI render style, lock the world in one clause, state what to keep about the person's face, give the aspect ratio, and end with the negatives.

Why does Nano Banana Pro change my face?

Because image models default to 'improving' faces into generic attractive ones unless you forbid it. Add the instruction explicitly: keep my face, hairstyle and skin tone clearly recognizable, do not beautify or replace my features. Also upload a clear, well-lit, front-facing photo — heavy shadow, sunglasses, or a small face in frame all cost likeness before the prompt even runs.

Is Nano Banana Pro better than ChatGPT for GTA prompts?

For the photoreal look and for keeping your face, yes. For anything with text baked into the image — loading screens, wanted posters, mission cards, menus — GPT Image 2 inside ChatGPT is clearly better, because its in-image typography is more accurate. Pick by shot type rather than by loyalty to one tool.

What is the 'real to GTA' trend?

It is the newer of the two GTA AI looks. The original is the illustrated 'GTA Me' loading-screen portrait — cel-shaded, bold outlines, painted cover-art finish. The real-to-GTA version instead makes your photo look like a photorealistic in-game screenshot from a modern console engine, with no HUD. Nano Banana Pro is the model that made it work reliably.

Can Gemini make a GTA 6 video?

Not through Nano Banana Pro, which is an image model. Google's video model is Veo 3.1 — 4K with native audio, capped around 8 seconds per shot, so it suits single cinematic beats rather than full trailers. For anything trailer-length use Kling 3.0, which handles multi-shot sequences up to 15 seconds.

About the author

Brian Bautista · Co-Founder & Creative Director

Brian is co-founder and creative director at Starrd, working as a creative technologist and data scientist. He tracks viral AI-video trends, designs Starrd's scene templates, and writes the deep-dive model comparisons and prompting breakdowns.

Part of GTA 6 AI Videos

More guides in this series

Related Articles

Ready to create your own video?

Pick a template, upload your photos, and generate a cinematic AI video in minutes.

Browse Templates