Quick answer
Nano Banana Pro inside Gemini is the strongest model for the photoreal 'real to GTA' look — your photo re-rendered as an actual in-game screenshot rather than a drawing. Direct it in plain sentences, not keyword stacks or --flags, which it ignores. Upload one clear, well-lit selfie, explicitly tell it to keep your face recognizable and not to beautify you, name the Grand Theft Auto VI render style, lock Vice City in one clause, state the aspect ratio, and end with 'no HUD, no on-screen text, no logos.'
Nano Banana Pro is the "real to GTA" model
There are two GTA AI looks, and people constantly try to make one model do both.
The illustrated look is the original "GTA Me" loading-screen portrait — cel-shaded, bold outlines, painted cover-art finish. Midjourney still owns that one.
The photoreal look — the currently-viral "real to GTA" trend — turns your photo into what reads as an actual in-game screenshot from a modern console engine. Not a drawing. That's Nano Banana Pro, inside Gemini, and it is not close.
Two reasons it wins here:
- It holds a real face. GPT Image 2 and Midjourney both drift your likeness toward a generic attractive average. Nano Banana Pro holds bone structure, skin tone, and hairline far more faithfully — which is the entire point when the shot is you in Vice City.
- It renders materials, not illustration. Wet asphalt, subsurface skin, screen-space reflections, physically based surfaces. That's what sells "in-game screenshot" over "AI picture."
Where it loses: text inside the image. Wanted posters, HUD overlays, mission cards, menu screens — Nano Banana Pro smears the lettering. Use ChatGPT and GPT Image 2 for those.
The prompting rule people get wrong
Nano Banana Pro ignores keyword stacks.
Prompts that look like this do almost nothing:
GTA 6, vice city, neon, 8k, hyperrealistic, cinematic, unreal engine, masterpiece, trending on artstation, --ar 16:9 --v 6
That's Midjourney grammar. Nano Banana Pro doesn't parse --flags at all, and comma-separated tag soup gets averaged into mush. It responds to plain directing sentences — the way you'd brief a photographer.
Same shot, phrased for the model:
Re-render this person as a photorealistic character in Grand Theft Auto VI. Third-person game camera, standing on a wet neon-lit Vice City street at night. Keep their face clearly recognizable. 16:9, no HUD.
If you're porting a prompt from Midjourney, strip every flag and rewrite the tag list as one or two sentences. That single edit fixes most "Nano Banana ignored my prompt" complaints.
Copy-paste GTA 6 prompts for Gemini
Photoreal in-game screenshot (the real-to-GTA look)
Re-render the person in this photo as a photorealistic in-game character from Grand Theft Auto VI. Third-person over-the-shoulder game camera, mid-shot. They stand on a wet neon-lit Vice City street at night — art-deco storefronts, palm trees, a parked convertible, puddle reflections, light rain in the air. Render it like a modern console game engine: physically based materials, subtle subsurface scattering on skin, screen-space reflections, soft volumetric haze around the neon, filmic magenta-and-teal colour grade, slight film grain. Keep my face, hairstyle and skin tone clearly recognizable — do not beautify or replace my features. Give me a cool, confident, moody main-character look: hard cold stare into the camera, mouth closed, absolutely not smiling. 16:9. No HUD, no on-screen text, no logos, no watermark.
Golden-hour version
Re-render the person in this photo as a photorealistic Grand Theft Auto VI character, shot at golden hour instead of at night. Third-person game camera, low angle. They lean against an invented convertible muscle car on a palm-lined Vice City beachfront boulevard, ocean glittering behind, pastel art-deco motels down the strip, long warm shadows across the asphalt. Modern console engine render — physically based materials, realistic skin, warm saturated grade with blown highlights, heat haze, film grain. Keep my face, hairstyle and skin tone clearly recognizable, do not beautify. Confident neutral expression, mouth closed. 16:9. No HUD, no on-screen text, no logos.
Couple / two-shot
Re-render the two people in these photos as photorealistic Grand Theft Auto VI characters standing together in a cutscene two-shot. Medium shot, waist up, both facing camera on a neon-lit Vice City rooftop at night, city skyline and palm trees behind, wet ground reflecting magenta and teal signage. Modern console game engine render — physically based materials, realistic skin, cinematic depth of field, filmic grade, film grain. Keep both faces, hairstyles and skin tones clearly recognizable — do not beautify or replace their features. Both look straight into camera, confident, mouths closed, not smiling. 16:9. No HUD, no on-screen text, no logos.
Your dog as a Vice City character
Re-render the pet in this photo as a photorealistic character in Grand Theft Auto VI. It sits upright in the passenger seat of an invented convertible on a neon-lit Vice City boulevard at night, one paw on the door, city lights streaking past behind. Modern console game engine render — realistic fur with individual strand detail, physically based materials, magenta-and-teal neon grade, shallow depth of field, film grain. Keep the pet's exact breed, markings, fur colour and face clearly recognizable — do not change its coat pattern. 16:9. No HUD, no on-screen text, no logos.
Getting your face to survive
Likeness is the whole game with Nano Banana Pro, and most of it is decided before the prompt runs.
The upload matters more than the wording:
- Front-facing, eyes visible, no sunglasses
- Even light — hard shadow across half the face costs likeness immediately
- Face large in frame; a full-body shot from ten feet away gives the model almost nothing to work with
- One person per photo unless the prompt is explicitly a two-shot
The wording that matters:
Keep my face, hairstyle and skin tone clearly recognizable — do not beautify or replace my features.
Include that sentence verbatim. Without it, the model quietly "improves" you into someone else and the result stops being funny.
Don't ask for a smile. GTA leads never grin, and a smiling render instantly reads as an AI photo rather than a game character. "Mouth closed, not smiling" is doing real work in every prompt above.
Which Google model for which job
| You want | Use |
|---|---|
| Photoreal "real to GTA" stills, your face | Nano Banana Pro (Gemini) |
| Text inside the image — posters, HUD, menus | GPT Image 2 via ChatGPT |
| Illustrated cel-shaded loading-screen art | Midjourney V8.1 |
| A single cinematic 4K shot with audio | Veo 3.1 |
| A multi-shot trailer up to 15 seconds | Kling 3.0 |
Gemini has no path to a full GTA trailer. Veo 3.1 caps around 8 seconds per shot, which is one beat, not a sequence — and stitching beats while keeping one face consistent across every cut is where prompt-by-prompt workflows fall apart.
GTA 6 Trailer
Upload one selfie and get a cinematic 15-second GTA 6 trailer starring you — neon Vice City, the crew, the heist, the getaway. Your face stays consistent across every cut. No prompts, no model-picking.
Keep going
- All 38 GTA 6 AI prompts — every prompt shown with the render it produced
- GTA 6 prompts for ChatGPT — and why ChatGPT refuses GTA prompts
- Turn your photo into a GTA 6 character — nine styles, real before-and-afters
- Make a GTA 6 poster with AI — cover-art prompts specifically
Fan-made parody prompts in the style of Grand Theft Auto VI. Not affiliated with or endorsed by Rockstar Games. None of these prompts request official assets, logos, or wordmarks.