Quick answer
To turn a photo into a song, you describe the photo in words and pair that description with a stance: hype (treat the subject like a champion), roast (target the choices in the frame, never the body), love (sincere and warm), or narrator (nature-documentary voice over something mundane). Feed the description plus the stance into an AI song generator as your lyric brief. The description is the real input, because the model cannot see your picture.
One photo, one stance, one genre. Here is what that combination actually produces:


Two Strings and a Dream
UK Drill
That vest is a suggestion…the tank top has no tank…them shorts got no follow through
The Problem With Photos
A screenshot of a group chat basically writes its own song. The words are already there, the timeline is already there, and two personalities are already fighting on the page. The model is arranging material, not inventing it. That is why turning your texts into a song works so reliably.
A photo gives you none of that. It is one frozen instant with no dialogue, no before, and no after. Hand a picture to a song model with no further instruction and you get exactly what you asked for: a song about a picture. Pleasant, forgettable, about nobody.
The fix is to stop thinking of the photo as the content. The photo is the subject. The content is your stance toward it.
The Four Stances
There are four angles that consistently produce a song worth sending to someone. Pick one before you write a single word of the brief, and commit to it. A song that is half roast and half love song is neither.
1. Hype
The photo is heroic. Whatever is in the frame, you treat that person like a champion walking out into an arena with the crowd on its feet.
This works because it is almost always slightly untrue, and that gap is the joke. A friend holding a slightly burnt tray of cookies, scored like the last five minutes of a title fight. The stance is total sincerity applied to a small moment. Do not wink at it. The instant the lyrics acknowledge that this is not actually a big deal, the whole thing deflates.
Good for: gym photos, graduations, someone finishing something hard, someone finishing something trivial with great effort.
2. Roast
The photo is evidence. The song is the cross-examination.
This is the funniest of the four and the only one that can go badly wrong, so it has a rule, below.
Good for: outfits, poses, the six-monitor battlestation, the LinkedIn headshot, the car everyone has opinions about, the haircut phase.
3. Love
Sincere, warm, no punchline. A partner, a kid, a parent, a pet who is getting old.
The trap is generic sweetness. Words like always and forever produce a greeting card. Specifics are what make someone cry: the particular chair, the ugly coat, the way they hold a mug. Feed the small true things and let the model handle the sentiment. It is much better at sentiment than at specifics.
4. Narrator
Nature-documentary voice, applied to something completely mundane. Hushed, reverent, faintly amazed at ordinary behavior.
Here, in the soft light of the kitchen, the creature approaches the bowl. It has done this four thousand times. It approaches with the same caution as the first.
This one is extremely strong on pets and surprisingly strong on people doing office work. The humor is entirely in the mismatch between the gravity of the voice and the nothingness of the event.
The Roast Rule
Roast the choices. Never the body.
Target the outfit, the pose, the caption, the sunglasses indoors, the gaming setup with more screens than a small airport. Do not touch weight, height, skin, hair loss, teeth, or a face.
This is not only a decency rule, though it is that. It is a craft rule, and it is the difference between a song someone reposts and a song that quietly ends a friendship.
A choice is something the subject decided. Making fun of a decision is a joke the person who made it can laugh at with everyone else, because the joke says look at this thing you did, and they did do it, and it was funny. Everyone stays in the room.
A body is not a decision. A joke about one says look at what you are, and there is no version of that the subject gets to enjoy. It stops being a bit and becomes an attack, and the audience feels it immediately. That is the fastest possible route to something nobody will share, including the person who made it.
The practical upside: choice-based roasts are simply funnier, because they are specific. "Bad outfit" is nothing. "Two shoelaces holding up a whole personality" is a song.
Here is a real example. A gym mirror selfie: bodybuilder, string tank top with straps roughly the width of shoelaces, extremely short shorts. Roast stance, UK drill. It came back as a track called "Two Strings and a Dream", and the standout line was "the tank top has no tank."
He is enormous and the song never mentions it. Every hit lands on the garment, which he chose, and it is funnier than anything about his body would have been.
Describing the Photo
The model cannot see your picture. Whatever description reaches it is the photo, as far as the song is concerned. This is the step people rush and then blame the model for.
Write it so a stranger could picture the scene. Include:
- Setting. Gym mirror, kitchen at night, office desk, back seat of a car.
- The subject and what they are doing. Standing, flexing, mid-sentence, asleep, eating.
- Clothing and grooming, in detail. This is where roast material lives.
- Expression. Deadly serious is comedy gold. Say so if it is there.
- Objects in frame. The energy drink, the seven cables, the single sad plant.
- The odd detail. The thing that makes this photo different from a thousand near-identical ones.
That last one matters most. The odd detail is nearly always where the best line comes from. Shoelace straps were the odd detail, and they became the title.
Specific nouns beat piles of adjectives. Ten concrete things in the frame will outperform two paragraphs about the vibe every time.
Write the description, then delete every word that could apply to any other photo. Whatever survives is the actual song.
Pick a Genre That Fights the Photo
The reflex is to match: gym photo, gym music. Resist it. Matching produces an advertisement.
Contrast produces a joke, and the joke is the mismatch in emotional scale.
- Toddler with spaghetti on their face, scored as a soaring operatic aria.
- Someone's disastrous parallel park, as a mournful acoustic ballad.
- A cat sitting in a cardboard box, as an eight-minute prog epic with a key change.
- The gym selfie above, as UK drill, because drill is deadly serious and the subject of the song is two shoelaces.
The rule of thumb: pick the genre whose emotional register is the wrong size for what is happening in the frame. Too big is funnier than too small.
The exception is the love stance. Sincerity should not fight anything. Match the mood there and let the details do the work.
Why Pets Win
If you want the thing people actually send to each other, use a pet.
A roast of a friend has a target, and that target may or may not be in on it. A love song about a partner is often a private thing that feels strange on a public feed. Both come with a small social cost before you post.
A pet has none. Nobody can be embarrassed on a dog's behalf. A grave documentary narration about a cat guarding an empty box has no victim and nothing to explain. The Narrator and Hype stances in particular were built for animals: the comedic engine is treating a creature with no idea what is happening as a figure of great consequence. That is why the pet version of nearly every one of these formats outruns the human version.
A Working Brief
Put together, the brief you hand any song tool looks like this:
Stance: Roast. Genre: UK drill, serious delivery, no comedy voice. Photo: Gym mirror selfie. Very large bodybuilder, mid-flex, dead serious expression. String tank top with straps about as thick as shoelaces, barely covering anything. Extremely short shorts. Phone covering part of his face. Rows of dumbbells behind him, fluorescent lighting. Target: The outfit only. Nothing about his body. He is genuinely huge and the song should never mention it. Hook idea: The tank top has no tank.
Stance, genre, description, target, hook. That is the whole format, and it transfers to any tool you already use.
Where Starrd Fits
Starrd is an AI video app: you upload a photo and it makes short cinematic videos from templates. A song feature built around these four stances is in development, but it has not launched, so there is nothing to sign up for yet.
The video templates are live now, and the same thinking applies to them. Pick your stance before you pick your template, describe the photo like the model has never seen it, and if you are going for a laugh, aim it at a choice.
Related Reading
- How to Turn Your Texts Into a Song: the pillar guide for this cluster, and the easier place to start.
- Viral AI Video Trends of 2026: what is climbing right now.