- What "aesthetic" means to a model
- The five-layer structure I use
- A real example from a Christmas campaign that went sideways first
- The words that quietly wreck your prompts
- Lighting and camera language that moves the needle
- How many attempts it really takes
- Matching the prompt style to the tool you're using
- Quick reference: aesthetics people ask for most, translated
- The one habit that changes everything
- Frequently asked questions
- Where to check the details
The short version: aesthetic photo prompts work when you stop using the word "aesthetic" and start describing wardrobe, lighting, lens, colour grade and mood separately, in that order. Vague mood words give you generic, plasticky output. Specific, layered detail gives you something you would post or sell.
What "aesthetic" means to a model
Here is the bit nobody tells you when you first start typing prompts into Midjourney or ChatGPT's image tool: the word "aesthetic" means nothing to the model. It is a human shorthand for a whole bundle of decisions we make instinctively, wardrobe, colour palette, lighting quality, camera angle, texture, era. When you type "aesthetic photo of a woman in a cafe," the model has to guess at all of that, and it guesses at the average of every stock photo and Pinterest board it was trained on. That average is smooth skin, warm filter, soft focus and a slightly plastic look that anyone who spends time online can spot in under a second.
So the real job of writing a good prompt is translation. You are translating a vague feeling in your head into concrete, filmable detail. If you can't describe what you want to a human photographer in a sentence, the AI tool has no chance either.
The five-layer structure I use
After running probably four hundred prompts through Midjourney, ChatGPT's image generator and Google's Gemini image tools over the past year, for my own LinkedIn content and for client campaigns, I settled on a structure with five layers. I write them in this order every single time:
- Subject and pose: who or what, doing what, from which angle (three-quarter, straight on, over the shoulder)
- Wardrobe and setting: specific fabric, colour, and location detail, not "stylish outfit" but "cream linen blazer, gold hoop earrings, standing against a rain-streaked window"
- Lighting: golden hour, overcast soft light, single window light, neon backlight, whatever it is
- Camera and lens language: 35mm, 85mm portrait lens, shallow depth of field, shot on film, medium format
- Mood and colour grade: muted earth tones, high contrast black and white, warm film grain, desaturated blue
Here's a prompt built that way, for a personal brand photo I use as an example when clients ask what I mean: "Woman in her early fifties, three-quarter angle, laughing naturally, wearing a tailored charcoal wool coat over a cream silk shirt, standing in a London street at dusk, streetlights just switching on, shot on an 85mm lens with shallow depth of field, film grain, warm amber and cool blue colour grade, editorial magazine style." That prompt gives the model almost nothing to guess at. Compare it to "aesthetic professional photo of a confident businesswoman" and you'll see immediately why one produces something usable and the other produces a stock photo you have seen a thousand times.
If you want a deeper breakdown of building prompts specifically for photos of yourself, I go through the whole process, including how to handle your own face and body accurately, in my guide to aesthetic AI generated pictures of yourself.
A real example from a Christmas campaign that went sideways first
Last November I was building a set of festive images for a client's email campaign, cosy, warm, "Christmas market" energy. My first ten prompts all said some version of "aesthetic Christmas photo, warm and cosy, festive." Every single one came back looking like a generic greetings card, string lights blown out to white blobs, faces slightly waxy, snow that looked like static.
What fixed it was cutting every mood word and replacing it with things a set designer would write on a call sheet: "wooden market stall with strings of warm white fairy lights, mulled wine steam visible in cold air, wool mittens holding a ceramic mug, shallow depth of field, shot at dusk with the light just turning blue, 50mm lens, slight film grain." Same festive feeling, completely different, usable result. I now keep a running note of exactly which phrases worked for seasonal content, and I've written up the fuller version of that process in my post on writing festive photo prompts with ChatGPT for Christmas content, because it turns out the same principle applies to Easter, Halloween, any seasonal push you are planning for 2026.
The words that quietly wreck your prompts
This is the part most people writing about prompts skip, because it feels like it makes the tools sound worse than they're marketed to be, and I'll say it plainly anyway: adjectives like "beautiful," "stunning," "gorgeous" and "aesthetic" are actively making your images worse, not better. They eat up your word budget without giving the model anything to render. Every time you add one of those words instead of a specific detail, you are trading a instruction the model can use for a vibe it has to invent, and it invents the same over-smoothed, overly symmetrical, slightly uncanny face every single time because that is the statistical middle of "beautiful woman" across its whole training set.
Drop the adjectives. Replace every one with a fact. Instead of "beautiful lighting," write "single soft window light from the left, slight shadow under the chin." Instead of "stunning outfit," write the fabric, the colour, the cut. Your prompts will look less poetic and more like a props list. That is exactly the point. The model does better with a props list than with poetry.
Lighting and camera language that moves the needle
If I had to pick the two things that make the biggest visible difference between an amateur-looking AI image and one that reads as intentional, it is lighting description and lens language. A few phrases I reach for constantly, because they consistently produce a specific, recognisable look rather than a generic one:
- "golden hour, low sun, long shadows" for warmth without the plastic glow
- "overcast soft light, no harsh shadows" for a calmer, editorial feel
- "single window light, moody, slight underexposure" for something closer to a portrait studio
- "35mm lens, deep depth of field, everything in focus" when you want a documentary feel
- "85mm lens, shallow depth of field, blurred background" when you want the subject to pop
- "shot on 35mm film, visible grain, slightly faded colours" for a nostalgic, less digital look
Those six phrases, used and combined, cover about 80% of what people mean when they say "aesthetic." Try dropping just one of them into your next prompt and comparing the output side by side with your usual wording. The difference is bigger than people expect, and once you see it you stop trusting your own instinct to reach for "beautiful" again. For portrait work specifically, where the small details of angle and lens matter even more than in landscape or product shots, I've broken down the exact phrasing that works in my guide to writing portrait style prompts for ChatGPT image generation.
How many attempts it really takes
Nobody talks about the number, so here it is: on a typical client shoot, I generate somewhere between 30 and 50 images to get 4 or 5 I'd use. That is not a failure rate to be embarrassed about, that is the normal cost of working with a tool that is fundamentally a slot machine with very good taste in composition. What changes with a well-built prompt is not that you get it right first time, it's that your "miss" rate drops from something like 9 in 10 down to maybe 6 in 10, and the misses look closer to what you wanted, so you're iterating instead of starting over.
A practical process I use with every new prompt:
Want AI doing the heavy lifting in your marketing?
I build the systems that handle the boring 80 percent, so you get your week back. Done properly, with the human kept in.
- Write the five-layer prompt as above and generate a batch of 4
- Pick the closest one and note exactly which element is wrong, pose, lighting, colour, framing
- Change only that one element in your next prompt, keep everything else word for word
- Repeat for two or three rounds, not ten
- If nothing in three rounds is close, the subject or setting description is the problem, not the styling, so go back and rewrite that first
This single-variable approach is the difference between someone who "plays around" with prompts for an hour and gets nothing, and someone who gets a usable image in 15 minutes because they know exactly what to change each time.
Matching the prompt style to the tool you're using
Aesthetic prompts do not transfer word for word between tools, and this trips people up constantly. Midjourney rewards dense, comma-separated detail and responds well to explicit style references and aspect ratios. ChatGPT's built-in image generation, which most people now reach for because it's inside a tool they already use daily, responds better to plainer, more conversational sentences and is far stronger at following instructions about text, layout and specific edits than Midjourney is. Google's Gemini image tools sit somewhere in between and are noticeably better at photorealistic skin texture out of the box.
If you are choosing between these for ongoing business use, budget and volume matter as much as prompt style, since some tools charge per image and others bundle generation into a subscription you're already paying for. I've laid out the practical differences, including where each one falls down, in my comparison of AI image generators for business use. And if you're specifically using ChatGPT and keep hitting generation limits mid-project, it's worth knowing the daily caps before you plan a whole campaign around it, which I cover in my broader ChatGPT prompts and limits guide.
Quick reference: aesthetics people ask for most, translated
A shortlist of the aesthetic categories I get asked about most often, translated into the concrete language that produces them:
- Cottagecore: linen and cotton fabrics, muted sage and cream palette, natural window light, wildflowers, wooden furniture, slightly overexposed highlights
- Dark academia: tweed and wool, deep burgundy and forest green, low warm lamp light, old books, dust particles visible in light beam, 50mm lens
- Editorial fashion: bold single colour background, hard directional studio light, sharp shadows, 85mm lens, high contrast, minimal styling
- Vintage film: shot on 35mm film, visible grain, slightly faded warm colours, soft focus corners, 1970s colour palette
- Cinematic: wide aspect ratio, teal and orange colour grade, dramatic side lighting, slight lens flare, anamorphic feel
That last one, cinematic, is popular enough with clients wanting punchier social content that I've written a full walk-through of the exact wording and edits that produce it, in my post on using ChatGPT prompts for cinematic style photo edits.
The one habit that changes everything
If you take one thing away from all of this, let it be a habit rather than a prompt template: keep a running document of exactly what worked, word for word, and what didn't. I keep mine in a plain notes file, no fancy system, just the prompt and a one-line note on the result. Six months in, that document is worth more than any prompt guide you'll read online, mine included, because it's built from your specific taste and your specific tools, not someone else's Pinterest board.
Frequently asked questions
Do I need to use the word "aesthetic" in my prompt at all?
No, and you're usually better off leaving it out entirely. The word tells the model nothing concrete, so replace it with the specific wardrobe, lighting, lens and colour detail you mean, and you'll get a sharper, less generic result.
How long should an aesthetic photo prompt be?
Aim for one to three sentences covering subject and pose, wardrobe and setting, lighting, camera or lens language, and mood or colour grade. That's usually 40 to 80 words, dense enough to give the model real instructions without burying the important details.
Why do my AI photos all look the same even when I change the prompt?
This usually means you're only changing the adjectives, "beautiful" to "stunning," rather than the concrete details. Change one specific, filmable element at a time, lighting, lens, wardrobe fabric, and you'll see far more variation between generations.
Which AI tool gives the most aesthetic photos for someone starting out?
Midjourney tends to produce the most consistently stylish output for pure aesthetic work, while ChatGPT's image tool is easier to iterate with conversationally and better for following exact instructions. Try the same well-built prompt in both before committing to one for ongoing work.
Where to check the details
Related reading: ChatGPT Prompts for Halloween Costume Ideas: 30 Copy and Paste Prompts (.