How to write an AI image prompt that actually works well
A working method for AI image prompts. Order the elements, name equipment and fabrics instead of adjectives, state each thing once, and split out negatives.
An AI image prompt works when it reads like a shot list rather than a wish. Name the subject first, then pose, wardrobe, setting, light, camera, and finish, in that order, and state each element exactly once. Most of the words people pile onto a failing prompt, the stunning, highly detailed, masterpiece, 8k kind, do nothing useful, because they describe a reaction instead of a thing the model can draw. Swap them for equipment, fabrics, and light you could actually set up, and the output stops wandering.
Why most prompts fail
A prompt is a set of constraints. The model fills every gap you leave, and it fills them with the average of everything it has seen. That average is what people mean by the generic look: symmetrical face, glossy skin, soft rim light from nowhere, shallow background blur, no weight in the clothing.
Three failure patterns cover most bad prompts.
- Adjective stacking. Twenty quality words in a row give the model no new information, and they crowd out the words that carry real instruction. Attention is finite. Every token spent on ultra realistic is a token not spent on the lens.
- Contradiction. People describe the same thing twice in different words. Soft diffused light near the start and dramatic hard shadows near the end is not a richer description, it is an argument. The model resolves it by picking one, and you do not get to choose which.
- Missing anchors. No location, no time of day, no lens, no surface. The model has nothing to hold on to, so the image drifts toward whatever that subject usually looks like in stock photography.
The fix is not a longer prompt. It is a prompt that is ordered, specific, and internally consistent.
The order that matters
Read the prompt aloud. It should move from the person outward, the way a photographer builds a frame.
1. Subject and identity
Open with who or what this is, and the identity facts that must survive: apparent age range, build, hair length and texture, skin tone in plain terms, facial hair, glasses. Do not bury this at the end. Whatever comes first tends to be treated as the anchor the rest of the image serves.
2. Pose
Say what the body is doing, not how it feels. Seated, weight on the left hip, torso turned three-quarters to camera, chin level, hands loose in the lap is instruction. Confident pose is not.
3. Wardrobe
Name the garment and the fabric. A jacket renders as an average jacket. A single-breasted charcoal wool jacket with notched lapels and visible pick stitching renders as a specific one. Fabric words carry weight information: linen creases, silk falls, denim holds its shape, wool has body.
4. Setting
One place, described with two or three concrete nouns. A narrow kitchen with open shelving, a scratched steel worktop, and a window on the left beats a beautiful modern kitchen. Give the model surfaces and it gives you reflections and shadows that agree with each other.
5. Light
State the source, the direction, and the quality. Single window, camera left, overcast midday, soft falloff, no fill is a lighting setup. Time of day helps, because it implies colour temperature without you naming numbers.
6. Camera
Focal length does more work than almost anything else in the prompt, because it sets perspective and therefore the shape of the face. A 35mm frames a room and slightly exaggerates whatever is closest. An 85mm flattens features and separates the subject. A 135mm compresses hard, which is why long-lens portraits look expensive. Add aperture if depth matters, and shooting distance if framing does.
7. Finish
The last clause is the medium and the grade: colour negative film, slight halation in the highlights, fine grain, or clean digital capture, neutral grade, no vignette. This is where texture lives. It is also where you decide whether skin keeps its pores.
Name equipment and fabrics, not adjectives
The highest-return edit you can make is replacing evaluative words with nameable things.
- Not dramatic lighting. Say one hard key from a bare bulb, high and camera right, black flag on the shadow side.
- Not luxurious dress. Say bias-cut silk crepe, floor length, matte surface, drape breaking at the ankle.
- Not professional photo. Say 85mm at f/2, one metre from the subject, eye level.
- Not sharp detail. Say visible skin texture, individual eyelashes, fabric weave readable at the shoulder.
Equipment and material names are dense. Each one implies a whole set of visual consequences the model already knows. Adjectives are thin, because thousands of unrelated images were captioned with the same word.
State each element once
Pick one word per idea and use it once. If you have said overcast, do not later add soft natural light, because now the model is weighing two descriptions of the same lighting and may render a compromise that is neither.
This matters most for the things a model is eager to change: age, hair, expression, and background. Say it once, clearly, in its proper position, and move on.
When an output is wrong, the instinct is to add a corrective clause at the end. Resist it. Edit the original clause instead. A prompt with a patch bolted to the end is a prompt arguing with itself.
Keep negatives in a separate block
Do not thread exclusions through the description. A woman in a red coat, no hat, standing on a bridge, not blurry, in winter, no other people has to be untangled before it can be read.
Put the positive description in one paragraph and the exclusions on their own line, in whatever form your tool supports. If the tool has a dedicated negative field, use it. If it does not, end with a single clause: exclude: hats, additional people, text, motion blur.
Two rules about negatives. Keep them short, because a long exclusion list drags attention toward the very things you are trying to remove. And never write a negative for something you can state positively. Bare head works better than no hat, and empty street works better than no crowds.
Aspect ratio is a setting, not prose
Writing vertical 9:16 format inside the sentence asks the model to read a framing instruction as scene content. Sometimes it obeys, sometimes it draws a phone. Use the parameter field your tool provides, or the flag it documents, and keep the prose free of it.
The same goes for model name, version, seed, and sampler. Those are settings. Every entry in the image prompt library keeps them as structured fields for exactly this reason, and the habit is worth copying when you write your own.
Worked example one: a portrait
The vague version, which is where most people start:
beautiful woman portrait, professional photography, stunning lighting,
highly detailed, 8k, masterpiece, cinematic, elegant dress,
amazing background, sharp focus, award winning
Nothing there is a constraint. Every word is a rating. The result will be a symmetrical face, plastic skin, an unplaceable blurred background, and a dress that is nothing in particular.
The rewritten version:
A woman in her early thirties, shoulder-length dark brown hair with a
loose wave, warm mid-brown skin, no makeup beyond a matte lip.
Seated on a wooden stool, weight on the left hip, torso turned
three-quarters to camera, head turned back to the lens, chin level,
hands resting one over the other.
Wearing a bias-cut ink-blue silk crepe dress, narrow straps, matte
surface, fabric breaking softly at the hip.
Setting: an empty room, lime-plastered wall, bare floorboards,
nothing else in frame.
Light: one large window camera left, overcast midday, soft falloff
across the face, no fill, shadow side allowed to go dark.
Camera: 85mm at f/2.8, eye level, one and a half metres from the
subject, framed from mid-thigh up.
Finish: colour negative film, fine grain, visible skin texture,
neutral grade.
Exclude: jewellery, text, additional people.
Same subject, same intent. The second prompt is longer, but not one word of it is decoration. Every clause removes a decision the model would otherwise make for you.
Worked example two: a product shot
The vague version:
amazing coffee cup product photo, professional studio, beautiful
lighting, ultra realistic, trending, high quality, 4k, perfect
The rewritten version:
A matte black stoneware coffee cup, unglazed rim, faint throwing
rings on the body, half full of black coffee.
Standing on a honed grey limestone slab, one water ring on the stone
beside it.
Light: a single softbox directly overhead and slightly behind, a
white bounce card just outside frame at camera left, specular
highlight running down the left side of the cup, shadow pooling
short and hard under the base.
Camera: 100mm macro at f/8, slightly above the rim line so the
coffee surface reads as an ellipse.
Finish: clean digital capture, neutral grade, no vignette, no added
contrast.
Exclude: steam, saucer, spoon, beans, text, props.
The exclusion block earns its place here. Coffee prompts attract steam and scattered beans the way portrait prompts attract heavy background blur, because that is what the training captions were full of. Naming them is the only reliable way to keep them out.
Before you generate
Read your prompt once against this list.
- Does it open with the subject and the identity facts that must survive?
- Does it move outward in order: pose, wardrobe, setting, light, camera, finish?
- Is every fabric and every piece of equipment named, with no evaluative adjectives left?
- Is each element stated exactly once, with no late patches?
- Are the exclusions in their own block, kept short, and phrased positively where possible?
- Is the aspect ratio in the settings rather than in the sentence?
If the answer to all six is yes and the image is still wrong, the problem is usually one clause carrying two instructions. Split it and try again.
This structure carries over to motion work almost unchanged, with the addition of what moves and how the camera travels. You can see it applied to clips in the video prompt library.