wowprompt

Using a reference photo without losing the face in AI art

Why AI models drift a face across generations, the clauses that hold identity, what makes a good source photo, and what to do when the output still drifts.

A reference photo holds the face when three things line up: the source image is clean and frontal enough to be read, the prompt names the identity features you want preserved instead of assuming the photo covers them, and nothing else in the prompt asks the model to restyle the person. Drift is almost never random. It is usually the prompt competing with the reference, and the prompt winning.

What the model is actually doing with your photo

An image-to-image or reference-guided generation does not paste your face into a new scene. It builds a new image from scratch and uses the reference as a pull toward certain features. How strong that pull is depends on a weight setting, on how much of the frame the face occupies in your source, and on how much the text prompt is asking to change.

That last point is the one people miss. Identity preservation and scene transformation trade off against each other. A prompt that asks for a new outfit, a new location, new lighting, and a new camera angle all at once is asking the model to rebuild most of the pixels. The face is rebuilt along with everything else, and you get someone who looks like a sibling.

Why faces drift

Five causes account for most of it.

The face is too small in the source. If the head occupies a small part of a wide frame, there is not much identity information to work with. The model fills the rest from its average of human faces, which is younger, more symmetrical, and smoother than any real person.

The angle is too far from the target. A source shot in hard three-quarter profile, generated into a frontal portrait, forces the model to invent the half of the face it never saw. Invented halves are generic halves.

The prompt restates the face. Writing beautiful woman with perfect features alongside a reference photo tells the model to move toward the average, in direct opposition to the photo. Any evaluative face description in the prompt is a vote against the reference.

Style words override structure. Cinematic, editorial, glamour, retouched, and airbrushed all carry face-shaping consequences. Retouching is what removes the asymmetries that make a face recognisable.

Reference weight is too low. Many tools default to a middle setting that favours prompt adherence. If your priority is the face, the setting needs to move, and you need to accept less freedom in the scene.

The clauses that hold identity

Treat the reference as the source for structure, and use the prompt to name the specific, non-evaluative features that must survive. These are the clauses that actually do work.

  • Preserve the facial structure from the reference image. Simple, and worth stating explicitly in tools that accept plain instruction.
  • Name the asymmetries. Slightly higher left eyebrow, a small mole below the right cheekbone, a scar through the left eyebrow. Asymmetry is identity. A model told to keep one will keep several nearby.
  • Fix age in a range. Mid-forties, visible nasolabial lines, crow's feet at the outer corners. Without this, output skews five to fifteen years younger, every time.
  • Fix the hairline and hair texture. Receding at the temples, tight coils cut close, centre part with a widow's peak. Hair is one of the first things a model reassigns.
  • Fix skin. Visible pores, uneven tone across the cheeks, no retouching. This single clause stops most of the plastic look.
  • Fix distinguishing marks. Freckle patterns, birthmarks, gaps between teeth, a broken nose that healed with a bend.

And the clauses to remove entirely: any beauty adjective, any reference to a celebrity or a style of face, and anything that implies a makeover. Glow up, flattering, and idealised are all instructions to change the person.

One more rule worth following: do not describe the eyes, nose, and mouth in prose when you have a reference. Your description will be less accurate than the photo, and the two will fight.

What a good source photo looks like

The gap between a usable reference and a bad one is larger than the gap between models.

Lighting. Even, soft, frontal or slightly off-axis. Overcast daylight or a large window is close to ideal. Avoid hard side light that hides half the face in shadow, avoid direct on-camera flash, which flattens structure and blows out skin, and avoid coloured light, which the model may read as skin tone.

Angle. Frontal or near-frontal, within about thirty degrees. The chin should be roughly level. Upward angles distort the nose and jaw, and downward angles shrink the jaw and enlarge the forehead. The model will learn those distortions as facts about the face.

Resolution and framing. The head should fill a meaningful part of the frame, roughly collarbone to just above the hair. A large image of a tiny face is not a high-resolution reference, and sharp focus on the eyes matters more than pixel count.

Single subject. One person, clearly foremost. A second face anywhere in the frame is a risk, and cropping is the easy fix.

Expression. Neutral to slightly relaxed, mouth closed or barely open. Extreme expressions bake in muscle positions that then appear in every generation.

No heavy processing. Skip anything with a beauty filter, skin smoothing, a face-slimming edit, or a strong colour grade. Those already moved the face toward an average, so you are asking the model to preserve a face that is not quite the real one. Light colour correction is fine.

Clothing and background. Plain is better. A busy pattern near the jaw can bleed into the generated collar, and a strong background colour can shift the skin tone in the result.

If you have a choice of several photos, pick the boring one. A plain, well-lit, straight-on photo with a dull expression outperforms a good-looking photo taken at an angle in mixed light.

When the output still drifts

Change one thing at a time, in this order.

  1. Raise the reference weight and generate again with nothing else changed. If identity improves but the scene collapses toward the source photo, you have found the range, and the answer sits just below where you are.
  2. Cut the prompt back. Remove every word that is not the scene. Generate. If the face returns, add clauses back in groups and find the one that was doing the damage. It is usually a style word.
  3. Reduce how much you are asking to change. Keep the outfit or keep the lighting, rather than replacing both. Large transformations and tight identity are not available at the same time.
  4. Match the head angle in the prompt to the head angle in the source. If the reference is frontal, do not ask for a profile.
  5. Swap the source photo. If you have a closer, flatter-lit frame, this is often faster than any amount of prompt work.
  6. Generate a batch and select. Seed variance is real in identity work. Six outputs from one good prompt usually contain one that holds the face better than the rest.
  7. Go in two passes. Generate the scene with the reference at a high weight and a minimal prompt, then run a second pass on the result for wardrobe and grade. Splitting the transformation keeps each step small enough that the face survives.

If none of that works, look for a contradiction between the photo and the request. A source photo lit from one side and a prompt asking for flat frontal light means the model must relight the face, and relighting is rebuilding.

Say so on the page

If you publish or share a prompt that depends on an uploaded photo, state that above the copy action. A reference-dependent prompt run without a reference produces a generic face and a frustrated user, and there is no way to tell from the prompt text alone. Prompts in the image library that need an upload are flagged for that reason, and prompts built around traditional garments are collected under traditional wear, where identity and fabric work usually have to hold together at the same time.

The short version: give the model a clean, flat, frontal photo, describe only the features it cannot infer, delete every beauty word, and change one thing per generation until the face holds.

reference photoidentityimage promptsportraits

Read next