Why Your AI Images Look Fake (and How to Fix It)
You can usually tell in under a second. Before you've consciously noticed anything specific, some part of your brain has already filed the image under "a computer made this." Only afterwards do you go looking for reasons. The skin is too smooth. The eyes are doing something odd. Everything is lit like a car showroom.
I run an image prompting tool, so I spend an unreasonable amount of time staring at generated images and asking why one reads as a photo while the next one, made with the same model, reads as a cutscene from a video game. The good news: the fake look isn't random, and these days it isn't really a model problem either. Current models can produce images that pass for photography. When they don't, it's almost always because the prompt pushed them somewhere else.
Here are the tells I see most often, where they come from, and what actually fixes them.
Plastic skin is the biggest tell
Waxy, poreless, evenly toned, like the person was dipped in resin. Two things cause it.
Diffusion models are trained to remove noise, and when they're unsure, they hedge toward smoothness, because smooth is the statistical average of every face they've ever seen. And the faces they've seen skew heavily toward retouched ones: editorial shoots, beauty-filtered selfies, stock photography. The training data was airbrushed before the model ever got to it.
Then we make it worse. "Flawless skin, perfect complexion, beautiful face" — I see prompts like this constantly, and every one of those words is an instruction to airbrush harder.
The fix is to ask for what you actually want, which is texture. "Visible pores", "fine lines", "unretouched skin" all work. So does camera language, for a less obvious reason: a phrase like "shot on a 50mm lens, natural window light" pulls the model toward the photojournalism and documentary photos in its training data, and those were far less retouched than the ad campaigns. You're not really describing hardware. You're choosing which pile of reference photos the model reaches into.
Dead eyes
Real eyes have catchlights, the small bright reflections of whatever is lighting the scene. When a generated face looks lifeless, the catchlights are usually missing, or worse, present but different in each eye, which your brain notices even when you can't say why.
The other half of the problem is the stare. Models love a perfectly centered, perfectly symmetrical gaze straight into the lens, and almost no real photograph looks like that.
Give the eyes something to do. "Looking just past the camera" or "glancing toward the window" instantly reads more candid than a head-on stare. And if your prompt names a light source (more on that next), the catchlights tend to sort themselves out, because the model now knows what the eyes should be reflecting.
Everything is lit like a showroom
No shadows anywhere. Every surface evenly, generously illuminated, with that faint HDR glow where the shadows have been lifted and the sky never clips. Real photographs are mostly shadow. A room lit by one window is dark in the corners, and your eye expects that darkness even if you've never thought about it.
The cause is partly training data (professional photos are well lit, and heavily processed HDR shots were wildly popular online for years) and partly vague prompting. If you say nothing about light, you get the average of all light, which is everything lit from everywhere.
So name one light source and let the rest fall away: "lit by a single window on the left, the far side of the room in shadow." Words like "overcast", "dusk" and "lamplight" each carry a whole lighting setup with them. And don't be afraid of "underexposed" or "slight grain". A little darkness does more for realism than any resolution keyword ever will.
Delete the magic words
"8k, ultra-realistic, hyperdetailed, masterpiece, trending on artstation." At some point a prompt list from 2022 convinced everyone these are the seasoning you sprinkle on for realism, and they've been copy-pasted ever since.
Think about where those words actually appear in the training data. No photographer has ever captioned their work "ultra realistic photo". The images that carry captions like that are digital paintings and 3D renders, because that's the community where calling something realistic is a compliment. "Masterpiece" and "artstation" point the same direction. You're asking for realism in a vocabulary that lives almost exclusively next to concept art, so what you get back is very polished concept art.
A photograph gets described the way a photo editor would describe it: who, where, what light, what lens, what moment. That's the whole trick.
Nothing is ever worn, stained, or crooked
Generated interiors look like show homes. Streets are freshly washed, t-shirts have never been through a dryer, desks hold a laptop and nothing else. Reality accumulates mess, and its absence is quietly loud.
The model won't add clutter on its own, because clutter is specific and the average of all desks is a clean desk. So put the entropy in yourself: a coffee ring, a tangle of cables, a picture frame hanging slightly crooked, scuffed sneakers. Details like these do double duty. Each one is also evidence, to your viewer's brain, that a real moment happened here.
The same face, again
Generate ten portraits of a "beautiful woman" and you'll meet roughly the same person ten times: symmetrical, mid-twenties, features from everywhere and nowhere. It's the averaging problem wearing a different hat. "Beautiful" has no information in it, so the model falls back on the mean of every face it knows.
Specificity is what breaks the spell. An age. Deep-set eyes. A slightly crooked nose, sun around the eyes, a gap in the teeth. This is honestly the reason we built Image Prompt Maker around personas made of dozens of concrete traits instead of one pile of adjectives. "Beautiful" produces the average. "52 years old, deep-set eyes, weathered skin, a nose broken once long ago" produces a person.
Describe a photograph, not a wish
Every fix above is the same fix, applied to a different symptom. Prompts fail when they describe how you want to feel about the image ("stunning, gorgeous, perfect, masterpiece") and they work when they describe a photograph that could exist.
Compare:
beautiful woman, perfect face, ultra realistic, 8k, hyperdetailed, masterpiece
a woman in her late 30s at a kitchen table, morning light from a window on the right, unretouched skin with visible pores, looking just past the camera, 50mm lens, slight grain
Same model, same settings, and they might as well have come from different decades of technology. The second image won't be flawless. But nobody will scroll past it thinking "AI", which is the one thing "ultra realistic" was never going to buy you.
If you'd rather not hand-write forty concrete details every time, that's essentially what Image Prompt Maker automates: you build the person, the light and the lens once, and the builder writes prompts like the second one for every image after that. But the principle costs nothing and works in any tool. Wishes out, photographs in.