KitanaAI Photo & VideoStudio
photo maker · 5 min read

Why AI Avatars Don't Look Like You, and How Reference Photos Fix It

The single-photo problem behind most likeness failures, what extra reference photos give the model, and how to choose three that actually help.

By the Kitana team
WAPHOTO MAKER

The most common complaint about any AI portrait tool is the same six words: it does not look like me. It is not because the model is bad at faces. It is because a single photograph is a surprisingly poor description of one.

What one photo actually tells the model

A photo is a face plus everything that happened to it at that moment: the direction of the light, the lens, the angle, the expression, the compression, the phone's beauty processing. The model sees all of that as one bundle and has no way of knowing which parts are you and which parts are the Tuesday afternoon you took it.

When it then restyles you as a painting or a 3D character, it has to decide what to keep. It keeps the strongest signals, and in a single photo the strongest signals are often not your identity. A shadow under the cheekbone reads as a hollow cheek. A slightly wide lens reads as a broader nose. A tilted head reads as an asymmetry. Each is faithfully carried into the avatar, and the result is a stranger with your hair.

This is also why two runs from the same photo differ so much. The model is sampling from many plausible answers, and with weak evidence about which features are essential, the answers spread widely. The mechanics of that are covered in how to turn a selfie into anime art, which is the style where it shows most.

What extra references give it

Add a second photo from a different day, and something changes. The light is different, the angle is different, the expression is different, and yet some things are identical: eye spacing, nose shape, the line of the mouth, the shape of the jaw, the hairline. The model can now separate what persists from what varies, and it treats what persists as you.

Add a third, and the separation gets sharper again. Past three, the gains flatten, which is why the tools stop there.

The practical effect is that the restyle changes the surface and keeps the structure. Anime flattens the shading but the eye spacing stays. Pixar rounds the features but the nose is still your nose. Yearbook changes the hair entirely, and you are still recognisable underneath it.

In Kitana this is called identity lock, and it is available on the five tools where a face is re-rendered: AI Avatar, AI Headshot, 90s Yearbook, Age Transform and Cartoon Me. Each run stays a single creation regardless of how many references you attach.

Choosing three photos that help

The word is variety, within limits.

Vary the conditions. Different days, different light, different angles. A frontal shot by a window, a three-quarter view outdoors, a candid from a friend's phone. The model learns from what stays constant, so it needs things to change around it.

Keep the person constant. Same person, obviously, but also the same version of them: same era, same hairstyle, no drastic weight change, no beard in one and clean-shaven in the next. A reference that describes a different version of you pulls the result toward the average of the two.

Keep every photo usable on its own. Sharp, reasonably lit, face large enough in the frame, no heavy filter, no sunglasses, no hand over the mouth. A blurred or dark reference is not neutral; the model reads noise as features. If a photo would be a poor main image, it is a poor reference.

Do not add the same selfie three times. It is the most common mistake, and it is understandable: those are the photos on your phone. Three frames from one session share the same light, lens and expression, so they tell the model nothing new.

A good source set for the main image is the same as a good reference set, and the full checklist is in preparing photos for the best AI avatar results.

When references do not help

They do not fix a bad main photo. If the main image is blurred, badly lit or filtered, the references improve the model's idea of you but the run still starts from a weak source. Fix the main image first.

They do not help tools that keep your pixels. Background replacement, object removal, upscaling and style transfer applied to a scene are not re-rendering your face, so there is nothing for a reference to anchor. The tools that accept references are exactly the ones that need them.

They do not overcome a heavy style on their own. Some styles simply discard more information than others. A Pixar render keeps more of you than a flat comic-book style, whatever you feed it. The right expectation is "recognisably me in this style", not "a photograph of me painted".

A test that tells you whether it worked

Show the avatar to someone who knows you, without telling them who it is meant to be. If they name you, identity lock did its job. If they say "is that your brother?", the structure drifted, and the fix is usually one of three: a better main photo, a reference that is not really you, or simply running it again.

It is worth being deliberate about this because a recognisable avatar is reusable. Generate it once from a good source set and it can be cropped for every platform, which is the whole method in creating consistent AI avatars across platforms. An avatar that only sort of looks like you is one you will end up regenerating for every new profile.

Identity lock is built into five of the fourteen tools in the Kitana apps and adds nothing to the cost of a run.

Frequently asked questions

How many extra photos can I add?
Up to three, on the avatar, headshot, yearbook, age and cartoon tools. The other tools either do not restyle a face or do not need it: background replacement, object removal and upscaling keep your face pixels as they are.
Do the extra photos need to be recent?
They need to look like you now. A photo from ten years ago with different hair and a different face shape teaches the model a different person, and the result averages the two. Same person, same era, different days.
Should the reference photos be the same pose as the main one?
No, and it is better if they are not. The value of a reference is that it shows the same features under different conditions. Three identical selfies add almost nothing over one; a frontal, a three-quarter and a candid laugh add a great deal.
Does identity lock cost more?
No. A run with reference photos is still one creation. The extra photos are uploaded as inputs to the same job.
Why did the result get worse after I added references?
Usually because one of the references was a different person, was heavily filtered, or was so dark or blurred that the model treated the noise as a feature. Remove the weakest photo and run again. A bad reference is worse than none.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Try identity lock

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.