KitanaAI Photo & Video
Photo Maker · 7 min read

How to Write Text-to-Image Prompts: A Practical Framework

A five-part framework for text-to-image prompts, subject, setting, style, light and framing, with before-and-after examples and fixes for common failures.

By the Kitana team
HTPHOTO MAKER

A good text-to-image prompt says, in plain words, what the subject is, where it is, what style it should look like, how it is lit and how it is framed. Those five parts cover most of what a model needs. Write them in one to three sentences, put the most important part first, and change one part at a time when you revise. The rest of this guide shows the framework with examples and explains how to fix the common failures.

The five parts

Part Question it answers Example
Subject What is the picture of? "a red bicycle leaning against a wall"
Setting Where and when? "on a narrow cobbled street in the evening"
Style What kind of image? "35mm film photograph" or "flat vector illustration"
Light How is it lit? "warm low sunlight from the left, long shadows"
Framing How is it composed? "wide shot, bicycle in the lower third"

Put together: A red bicycle leaning against a wall on a narrow cobbled street in the evening, 35mm film photograph, warm low sunlight from the left with long shadows, wide shot with the bicycle in the lower third.

That is a complete prompt. It is not long, and every part changes the result.

Subject: be concrete

Vague nouns produce average pictures. "A dog" gives you the most typical dog. "A scruffy grey terrier sitting on a doorstep" gives you a particular one.

  • Name the thing precisely: "a ceramic espresso cup", not "a cup".
  • Give one or two identifying details: colour, material, age, condition.
  • Say what it is doing, if anything.
  • Limit the count. Three named objects is fine; eight is where models start dropping things.

Setting and style

Setting is location and time: indoors or out, season, time of day, weather. It shapes colour and mood more than people expect.

Style tells the model what kind of image to make. Useful, plain options:

  • Photographic: "product photograph", "documentary photo", "35mm film photograph", "studio portrait"
  • Illustrated: "watercolour illustration", "flat vector illustration", "pencil sketch", "children's book illustration"
  • Rendered: "3D render", "clay model", "isometric illustration"

Pick one. "Photorealistic watercolour" asks for two things at once. If you want to restyle a photo you already have, style transfer is the better tool; see the best AI style transfer apps compared.

Light and framing

Light is the most under-used part of most prompts. "Soft overcast daylight", "warm window light", "hard midday sun", "neon signs at night" and "candlelight" each change everything.

Framing controls composition: "close-up", "wide shot", "overhead view", "eye level", "subject centred", "lots of empty space on the right for text". If the image is for a thumbnail or a post, say so in the framing, and see AI social content sizes for the shapes.

Before and after

Weak: Beautiful cosy coffee shop, amazing, high quality, 4k.

Better: A small corner café with wooden tables and a large window, rainy afternoon, documentary photograph, soft grey daylight and warm lamps inside, wide shot from the doorway.

Weak: Logo for my bakery.

Better: A simple flat vector illustration of a croissant on a plain cream background, two colours, brown and cream, centred with plenty of space around it.

The better versions replace quality words with information.

Fixing common failures

  1. Something is missing. Move it to the start of the prompt and remove competing objects.
  2. Wrong style. Name one style clearly and remove style words that conflict.
  3. Too busy. Add "plain background" or "minimal composition", and cut objects.
  4. Text is garbled. Keep text short, put it in quotes, or add it later in a design tool.
  5. Hands or small details look wrong. Frame them smaller, or choose a composition that does not feature them.
  6. Every result looks the same. Change the light or the framing, not the subject.

Generation is sampled, so the same prompt gives different results. Run a good prompt twice before rewriting it. A failed generation is refunded.

What not to prompt

The Acceptable Use Policy bans NSFW content, non-consensual likeness, deepfakes and face swaps. Every prompt is screened by a content moderation check before generation. Do not describe real people by name, and avoid other brands' logos and trademarks. For an image of yourself, use a photo-based tool like the avatar tool instead of describing your face.

A template to copy

[Subject with one or two details], [setting and time], [one style], [light], [framing].

Save the prompts that work in a note. Most people end up with five or six reusable templates.

A worked example: a bakery's launch post

A baker wants an image for a new sourdough loaf announcement on Instagram. Here is the actual revision path.

  1. First attempt: Sourdough bread, delicious, professional. The result is a generic loaf on a generic counter.
  2. Add the subject details: A round sourdough loaf with a deep golden crust and a single curved score, dusted with flour. The loaf now looks like the baker's.
  3. Add setting and light: ...on a linen cloth on a wooden table, early morning, soft window light from the right. The mood improves.
  4. Add style and framing: ...product photograph, overhead view, loaf slightly off-centre with space at the top for text. Now it fits a post with a caption overlay.
  5. Keep text out. The shop name is added later in a design tool instead of in the prompt.

The baker then uses the social tool to prepare Instagram Post and Story-Reel versions, and saves the prompt as a template for future loaves, changing only the bread description each time.

A revision checklist

  1. Did the subject come out right? If not, fix the subject only.
  2. Is the style right? If not, change the style phrase only.
  3. Is the mood right? Change the light.
  4. Is the composition right? Change the framing.
  5. Run the same prompt twice before judging it.
  6. Save prompts that worked, with a note about what they are for.

Common mistakes

  • Stacking quality words. "Stunning, masterpiece, 8k" adds nothing the model can act on.
  • Contradicting yourself. "Minimal, busy street market" pulls in two directions.
  • Asking for long text. Keep words in images to a few, or add them afterwards.
  • Rewriting everything after one bad run. You lose what was working.
  • Describing a real person. It is against the Acceptable Use Policy and blocked by moderation. Use a photo-based tool for yourself.
  • Forgetting the format. A landscape image makes a poor Story. Say the framing.

Prompt patterns for common jobs

A few templates that cover most everyday requests:

  • Product shot: [product with material and colour] on [surface], [plain or styled background], product photograph, soft window light, [framing].
  • Blog header: [scene that illustrates the topic], flat vector illustration, limited palette of [two or three colours], wide shot with empty space on the left.
  • Poster background: [setting and mood], [style], dramatic light, lots of empty space at the top for a title.
  • Mood image: [place and time of day], 35mm film photograph, [weather], [light].

Fill the brackets, keep the rest. Templates make results more consistent across a series, which matters more than any single perfect image.

When a template stops working for a new subject, go back to the five parts and check which one the new subject needs changed.

Try it

Text-to-image runs in the browser studio and in the app. On the web you can make one image without an account and three more after signing up, which is enough to try the framework on a real idea. For background on how the models work, read what text-to-image AI is and how it works.

Frequently asked questions

How long should a text-to-image prompt be?
Usually one to three sentences. Long enough to cover subject, setting, style, light and framing; short enough that every word is doing something. Very long prompts tend to have parts that contradict each other.
Do I need special keywords or magic phrases?
No. Plain, specific description works better than stacks of quality words. 'A ceramic mug on a wooden table by a window, soft morning light, close-up' beats a list of adjectives.
Why does the model ignore part of my prompt?
Usually because two parts compete, the prompt is long, or the detail is small. Move the important part to the start, remove anything that contradicts it, and try again.
Can I make images of real people or brands?
Do not use text-to-image to make images of real people. The Acceptable Use Policy bans non-consensual likeness and deepfakes, and every prompt is screened by a content moderation check before generation. Avoid other companies' logos and trademarks too.
Where can I use text-to-image?
It is one of the four tools in the browser studio at photomaker.lol, and it is also in the Kitana app. On the web you can make one image without signing up and three more after sign-up.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Open the browser studio

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.