The structure that works
- Subject: who or what, doing what, where. a red fox sitting on a mossy rock in a pine forest
- Style or medium: oil painting, 35mm photograph, watercolor, 3D render, or an artist: by Ivan Shishkin.
- Lighting and mood: golden hour, soft diffused light, dramatic rim light, foggy.
- Framing and lens: close-up, wide shot, 85mm, shallow depth of field.
- Quality words: highly detailed, sharp focus, masterpiece. Two or three are enough.
Words at the start of the prompt weigh more than words at the end. Keep the subject first.
An annotated example
portrait of an old fisherman, weathered face, wool sweater, oil painting by Rembrandt, warm candlelight, dark background, close-up, 85mm, highly detailed
The first five words fix the subject. The artist token sets palette and brushwork (see how much one name changes on the artist styles pages). The light words decide the atmosphere, the framing words the composition, and the last group polishes. Negative prompt: blurry, deformed hands, extra fingers, watermark, text — more in what is a negative prompt.
Weights and emphasis
In Automatic1111 and most interfaces, parentheses raise the weight of a term: (red scarf:1.3) means 30 % more important, (background:0.7) less. Square brackets lower it. Stay between 0.6 and 1.4: beyond that the image breaks. Use weights to rescue a detail the model keeps ignoring, not on every word.
Common mistakes
- Full sentences and negations: "a room without windows" often produces windows. Put unwanted things in the negative prompt.
- Too many ideas: one subject per image; 75 tokens (about 60 words) is the limit of one CLIP chunk for SD 1.5 and SDXL, and what comes after is weaker.
- Contradictory styles: photorealistic watercolor gives a muddy result.
- Changing everything at once: fix the seed and change one word at a time to learn what each does.
Differences between models
SD 1.5 and SDXL respond to keyword lists like the example above. Stable Diffusion 3 and 3.5 understand natural sentences better and can place several subjects ("a cat on the left, a dog on the right") and render short text in quotes. Whatever the model, the 69 documented prompts on this site show the full parameters behind each image, and the prompt generator expands a few words into a complete prompt.