Kashi Keefe← All resources
Creative·ChatGPT·12 min read

The Ultimate Image-Prompting Guide

Everything you need to write image prompts that come out right the first time.

Most people type a vague idea into ChatGPT and hope for the best. They get something generic, or worse, something that almost works but has a weird hand or a floating eye. The problem is not the AI — it is the prompt. Image generation lives and dies on specificity. This guide gives you the exact framework, vocabulary, and prompt structure to get results that look intentional from the first try.

ChatGPT's image generation runs on DALL-E. It reads your prompt literally, so every word either earns its place or costs you. The good news: once you learn the layers of an effective prompt, you can apply them to any image — product photos, social content, lead magnets, slide visuals, pitch deck graphics, anything.

The Anatomy of an Image Prompt

Every strong image prompt has five layers. You do not need all five every time, but knowing them lets you diagnose exactly why an image missed. Think of it as a recipe: leave out an ingredient and you will notice.

  • Subject — What is in the image? Be exact. Not 'a woman' but 'a woman in her late 30s, dark curly hair, wearing a tailored charcoal blazer, looking directly at camera.'
  • Setting — Where is it happening? Indoor or outdoor, time of day, background details. 'In a minimalist home office with a single large window, morning light coming from the left.'
  • Style — What does it look like aesthetically? Photography, illustration, flat design, oil painting, cinematic still, editorial magazine photo. This one word changes everything.
  • Lighting — Soft natural light, harsh studio light, golden hour, backlit, rim lighting. Lighting is the difference between a snapshot and a visual that stops a scroll.
  • Technical parameters — Aspect ratio, camera angle, lens type (e.g. wide angle, 85mm portrait lens), depth of field. These give you control that most people never think to use.

The Prompt

Use this structure as your starting template. Fill in each bracketed section with your specifics. The example below is set up for a professional personal-brand photo — swap the details for whatever you are making.

Create a [style] image of [detailed subject description]. The setting is [specific location with environmental details]. Lighting is [lighting description — quality, direction, color temperature]. Shot with a [lens/camera description] producing [depth of field description]. The mood is [emotional tone]. The color palette leans [2–3 colors or a palette descriptor]. No text, no watermarks.

---

Example filled in:

Create a cinematic editorial photo of a woman in her late 30s with dark curly hair, wearing a tailored charcoal blazer and white linen shirt, looking directly at the camera with a calm, confident expression. The setting is a minimalist home office with exposed concrete walls, a large window behind her, and a blurred bookshelf to the right. Lighting is soft natural window light coming from the left side, creating gentle shadows across her face — warm 5500K tone. Shot with an 85mm lens at f/1.8, shallow depth of field with the background softly blurred. The mood is professional but approachable. Color palette leans charcoal, warm white, and muted sage. No text, no watermarks.

Tips That Change Your Results

  • Lead with style, not subject. ChatGPT's model weighs early words more heavily. If you write 'a woman...' first, it makes compositional decisions before it knows what you want aesthetically. Start with 'Cinematic photo of...' or 'Flat vector illustration of...' instead.
  • Name the camera or lens. Writing '85mm portrait lens' or 'wide-angle 24mm' signals photographic realism and gives you the compression or distortion effect that matches real photography.
  • Describe what you do NOT want — specifically. 'No text' is basic. Go further: 'No cartoonish features, no oversaturated colors, no generic stock-photo smile.' Negative constraints are underused and extremely effective.
  • Lock the color palette. Use hex-adjacent language or palette names: 'muted earth tones — terracotta, warm beige, and forest green.' This stops the AI from defaulting to oversaturated, generic color choices.
  • Use reference art styles carefully. Saying 'in the style of a 1970s National Geographic photo' or 'like a Wes Anderson film still' gives the AI a rich aesthetic shortcut. Avoid naming living artists — it gets inconsistent.
  • Separate the composition from the detail. If you want something complex, describe the foreground subject first, then midground, then background. This helps the model layer the scene instead of flattening it.
  • Iterate with one change at a time. If the image is 90% right, do not rewrite the whole prompt. Add one clause. 'Same image but the lighting is now cooler and blue-toned, like dusk' is more surgical than starting over.
  • Specify aspect ratio at the end. Add '16:9 landscape orientation' for slide decks and covers, '4:5 vertical' for Instagram, '1:1 square' for profile images. ChatGPT respects this when you state it clearly.
  • Ask for variations explicitly. End your prompt with 'Generate four variations with slightly different expressions and lighting angles.' You get options without having to reprompt from scratch.

Prompt Variations by Use Case

The core structure stays the same. What changes is the style layer and the parameters you emphasize. Here are four ready-to-use variations you can adapt immediately.

Variation 1: Social Media Content Graphic

Flat design illustration of a smartphone displaying a simple analytics dashboard with rising line graphs. Clean, minimal aesthetic with bold geometric shapes. Background is deep navy blue. Accent colors: electric teal and warm white. No gradients. No realistic textures. No text or labels on the screen. 1:1 square format.

Variation 2: Product Photo on Clean Background

Commercial product photography of a matte black ceramic coffee mug on a white marble surface. Studio lighting setup: soft box from the upper left, subtle fill light on the right, clean shadow beneath the mug. Shallow depth of field, 100mm macro lens. The mug is centered, slightly angled 30 degrees to the right. Background is pure white. No props, no text, no lifestyle elements. Hyper-realistic, high resolution.

Variation 3: Concept Illustration for a Slide Deck

Editorial vector illustration representing 'AI-powered business growth.' An abstract human figure (simplified, no face) interacting with a glowing neural network pattern that expands outward into upward-trending arrows and geometric shapes. Color palette: deep purple, electric blue, and bright white on a dark charcoal background. Style: modern tech illustration, similar to a high-end SaaS marketing asset. No text, no logos, no literal robot imagery. 16:9 landscape format.

Variation 4: Event or Course Cover Image

Cinematic wide-angle photograph of an empty modern conference room at night, floor-to-ceiling glass windows overlooking a lit city skyline. The room has a long dark wood table, minimalist pendant lights glowing warm amber, and a single projector screen at the far end emitting soft blue light. Mood: ambitious, focused, cinematic. Color palette: deep navy, warm amber, and cool blue highlights. No people. No text. 16:9 landscape.

Common Mistakes and How to Fix Them

  • Problem: The image looks generic. Fix: You did not define style or lighting. Add 'cinematic editorial photo' or 'flat design illustration' at the start and specify the light source.
  • Problem: Colors are too saturated and cartoonish. Fix: Add 'muted color palette' and explicitly name 2–3 desaturated tones. Also add 'no oversaturation, no neon colors.'
  • Problem: Hands or faces look wrong. Fix: Minimize complex hand poses in your prompt. If you need hands, say 'hands resting naturally on a desk, not holding anything.' For faces, specify 'realistic proportions, natural skin texture, neutral expression.'
  • Problem: The composition is cluttered. Fix: Add 'negative space in the upper third' or 'simple, uncluttered composition with one clear focal point.' Explicitly cut things: 'no props, no background clutter.'
  • Problem: The AI ignores part of your prompt. Fix: Shorten the prompt. ChatGPT's image model loses track of constraints when prompts are very long. Cut modifiers that are not load-bearing and keep the five core layers clean.
  • Problem: You keep getting the same style even when you change the prompt. Fix: Add a hard style reset line: 'This is NOT a stock photo. This is NOT a cartoon. Style: [your specific style].' The explicit negation breaks the default behavior.

Building a Prompt Library

Once you find a prompt structure that works, save it. Build a simple document — a Notion page, a Google Doc, anything — organized by image type: social graphics, product photos, people shots, abstract concepts, slide visuals. Every time you get a result you like, paste the prompt and a screenshot next to it. Within a few weeks you have a personal library that lets you spin up on-brand visuals in minutes, not hours.

You can also use ChatGPT to help you build prompts. Describe what you want in plain language and ask it to 'write a detailed image generation prompt using the five-layer structure: subject, setting, style, lighting, and technical parameters.' It will often catch specifics you would have forgotten. Then paste that prompt back into image generation.

One pattern worth locking in early: always end your prompts with 'No text, no watermarks, no logos' unless you specifically want those elements. ChatGPT will sometimes add placeholder text or UI chrome by default. This single line eliminates 80% of those surprises.

The Mindset Shift That Makes This Click

Stop thinking of prompting as asking. Think of it as directing. You are the creative director. The AI is a capable production team that will execute exactly what you describe — and nothing more. If your brief is vague, the output is generic. If your brief is specific, the output is specific. The tool does not guess your intent. It reads your words.

That shift — from 'let me see what happens' to 'here is exactly what I want' — is what separates people who get consistent, usable images from people who regenerate 30 times and still settle. Use the five-layer structure, specify everything, and iterate surgically. You will start hitting on the first or second try more often than not.