Cinematic AI Image Prompts: Depth, Lighting And Camera Angles
Depth in three layers, light with a source, a camera with a position and a lens. The cinematographer's questions I answer before I write any prompt, and the dark industrial house style they produce for my carousels and covers. Direction, not description.
Flat is the default. Ask an image model for a person in a room and you get the person pasted onto the room, lit from everywhere, dead centre, no air between them and the wall. It is not wrong. It is just a render, and a render is what your audience scrolls past.
A film frame is built, not described. There is something close to the lens, a subject in the middle distance, a world falling away behind. The light comes from somewhere and the shadow side proves it. The camera stands in a particular place with a particular lens, and that choice decides how the viewer feels about what they see.
My carousel and cover style came out of these rules: black dominating most of the frame, hard rim light and volumetric shafts, smoke, one exaggerated physical metaphor doing the work of a paragraph. Every one of those choices is a depth, light or camera decision. The frameworks for depth, light logic and camera language are ideas from a course I took, retold here in my own words and pointed at the tools I run: Nano Banana Pro and GPT Image 2 in Higgsfield, Soul Cinematic for film stills of my own face.
This guide is the checklist. Three layers, six lighting questions, three camera questions. Answer them, and the prompt writes itself.
The three levels
You write cinematic lighting and 8k in the prompt and get an evenly lit render with the subject in the middle.
Every prompt answers the checklist first: foreground, midground, background; light source, quality, direction, colour, material; camera position, shot size, lens. The frames have air in them.
Your agent holds the house style and the checklist, writes every prompt from them, keeps the lighting logic consistent across a set, and rejects any output that fails the creative test.
The mental model
Describe what is in the picture and add the word cinematic.
Direct how the frame is built. Depth in layers, light with logic, a camera with a position and a lens. The subject is the last thing you describe, not the first.
| Role | Talks to you | Job |
|---|---|---|
| Depth | First | Foreground close to the lens, subject in the midground, world falling away behind, atmosphere between |
| Light | Second | Source, hard or soft, direction, colour temperature, how materials react, consistent across the set |
| Camera | Third | Where it stands, how much we see, which lens, which camera character |
| The metaphor | Covers and carousels | One exaggerated physical image carries the idea; the framework lives inside the scene |
The checklist, in order
1. Build the frame in three layers
Write the foreground first: leaves, a railing, a shoulder, a window edge, smoke, slightly out of focus, close to the lens. Then the midground: the subject. Then the background: the street, the skyline, the machine room falling away. Then put something in the air between them: haze, dust in a shaft of light, low fog, exhaust. That air is what separates the layers and makes the space feel physical. A foreground object is the fastest fix for a flat image.
2. Take the subject off centre
Dead centre is for a hero shot or a direct stare. Otherwise place the subject on a third, and leave room in the direction they look or move. Then use what is already in the location, a road, a rail, a row of lights, a bridge edge, and angle it toward them. The eye follows lines. If nothing points at the subject, the frame is static.
3. Give the light a source, then let the shadow prove it
Cinematic lighting is not a style word, it is logic. Where does the light come from: a window, a lamp, headlights, a doorway, a neon sign, a practical bulb in the scene. Hard (bare sun, a sharp lamp, drama and tension) or soft (a big window, overcast, calm and expensive). Which side does it hit: side light carves shape, backlight separates and rims, front light flattens. Say what the shadow side does. If you cannot answer these, the model invents even light and you get a render.
4. Set the emotion with colour temperature, and motivate it
Warm feels alive and intimate. Cool feels distant and controlled. Green feels wrong. Red feels like danger, but only if something in the scene makes it: an exit sign, a brake light, a safelight. Warm subject against a cool background is the classic separation. Repeat one accent colour through the frame so it reads as intentional. Then make the materials react: sharp highlights on metal, light passing through glass, fabric absorbing, wet ground reflecting. If every surface reacts the same, it is a render.
5. Answer the three camera questions
Where is the camera: eye level is honest, low is power or threat, high is small or exposed, three quarter is natural and dimensional, over the shoulder puts the viewer in the conversation. How much do we see: wide for the world, full for posture and outfit, medium for behaviour, close up for emotion, extreme close up for one detail; contrast between them is where the impact lives. How does the space feel: a wide lens (16 to 24mm) is deep and immersive, 40 to 50mm is calm and human, 85mm separates a portrait, 200mm compresses and stacks. Describe the lens and the camera position both, because the lens alone does not decide perspective.
6. Assemble the prompt, then run the creative test
Subject, environment in layers, camera position and shot size and lens, light source and direction and quality and colour, one or two mood words, style only if it helps. Attach three to five reference images of the look; rules give the law, references give the taste. Generate four. Then the test I use on every cover: would someone stop scrolling on the image alone; could they get the idea in three seconds; would they want to zoom in. Fail any of the three and regenerate. If text on the image fights the scene, generate the scene with no text and add the typography afterwards.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
I want this image framed like a film, not described. Idea: [one sentence]. Canvas: [4:5 / 16:9 / 21:9]. Before writing the prompt, answer these and show me: Depth: what is in the foreground close to the lens, what is the subject in the midground, what falls away in the background, and what is in the air between them (haze, dust, fog, smoke). Composition: which third the subject sits on, where the lead room is, which lines in the location point at them. Light: the source, hard or soft, which side it hits, what the shadow side does, colour temperature and why, one accent colour repeated, how the materials react. Camera: where it stands, the shot size, the lens in mm and why, and a camera character if useful. Mood: one or two words. Then assemble one prompt in this order: subject, environment in layers, camera, light, mood, style only if it helps. No words like masterpiece, hyper detailed or perfect lighting. I will attach [N] reference images of the look; say what to take from them. Model: [Nano Banana Pro / GPT Image 2 / Soul Cinematic]. Generate four, then score each: would someone stop scrolling on the image alone, could they get the idea in three seconds, would they zoom in. Tell me which rule each failure broke.
Same scene, four frames: one wide that establishes the light source, then three closer shots. Keep the same light direction, softness and colour in every frame; only the camera position, shot size and lens change. Write all four prompts so the shadow side is on the same side of the subject throughout. Same canvas, same references.
A cover for the hook: [headline]. Pick one exaggerated physical metaphor for the idea; never the literal object. Dark environment, black dominating most of the frame, hard rim light and volumetric shafts, smoke, one accent colour that means something. Leave the top third empty for the headline. Write the frame with the checklist, then tell me whether to put the type in the model or generate with no text and add it afterwards.
Chat window vs the cinematic models in Higgsfield
Layer 1 is any image chat with the checklist prompt and three references attached. It gets you most of the way. Layer 2 is Higgsfield, where the model choice and the settings do the rest.
Model by frame
Nano Banana Pro (nano_banana_pro) for reference heavy frames with people and up to 14 references. GPT Image 2 (gpt_image_2) for covers with typography, plus a transparent background option for compositing type afterwards. Soul Cinematic (soul_cinematic) for film stills of a trained identity, with a 21:9 option. Soul Location for environments with no one in frame.
Canvas
Carousels 4:5 (1080 by 1350), one independent image per slide, never a collage. Presentations and YouTube 16:9. Widescreen film frames 21:9 on Nano Banana Pro or Soul Cinematic.
Resolution
1k to test a composition, 2k to ship, 4k for anything printed or heavily cropped. Both main models default to 2k.
Typography fallback
If the engine mangles the display type, generate with no text and composite the type afterwards. Crisp type beats in model type every time they fight.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
Claude Code with the Higgsfield CLI runs the checklist, the references and the creative test. The skill below is the instructions; add your own palette and never list to it and it becomes your house style.
- Install and log in to the CLI (commands below).
- Save the skill file. Put three to five reference images of the look in a folder the agent can read.
- Say "frame it like a film" with the idea and the canvas: "a founder choked by four glowing constraints, 4:5 cover, headline space top third." The agent answers the checklist, writes the prompt, attaches the references, generates four and returns what passes the creative test.
- For a set, say so: it anchors the light in the first frame and holds it across the rest.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh higgsfield auth login higgsfield model get nano_banana_pro higgsfield model get soul_cinematic
ChatGPT (a project or a custom GPT)
ChatGPT with GPT Image is a strong Layer 1 for this, and it is the engine I have used for covers with heavy typography. The checklist is the same; the references are attached to the project.
- Create a Project. Paste the skill as instructions. Upload three to five reference images of the look.
- Give it the idea and the canvas. Ask it to answer the checklist first (layers, light, camera) and show you the answers before it writes the prompt.
- Generate four. Ask it to score each against the creative test and say which rule the failures broke.
- If the type is mangled, ask for the scene with no text and add the headline in your design tool.
Anything with an API (a token and a curl call)
The CLI is the API. One command per frame; the prompt is the assembled checklist; the references are repeated flags.
- Cover with typography space: gpt_image_2, 4:5, references attached, headline described as negative space if you plan to composite.
- Reference heavy frame with a person: nano_banana_pro, 4:5 or 21:9, up to 14 references.
- Film still of your own face: soul_cinematic with your Soul id, 16:9 or 21:9.
P="$(cat prompt.txt)" # the assembled checklist prompt # cover, room for a headline, composite type afterwards if it fights higgsfield generate create gpt_image_2 --prompt "$P" \ --image refs/ref-1.png --image refs/ref-2.png --image refs/ref-3.png \ --aspect_ratio 4:5 --resolution 2k --quality high --wait # directed frame with a person and references higgsfield generate create nano_banana_pro --prompt "$P" \ --image refs/ref-1.png --image refs/me-front.jpg --aspect_ratio 21:9 --resolution 2k --wait # film still of a trained identity higgsfield generate create soul_cinematic --soul-id $SOUL_ID --prompt "$P" \ --aspect_ratio 21:9 --quality 2k --wait
Failure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| Subject looks pasted on the background | Add a foreground object near the lens and atmosphere between the layers |
| Evenly lit, no shadow, looks like a render | Name the light source and the shadow side; cut the word cinematic |
| Every subject dead centre | Put them on a third with lead room; point a line at them |
| Red light looks like a filter | Motivate it: an exit sign, a brake light, a safelight in the scene |
| Four frames, four different lighting setups | Anchor the source in the wide shot; close ups follow it |
| Face distorted on the wide lens | Distortion is camera distance, not the lens; move the camera back and say so |
| Headline mangled in the image | Generate with no text; composite the type afterwards |
| Style drifted to generic AI art | The references are stale; refresh the three to five best outputs |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Higgsfield | Nano Banana Pro, GPT Image 2, Soul Cinematic and Soul Location in one place, web and command line. | no link, just use it |
| GPT Image 2 (ChatGPT) | Covers with display typography; the engine behind my carousel style. | no link, just use it |
| Nano Banana Pro (Gemini) | Reference heavy directed frames with people, widescreen ratios. | no link, just use it |
| Claude | Answers the checklist, assembles the prompt, runs the creative test. | Open |
The free skill
It is the cinematographer's checklist as instructions for an agent: build the frame in three layers with atmosphere between them, place the subject off centre with lead room and leading lines, give the light a source, a quality, a direction and a colour temperature and make materials react to it, then answer where the camera is, how much we see and how the space feels. It assembles the prompt from those answers, keeps the lighting anchored across a set, and scores outputs against the creative test.
--- name: frame-it-like-a-film description: Directs AI image prompts like a cinematographer: builds the frame in foreground, midground and background with atmosphere, places the subject off centre with lead room and leading lines, gives light a source, quality, direction and colour temperature, answers where the camera stands, how much we see and which lens, then assembles the prompt and scores outputs against the creative test. Holds a dark cinematic house style for covers and carousels. Trigger on "frame it like a film", "make this cinematic", "this looks flat", "cover for this hook", "carousel art", "directed frame". --- # Frame It Like A Film You are directing how a frame is built, not describing what is in it. The subject is the last thing you write. Answer the checklist first, show the answers, then assemble the prompt. ## The checklist **Depth (three layers plus air)** - Foreground: something close to the lens, slightly out of focus (leaves, railing, shoulder, window edge, smoke, car hood). - Midground: the subject. - Background: the world falling away (street, skyline, machine room, landscape). - Atmosphere between layers: haze, dust in a light shaft, low fog, smoke, exhaust. **Composition** - Subject on a third, not dead centre, unless it is a deliberate hero or direct stare. - Lead room in the direction the subject looks or moves. - Leading lines from the location (roads, rails, rows of lights, architecture, light beams) angled toward the subject. **Light logic (answer all six)** 1. Source: window, sun, lamp, headlights, doorway, neon, practical bulb, screen. May be off frame; the shadows show where it is. 2. Quality: hard (small direct source, drama, tension) or soft (large diffused source, calm, expensive, intimate). 3. Direction: side light carves shape; backlight separates and rims; front light flattens (beauty only). Say what the shadow side does. 4. Colour temperature: warm (alive, intimate), cool (distant, controlled), green (wrong, industrial), red (danger, urgency). Warm subject against cool background for separation. Repeat one accent colour. 5. Motivation: coloured light needs a believable source in the scene. 6. Materials: metal throws sharp highlights, glass passes light, fabric absorbs, wet ground and paint reflect. If everything reacts the same, it is a render. **Camera (three questions)** - Where is the camera: eye level (honest), low (power or threat), high (small, exposed), Dutch (unstable); frontal, three quarter, profile, over the shoulder, POV. - How much do we see: extreme wide (environment is the subject), wide (person and place), full (posture, outfit), medium (behaviour), close up (emotion), extreme close up (one detail). Contrast between sizes is where impact lives. - How does the space feel: 16 to 24mm deep and immersive, 40 to 50mm calm and human, 85mm portrait separation, 200mm compressed and stacked. State lens and camera position both; distortion comes from camera distance, not the lens. - Optional camera character: ARRI Alexa (drama), Sony Venice (clean low light), RED (sharp modern), medium format (glossy product). **Mood**: one or two words. **Style**: only if it helps. Never "masterpiece", "hyper detailed", "ultra glossy", "perfect lighting", "8k". ## Assemble the prompt Order: subject, environment in layers, camera (position, shot size, lens), light (source, quality, direction, colour, materials), mood, style. One paragraph. Say what to take from each attached reference (three to five images of the look; rules give the law, references give the taste). ## Sets Anchor the light source in the first (widest) frame. Every closer frame keeps the same direction, softness and colour; only camera position, shot size and lens change. ## Models and settings (Higgsfield) - Reference heavy frame with people: `nano_banana_pro`, up to 14 references, ratios include 4:5, 16:9, 21:9, resolution 1k test, 2k ship, 4k print. - Cover with typography: `gpt_image_2`, 4:5 for carousels, 16:9 for decks. If the type is mangled, generate with no text and composite the type afterwards. - Film still of a trained identity: `soul_cinematic --soul-id <id>`, 16:9 or 21:9. - Environments with no one in frame: `soul_location`, prompt only. ```bash higgsfield generate create nano_banana_pro --prompt "<assembled prompt>" \ --image refs/1.png --image refs/2.png --image refs/3.png \ --aspect_ratio 4:5 --resolution 2k --wait ``` ## House style (covers and carousels) Black dominates most of the frame. Dark industrial or machine room environments, hard rim light, volumetric shafts, smoke, film grain. Off white primary text, gold for key words, oxblood for danger, category colour only when it means something. One exaggerated physical metaphor carries the idea; the framework lives inside the scene as signage, gauges, screens, plaques. Negative space reserved for the headline. One independent image per slide, never a collage. Every human figure is the founder as described in the user's identity descriptor unless told otherwise. Never: flat vector, floating icon soup, stock photo energy, cheap neon, tiny text, logos or watermarks unless asked. ## The creative test Generate four. Fail any output that misses one of three: would someone stop scrolling on the image alone; could they get the idea in three seconds; would they want to zoom in. For each failure name the rule it broke (layer, composition, light, camera). Return only what passes, with the prompt used.
What done looks like at thirty days
- You can name the missing rule on any flat frame in five seconds
- A four frame set exists with one light source held across all of it
- Three to five reference images of your look live in a folder and get refreshed
- The house style is a skill your agent writes every prompt from
- Every cover you post passes the creative test before it goes out
Want the house style built for your brand?
Inside the AI CEO Lab, the Content Engine module builds your reference set, your style rules and the agent that turns a hook into a directed frame every day.