AI CEO Lab← All free guides
On this page
Images · Cinematography

Cinematic AI Image Prompts: Depth, Lighting And Camera Angles

Depth in three layers, light with a source, a camera with a position and a lens. The cinematographer's questions I answer before I write any prompt, and the dark industrial house style they produce for my carousels and covers. Direction, not description.

Cinematic AI Image Prompts: Depth, Lighting And Camera Angles

Flat is the default. Ask an image model for a person in a room and you get the person pasted onto the room, lit from everywhere, dead centre, no air between them and the wall. It is not wrong. It is just a render, and a render is what your audience scrolls past.

A film frame is built, not described. There is something close to the lens, a subject in the middle distance, a world falling away behind. The light comes from somewhere and the shadow side proves it. The camera stands in a particular place with a particular lens, and that choice decides how the viewer feels about what they see.

My carousel and cover style came out of these rules: black dominating most of the frame, hard rim light and volumetric shafts, smoke, one exaggerated physical metaphor doing the work of a paragraph. Every one of those choices is a depth, light or camera decision. The frameworks for depth, light logic and camera language are ideas from a course I took, retold here in my own words and pointed at the tools I run: Nano Banana Pro and GPT Image 2 in Higgsfield, Soul Cinematic for film stills of my own face.

This guide is the checklist. Three layers, six lighting questions, three camera questions. Answer them, and the prompt writes itself.

The three levels

Level 1 · Manual

You write cinematic lighting and 8k in the prompt and get an evenly lit render with the subject in the middle.

Level 2 · AI + connections

Every prompt answers the checklist first: foreground, midground, background; light source, quality, direction, colour, material; camera position, shot size, lens. The frames have air in them.

Level 3 · Agents on cadence

Your agent holds the house style and the checklist, writes every prompt from them, keeps the lighting logic consistent across a set, and rejects any output that fails the creative test.

Connections for this guide: Higgsfield with the CLI logged in (or any image model in a chat window). Three to five reference images of the look you want, on disk. Claude to turn the checklist into prompts. The cinematography skill below.

The mental model

Wrong

Describe what is in the picture and add the word cinematic.

Right

Direct how the frame is built. Depth in layers, light with logic, a camera with a position and a lens. The subject is the last thing you describe, not the first.

RoleTalks to youJob
DepthFirstForeground close to the lens, subject in the midground, world falling away behind, atmosphere between
LightSecondSource, hard or soft, direction, colour temperature, how materials react, consistent across the set
CameraThirdWhere it stands, how much we see, which lens, which camera character
The metaphorCovers and carouselsOne exaggerated physical image carries the idea; the framework lives inside the scene

The checklist, in order

1. Build the frame in three layers

Write the foreground first: leaves, a railing, a shoulder, a window edge, smoke, slightly out of focus, close to the lens. Then the midground: the subject. Then the background: the street, the skyline, the machine room falling away. Then put something in the air between them: haze, dust in a shaft of light, low fog, exhaust. That air is what separates the layers and makes the space feel physical. A foreground object is the fastest fix for a flat image.

2. Take the subject off centre

Dead centre is for a hero shot or a direct stare. Otherwise place the subject on a third, and leave room in the direction they look or move. Then use what is already in the location, a road, a rail, a row of lights, a bridge edge, and angle it toward them. The eye follows lines. If nothing points at the subject, the frame is static.

3. Give the light a source, then let the shadow prove it

Cinematic lighting is not a style word, it is logic. Where does the light come from: a window, a lamp, headlights, a doorway, a neon sign, a practical bulb in the scene. Hard (bare sun, a sharp lamp, drama and tension) or soft (a big window, overcast, calm and expensive). Which side does it hit: side light carves shape, backlight separates and rims, front light flattens. Say what the shadow side does. If you cannot answer these, the model invents even light and you get a render.

4. Set the emotion with colour temperature, and motivate it

Warm feels alive and intimate. Cool feels distant and controlled. Green feels wrong. Red feels like danger, but only if something in the scene makes it: an exit sign, a brake light, a safelight. Warm subject against a cool background is the classic separation. Repeat one accent colour through the frame so it reads as intentional. Then make the materials react: sharp highlights on metal, light passing through glass, fabric absorbing, wet ground reflecting. If every surface reacts the same, it is a render.

5. Answer the three camera questions

Where is the camera: eye level is honest, low is power or threat, high is small or exposed, three quarter is natural and dimensional, over the shoulder puts the viewer in the conversation. How much do we see: wide for the world, full for posture and outfit, medium for behaviour, close up for emotion, extreme close up for one detail; contrast between them is where the impact lives. How does the space feel: a wide lens (16 to 24mm) is deep and immersive, 40 to 50mm is calm and human, 85mm separates a portrait, 200mm compresses and stacks. Describe the lens and the camera position both, because the lens alone does not decide perspective.

6. Assemble the prompt, then run the creative test

Subject, environment in layers, camera position and shot size and lens, light source and direction and quality and colour, one or two mood words, style only if it helps. Attach three to five reference images of the look; rules give the law, references give the taste. Generate four. Then the test I use on every cover: would someone stop scrolling on the image alone; could they get the idea in three seconds; would they want to zoom in. Fail any of the three and regenerate. If text on the image fights the scene, generate the scene with no text and add the typography afterwards.

Starter prompts

Paste these as written. They are short on purpose, because the long ones drift.

The prompt: direct the frame

I want this image framed like a film, not described. Idea: [one sentence]. Canvas: [4:5 / 16:9 / 21:9]. Before writing the prompt, answer these and show me: Depth: what is in the foreground close to the lens, what is the subject in the midground, what falls away in the background, and what is in the air between them (haze, dust, fog, smoke). Composition: which third the subject sits on, where the lead room is, which lines in the location point at them. Light: the source, hard or soft, which side it hits, what the shadow side does, colour temperature and why, one accent colour repeated, how the materials react. Camera: where it stands, the shot size, the lens in mm and why, and a camera character if useful. Mood: one or two words. Then assemble one prompt in this order: subject, environment in layers, camera, light, mood, style only if it helps. No words like masterpiece, hyper detailed or perfect lighting. I will attach [N] reference images of the look; say what to take from them. Model: [Nano Banana Pro / GPT Image 2 / Soul Cinematic]. Generate four, then score each: would someone stop scrolling on the image alone, could they get the idea in three seconds, would they zoom in. Tell me which rule each failure broke.

The set prompt

Same scene, four frames: one wide that establishes the light source, then three closer shots. Keep the same light direction, softness and colour in every frame; only the camera position, shot size and lens change. Write all four prompts so the shadow side is on the same side of the subject throughout. Same canvas, same references.

The cover prompt

A cover for the hook: [headline]. Pick one exaggerated physical metaphor for the idea; never the literal object. Dark environment, black dominating most of the frame, hard rim light and volumetric shafts, smoke, one accent colour that means something. Leave the top third empty for the headline. Write the frame with the checklist, then tell me whether to put the type in the model or generate with no text and add it afterwards.

Layer 2

Chat window vs the cinematic models in Higgsfield

Layer 1 is any image chat with the checklist prompt and three references attached. It gets you most of the way. Layer 2 is Higgsfield, where the model choice and the settings do the rest.

Model by frame

Nano Banana Pro (nano_banana_pro) for reference heavy frames with people and up to 14 references. GPT Image 2 (gpt_image_2) for covers with typography, plus a transparent background option for compositing type afterwards. Soul Cinematic (soul_cinematic) for film stills of a trained identity, with a 21:9 option. Soul Location for environments with no one in frame.

Canvas

Carousels 4:5 (1080 by 1350), one independent image per slide, never a collage. Presentations and YouTube 16:9. Widescreen film frames 21:9 on Nano Banana Pro or Soul Cinematic.

Resolution

1k to test a composition, 2k to ship, 4k for anything printed or heavily cropped. Both main models default to 2k.

Typography fallback

If the engine mangles the display type, generate with no text and composite the type afterwards. Crisp type beats in model type every time they fight.

Give it to your agent, three ways

Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.

An agent with connectors (Claude Desktop, Claude Code, Grok Bot)

Claude Code with the Higgsfield CLI runs the checklist, the references and the creative test. The skill below is the instructions; add your own palette and never list to it and it becomes your house style.

  1. Install and log in to the CLI (commands below).
  2. Save the skill file. Put three to five reference images of the look in a folder the agent can read.
  3. Say "frame it like a film" with the idea and the canvas: "a founder choked by four glowing constraints, 4:5 cover, headline space top third." The agent answers the checklist, writes the prompt, attaches the references, generates four and returns what passes the creative test.
  4. For a set, say so: it anchors the light in the first frame and holds it across the rest.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
higgsfield model get nano_banana_pro
higgsfield model get soul_cinematic

ChatGPT (a project or a custom GPT)

ChatGPT with GPT Image is a strong Layer 1 for this, and it is the engine I have used for covers with heavy typography. The checklist is the same; the references are attached to the project.

  1. Create a Project. Paste the skill as instructions. Upload three to five reference images of the look.
  2. Give it the idea and the canvas. Ask it to answer the checklist first (layers, light, camera) and show you the answers before it writes the prompt.
  3. Generate four. Ask it to score each against the creative test and say which rule the failures broke.
  4. If the type is mangled, ask for the scene with no text and add the headline in your design tool.

Anything with an API (a token and a curl call)

The CLI is the API. One command per frame; the prompt is the assembled checklist; the references are repeated flags.

  1. Cover with typography space: gpt_image_2, 4:5, references attached, headline described as negative space if you plan to composite.
  2. Reference heavy frame with a person: nano_banana_pro, 4:5 or 21:9, up to 14 references.
  3. Film still of your own face: soul_cinematic with your Soul id, 16:9 or 21:9.
P="$(cat prompt.txt)"   # the assembled checklist prompt

# cover, room for a headline, composite type afterwards if it fights
higgsfield generate create gpt_image_2 --prompt "$P" \
  --image refs/ref-1.png --image refs/ref-2.png --image refs/ref-3.png \
  --aspect_ratio 4:5 --resolution 2k --quality high --wait

# directed frame with a person and references
higgsfield generate create nano_banana_pro --prompt "$P" \
  --image refs/ref-1.png --image refs/me-front.jpg --aspect_ratio 21:9 --resolution 2k --wait

# film still of a trained identity
higgsfield generate create soul_cinematic --soul-id $SOUL_ID --prompt "$P" \
  --aspect_ratio 21:9 --quality 2k --wait

Failure modes

Every one of these has happened to me or to someone I set this up for.

FailureFix
Subject looks pasted on the backgroundAdd a foreground object near the lens and atmosphere between the layers
Evenly lit, no shadow, looks like a renderName the light source and the shadow side; cut the word cinematic
Every subject dead centrePut them on a third with lead room; point a line at them
Red light looks like a filterMotivate it: an exit sign, a brake light, a safelight in the scene
Four frames, four different lighting setupsAnchor the source in the wide shot; close ups follow it
Face distorted on the wide lensDistortion is camera distance, not the lens; move the camera back and say so
Headline mangled in the imageGenerate with no text; composite the type afterwards
Style drifted to generic AI artThe references are stale; refresh the three to five best outputs

The tools I use for this

ToolWhat it is for here
HiggsfieldNano Banana Pro, GPT Image 2, Soul Cinematic and Soul Location in one place, web and command line.no link, just use it
GPT Image 2 (ChatGPT)Covers with display typography; the engine behind my carousel style.no link, just use it
Nano Banana Pro (Gemini)Reference heavy directed frames with people, widescreen ratios.no link, just use it
ClaudeAnswers the checklist, assembles the prompt, runs the creative test.Open
Some links are affiliate links. I only recommend tools I run in my own accounts.

The free skill

It is the cinematographer's checklist as instructions for an agent: build the frame in three layers with atmosphere between them, place the subject off centre with lead room and leading lines, give the light a source, a quality, a direction and a colour temperature and make materials react to it, then answer where the camera is, how much we see and how the space feels. It assembles the prompt from those answers, keeps the lighting anchored across a set, and scores outputs against the creative test.

How to use it: copy the whole thing, paste it into your bot (or save it as a skill file if you use Claude Code), and say “frame it like a film”. It walks you through the rest. Works with any agent that can read your files.
cinematic-ai-image-prompts.md
---
name: frame-it-like-a-film
description: Directs AI image prompts like a cinematographer: builds the frame in foreground, midground and background with atmosphere, places the subject off centre with lead room and leading lines, gives light a source, quality, direction and colour temperature, answers where the camera stands, how much we see and which lens, then assembles the prompt and scores outputs against the creative test. Holds a dark cinematic house style for covers and carousels. Trigger on "frame it like a film", "make this cinematic", "this looks flat", "cover for this hook", "carousel art", "directed frame".
---

# Frame It Like A Film

You are directing how a frame is built, not describing what is in it. The subject is
the last thing you write. Answer the checklist first, show the answers, then assemble
the prompt.

## The checklist

**Depth (three layers plus air)**
- Foreground: something close to the lens, slightly out of focus (leaves, railing,
  shoulder, window edge, smoke, car hood).
- Midground: the subject.
- Background: the world falling away (street, skyline, machine room, landscape).
- Atmosphere between layers: haze, dust in a light shaft, low fog, smoke, exhaust.

**Composition**
- Subject on a third, not dead centre, unless it is a deliberate hero or direct stare.
- Lead room in the direction the subject looks or moves.
- Leading lines from the location (roads, rails, rows of lights, architecture, light
  beams) angled toward the subject.

**Light logic (answer all six)**
1. Source: window, sun, lamp, headlights, doorway, neon, practical bulb, screen. May be
   off frame; the shadows show where it is.
2. Quality: hard (small direct source, drama, tension) or soft (large diffused source,
   calm, expensive, intimate).
3. Direction: side light carves shape; backlight separates and rims; front light
   flattens (beauty only). Say what the shadow side does.
4. Colour temperature: warm (alive, intimate), cool (distant, controlled), green
   (wrong, industrial), red (danger, urgency). Warm subject against cool background
   for separation. Repeat one accent colour.
5. Motivation: coloured light needs a believable source in the scene.
6. Materials: metal throws sharp highlights, glass passes light, fabric absorbs, wet
   ground and paint reflect. If everything reacts the same, it is a render.

**Camera (three questions)**
- Where is the camera: eye level (honest), low (power or threat), high (small,
  exposed), Dutch (unstable); frontal, three quarter, profile, over the shoulder, POV.
- How much do we see: extreme wide (environment is the subject), wide (person and
  place), full (posture, outfit), medium (behaviour), close up (emotion), extreme close
  up (one detail). Contrast between sizes is where impact lives.
- How does the space feel: 16 to 24mm deep and immersive, 40 to 50mm calm and human,
  85mm portrait separation, 200mm compressed and stacked. State lens and camera
  position both; distortion comes from camera distance, not the lens.
- Optional camera character: ARRI Alexa (drama), Sony Venice (clean low light), RED
  (sharp modern), medium format (glossy product).

**Mood**: one or two words. **Style**: only if it helps. Never "masterpiece", "hyper
detailed", "ultra glossy", "perfect lighting", "8k".

## Assemble the prompt

Order: subject, environment in layers, camera (position, shot size, lens), light
(source, quality, direction, colour, materials), mood, style. One paragraph. Say what
to take from each attached reference (three to five images of the look; rules give the
law, references give the taste).

## Sets

Anchor the light source in the first (widest) frame. Every closer frame keeps the same
direction, softness and colour; only camera position, shot size and lens change.

## Models and settings (Higgsfield)

- Reference heavy frame with people: `nano_banana_pro`, up to 14 references, ratios
  include 4:5, 16:9, 21:9, resolution 1k test, 2k ship, 4k print.
- Cover with typography: `gpt_image_2`, 4:5 for carousels, 16:9 for decks. If the type
  is mangled, generate with no text and composite the type afterwards.
- Film still of a trained identity: `soul_cinematic --soul-id <id>`, 16:9 or 21:9.
- Environments with no one in frame: `soul_location`, prompt only.

```bash
higgsfield generate create nano_banana_pro --prompt "<assembled prompt>" \
  --image refs/1.png --image refs/2.png --image refs/3.png \
  --aspect_ratio 4:5 --resolution 2k --wait
```

## House style (covers and carousels)

Black dominates most of the frame. Dark industrial or machine room environments,
hard rim light, volumetric shafts, smoke, film grain. Off white primary text, gold for
key words, oxblood for danger, category colour only when it means something. One
exaggerated physical metaphor carries the idea; the framework lives inside the scene
as signage, gauges, screens, plaques. Negative space reserved for the headline. One
independent image per slide, never a collage. Every human figure is the founder as
described in the user's identity descriptor unless told otherwise. Never: flat vector,
floating icon soup, stock photo energy, cheap neon, tiny text, logos or watermarks
unless asked.

## The creative test

Generate four. Fail any output that misses one of three: would someone stop scrolling
on the image alone; could they get the idea in three seconds; would they want to zoom
in. For each failure name the rule it broke (layer, composition, light, camera). Return
only what passes, with the prompt used.

What done looks like at thirty days

  • You can name the missing rule on any flat frame in five seconds
  • A four frame set exists with one light source held across all of it
  • Three to five reference images of your look live in a folder and get refreshed
  • The house style is a skill your agent writes every prompt from
  • Every cover you post passes the creative test before it goes out

Want the house style built for your brand?

Inside the AI CEO Lab, the Content Engine module builds your reference set, your style rules and the agent that turns a hook into a directed frame every day.

Pick a side.

Most people read this and forget it by Friday.

The other kind builds the thing that week. They stop needing free guides, because they are too busy running actual systems.

Free guides stay free. The room is where the builds happen.