AI CEO Lab← All free guides
On this page
Images · Realism

How To Make AI Images Look Real: Prompt, Model, Settings, References

Four things decide whether an image reads as a photo or as a render: the prompt, the model, the settings, and the references. Here is how I run all four together in Higgsfield with Nano Banana Pro and GPT Image 2, and the checklist I use before I let anything go out.

How To Make AI Images Look Real: Prompt, Model, Settings, References

Most AI images fail in the first half second. Not because the idea was bad. Because the skin is plastic, the light comes from nowhere, and the background is too clean for the situation. Your patient, your prospect, your feed all clock it instantly, and once they clock it, they stop trusting the words next to it.

The fix is not a magic prompt. I learned this the expensive way. An image is four decisions made together: what you ask for, which engine you ask, the format and quality you set, and what you show it as a reference. Change one without the others and you are back to guessing.

The idea that image work comes down to those four things comes from a course I took. What follows is my version of it, with the tools I actually run: Higgsfield as the studio, Nano Banana Pro for anything with a real person or a reference, GPT Image 2 for anything with text or design, and Claude to write the prompt before I touch a generate button.

This guide is the day to day routine. Set the format first, write the prompt in six parts, pick the engine by the job, generate more than one, and run the realism check before you keep anything.

The three levels

Level 1 · Manual

You type one line into whatever image tool is open, get something plastic, tweak the words six times, and post the least bad one.

Level 2 · AI + connections

Every image goes through the same routine: format locked, six part prompt written by your agent, model chosen by the job, four outputs, realism check, keep one. Twenty minutes, repeatable.

Level 3 · Agents on cadence

Your agent runs the routine from the command line on a batch: it writes the prompts, submits them to Higgsfield, returns the URLs, and flags anything that fails the realism check before you look.

Connections for this guide: A Higgsfield account with the CLI logged in (or the web studio if you are doing it by hand). Claude or ChatGPT to write prompts. A folder of your own reference photos: your face, your product, your clinic. The image skill below.

The mental model

Wrong

Find the perfect prompt and reuse it everywhere.

Right

Four decisions made together, every time: prompt, model, settings, references. When an image looks fake, one of the four is wrong. Find which.

RoleTalks to youJob
PromptFirstSix parts: subject, environment, camera, lighting, mood, style. Written by your agent, edited by you
ModelSecondNano Banana Pro for people, references and edits. GPT Image 2 for text, design and product labels
SettingsThirdOne aspect ratio per project. 1k while exploring, 2k or 4k when the direction is right. Four outputs
ReferencesAlwaysYour real photos, with the prompt saying exactly what to keep from each one

The routine, step by step

1. Lock the format before anything else

Decide where the image lives and set the aspect ratio once. 9:16 for Reels and Stories, 4:5 for the Instagram feed and carousels, 1:1 for profile and static ads, 16:9 for YouTube and slides. The same prompt looks completely different on a different canvas, so never generate half a project in one ratio and switch. In Higgsfield that is the aspect_ratio flag; Nano Banana Pro accepts 1:1, 3:2, 2:3, 4:3, 3:4, 4:5, 5:4, 9:16, 16:9 and 21:9.

2. Write the prompt in six parts, with your agent

Do not write the prompt yourself in the box. Give Claude the rough idea and the six part shape: who or what is in frame, where it is, how the camera sees it, where the light comes from, one or two mood words, and a style only if it helps. Read what comes back and cut anything you did not ask for. Words like masterpiece, hyper detailed and perfect lighting push an image toward fake. Shot on an iPhone, candid, natural framing pull it back.

3. Pick the engine by the job

Nano Banana Pro is my default for anything with a person, a reference photo or an edit; it takes up to 14 reference images. GPT Image 2 is my default when there is text on the image, a label, a layout or a design element; add the word photorealism at the end when realism matters. Soul 2.0 or Soul Cinematic when I want editorial or film stills of a trained identity. If one engine will not give it to you after two tries, switching engines is a legitimate fix, not a failure.

4. Set quality by stage

Explore at 1k with four outputs. Once the direction is right, rerun the winner at 2k, or 4k for anything that will be printed or cropped. Credits go to the final, not the exploration. Higgsfield defaults to 2k on both Nano Banana Pro and GPT Image 2, so set 1k on purpose while testing.

5. Reference with instructions, never bare

Upload your own photo, your own product, your own room. Then tell the model what the reference is for. Keep the face, hair and skin tone but change the location. Use this only as a pose reference, not an identity reference. Keep the packaging, logo, shape and colour exactly. A reference without instructions gets the wrong thing copied.

6. Generate four, keep one, then run the realism check

One output should never judge an idea. Generate four, put them side by side, and score each against the giveaways: airbrushed skin with zero texture, studio light in a place that would not have it, a background too tidy for the moment, hands and eyes that are not quite right, text that is almost right. Anything that fails goes back to step two with the failing part changed, not the whole prompt.

Starter prompts

Paste these as written. They are short on purpose, because the long ones drift.

The prompt: build the image

I need an image that does not look AI generated. Job: [what it is for and where it will be posted]. Aspect ratio: [9:16, 4:5, 1:1 or 16:9]. Rough idea: [one sentence]. Write me a prompt in six parts: subject (who or what, specific, including realistic skin texture and small imperfections if a person), environment (a real place, time of day, what is around), camera (shot size, angle, lens, and shot on an iPhone if it is social content), lighting (where the light comes from and what it does to the shadow side; it must match the place), mood (one or two words), style (only if it helps; no words like masterpiece, hyper detailed or perfect lighting). I will attach reference images: [list each one and what to keep from it]. Then tell me which model to run it on, Nano Banana Pro for people and references or GPT Image 2 for text and design, and give me the settings: resolution 1k to explore, 2k to finish, four outputs.

The realism check

Score these four images against this list and tell me which to keep: airbrushed or plastic skin with no texture; studio quality light in a place that would not have it; a background too clean or too staged for the moment; hands, eyes or teeth that are wrong; text that is almost right; anything that looks like a render instead of a photo. For each failure, tell me which of the six prompt parts to change, not a whole new prompt.

The batch prompt

Using the same six part structure and the same references, write me [N] prompts for [the content type], one per scene, all in [aspect ratio], keeping the lighting logic consistent across them. Return them as a numbered list I can paste one at a time, each with the model to run it on.

Layer 2

Chat window vs the Higgsfield CLI

Layer 1 is the Gemini app or ChatGPT: paste the prompt, attach the references, take what you get. It works and it is how I still test a quick idea. Layer 2 is Higgsfield from the command line, which is how the agent runs it. These are the settings that actually change.

Model ID, not model name

Nano Banana Pro is nano_banana_pro. Nano Banana 2 is nano_banana_flash. GPT Image 2 is gpt_image_2. Soul 2.0 is text2image_soul_v2. Run higgsfield model list if a name has moved.

Resolution is explicit

Both Nano Banana Pro and GPT Image 2 default to 2k and accept 1k, 2k and 4k. Pass resolution 1k while exploring. GPT Image 2 also has quality low, medium, high (default high) and a background flag with a transparent option for cutouts.

References are repeated flags

Each reference is its own image flag, path or upload id. Nano Banana Pro takes up to 14. Soul 2.0 and Soul Cinematic take at most one, because the identity comes from the trained Soul, not the photo.

Wait for the URL

Add wait to every create so the command blocks and prints the media URL. Without it you are polling by hand.

Give it to your agent, three ways

Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.

An agent with connectors (Claude Desktop, Claude Code, Grok Bot)

The agent path is the Higgsfield CLI driven by Claude Code, with the skill below as the instructions. Higgsfield also exposes an MCP, which is the route if you prefer the agent to talk to it as a connector.

  1. Install the CLI (command below), then run higgsfield auth login once in a terminal and confirm with higgsfield account status.
  2. Save the skill file from the bottom of this page into your skills folder, or paste it into the agent's instructions.
  3. Put your reference photos in a folder the agent can read. Chat pasted images are not reachable from the command line.
  4. Say the trigger with the job: "make me an image that does not look AI: me at the front desk of the clinic, 4:5, for the feed." The agent writes the six part prompt, submits, and returns four URLs with a pass or fail each.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
higgsfield model list
higgsfield model get nano_banana_pro   # aspect ratios, resolution, reference limits

ChatGPT (a project or a custom GPT)

ChatGPT has no connector into Higgsfield, but it has GPT Image built in, and that is a fine Layer 1 for anything with text or design. For people and references, use the Gemini app for Nano Banana in the same way.

  1. Create a ChatGPT Project. Paste the skill as the project instructions.
  2. Upload the reference photos to the project. Say what each one is for.
  3. Ask for the six part prompt first, edit it, then say generate, with the aspect ratio and the word photorealism at the end for real scenes.
  4. Ask for four variations, then ask it to score each one against the realism check in the skill before you download anything.

Anything with an API (a token and a curl call)

If your agent only speaks a shell, the CLI is the API. One command per image, references as paths, the result URL on stdout.

  1. Explore at 1k with the prompt your agent wrote and your references as repeated image flags.
  2. Rerun the winner at 2k or 4k with the same prompt and seed of references.
  3. For text on the image, switch the model to gpt_image_2 and keep everything else.
# explore: people + references
higgsfield generate create nano_banana_pro \
  --prompt "$(cat prompt.txt)" \
  --image ./refs/me-front.jpg --image ./refs/me-3q.jpg --image ./refs/clinic-desk.jpg \
  --aspect_ratio 4:5 --resolution 1k --wait

# finish: same prompt, higher resolution
higgsfield generate create nano_banana_pro --prompt "$(cat prompt.txt)" \
  --image ./refs/me-front.jpg --image ./refs/me-3q.jpg --image ./refs/clinic-desk.jpg \
  --aspect_ratio 4:5 --resolution 2k --wait

# text or design on the image
higgsfield generate create gpt_image_2 --prompt "$(cat prompt.txt) photorealism" \
  --aspect_ratio 4:5 --resolution 2k --quality high --wait

Failure modes

Every one of these has happened to me or to someone I set this up for.

FailureFix
Skin looks like a mannequinAdd texture, pores, a stray hair, asymmetry to the subject line; switch to Nano Banana Pro
Ring light glow on a beachName the real light source and let the shadow side prove it
Background looks like a hotel brochureAdd shot on an iPhone, candid, natural framing; remove studio and perfect from the prompt
Same prompt, three different canvasesLock the aspect ratio per project before the first generation
Reference copied the pose instead of the faceTell the model what the reference is for, every time
Text on the image is gibberishMove the job to GPT Image 2 and put the exact text in the prompt
Judged the idea from one outputFour outputs minimum; the idea is fine, the roll was not
Spent credits at 4k exploring1k to find it, 2k or 4k only for the keeper

The tools I use for this

ToolWhat it is for here
HiggsfieldThe studio. Every image model I use in one place, web and command line.no link, just use it
Nano Banana Pro (Gemini)My default for people, references and edits. Up to 14 references per generation.no link, just use it
GPT Image 2 (ChatGPT)My default for anything with text, a label or a designed layout.no link, just use it
ClaudeWrites the six part prompt and runs the routine as an agent.Open
Some links are affiliate links. I only recommend tools I run in my own accounts.

The free skill

It is the routine as instructions for an agent: lock the aspect ratio for the project, write the prompt in six parts (subject, environment, camera, lighting, mood, style), choose Nano Banana Pro for people and references or GPT Image 2 for text and design, set resolution by stage, generate four, then score each output against the realism check and return only what passes.

How to use it: copy the whole thing, paste it into your bot (or save it as a skill file if you use Claude Code), and say “make me an image that does not look AI”. It walks you through the rest. Works with any agent that can read your files.
make-ai-images-look-real.md
---
name: ai-images-that-dont-look-ai
description: Produces AI images that read as photographs, not renders, by making four decisions together: a six part prompt, the right model for the job, explicit settings, and instructed reference images. Runs on Higgsfield (Nano Banana Pro, GPT Image 2) from the CLI or by hand in a chat window. Trigger on "make me an image that does not look AI", "this looks fake", "generate a photo of me at", "make a realistic image of".
---

# AI Images That Don't Look AI

You are producing an image the viewer will accept as a photograph. A fake looking
image has one of four decisions wrong: prompt, model, settings, references. Your job
is to make all four on purpose, generate several, and return only what passes the
realism check.

## Before you start

- Confirm where the image will live. Set the aspect ratio from that and keep it for the
  whole project: 9:16 Reels and Stories, 4:5 Instagram feed and carousels, 1:1 profile
  and static ads, 16:9 YouTube and slides.
- Ask for reference images as files on disk (not pasted into chat) and ask what each
  one is for: identity, pose, outfit, product, location, or style.
- If the Higgsfield CLI is available, check `higgsfield account status`. If not logged
  in, ask the user to run `higgsfield auth login`. If the CLI is not available, produce
  the prompt and settings for the user to paste into the Gemini app or ChatGPT.

## Step 1: Write the prompt in six parts

Write it as one paragraph, in this order. Every part present, nothing decorative.

1. **Subject**: who or what, specific. For a person: age range, hair, skin, clothing,
   expression, posture, what they are doing. Include realistic skin texture and small
   imperfections. Never a generic default figure when a real person is referenced.
2. **Environment**: a real place, time of day, what is around, textures, weather.
3. **Camera**: shot size, angle, lens or phone. For social content, "shot on an iPhone,
   candid, natural framing" unless the brief is a studio shot.
4. **Lighting**: where the light comes from, hard or soft, which side, and what the
   shadow side does. The light must be possible in that place.
5. **Mood**: one or two words.
6. **Style**: only if it helps. Never "masterpiece", "hyper detailed", "ultra glossy",
   "perfect lighting"; these push the output toward a render.

For each reference, add one sentence saying what to keep and what to change.

## Step 2: Choose the model by the job

- Person, reference photo, or edit of an existing image: **Nano Banana Pro**
  (`nano_banana_pro`, up to 14 references).
- Text on the image, label, layout, design element: **GPT Image 2** (`gpt_image_2`);
  put the exact text in the prompt and add "photorealism" at the end for real scenes.
- Editorial or film still of a trained identity: Soul 2.0 (`text2image_soul_v2`) or
  Soul Cinematic (`soul_cinematic`), one reference maximum.
- Two failed tries on one engine: switch engines before rewriting the whole prompt.

## Step 3: Settings

- Aspect ratio from the brief, every generation.
- Resolution `1k` while exploring, `2k` to finish, `4k` for print or heavy crops.
- Four outputs per prompt while exploring. One output never judges an idea.

```bash
higgsfield generate create nano_banana_pro \
  --prompt "<six part prompt>" \
  --image ./refs/a.jpg --image ./refs/b.jpg \
  --aspect_ratio 4:5 --resolution 1k --wait
```

If a flag is rejected, run `higgsfield model get <model>` and use what it lists.

## Step 4: The realism check

Score every output. Fail on any of these:

- Airbrushed or plastic skin, no texture, no asymmetry.
- Studio quality light in a place that would not have it.
- Background too clean or too staged for the moment.
- Hands, eyes or teeth wrong.
- Text almost right.
- Overall "render" feel: every surface reacting to light the same way.

For each failure name the prompt part to change (subject, environment, camera,
lighting, mood, style) or the model to switch to. Do not rewrite the whole prompt.

## Step 5: Deliver

Return only outputs that pass, as URLs or file paths, with one line each on why it
passed. Then the exact prompt and settings used, so the winner can be rerun at a
higher resolution with the same references.

## House rules

- Real photos of the real person when a real identifiable person is needed.
- Never a fake patient or a fake clinician presented as real.
- Never invent a logo; composite the real file afterwards.

What done looks like at thirty days

  • A reference folder of your own face, products and spaces that the agent can read by path
  • Every image you post went through format, six part prompt, engine, four outputs, realism check
  • You can say which of the four decisions was wrong on any fake looking image
  • The agent runs the routine from the command line and returns scored candidates
  • You have not posted plastic skin in a month

Want the whole image system built around your brand?

Inside the AI CEO Lab, the Content Engine module sets up your reference library, your locked prompts and the agent that runs them, so every image you post looks like you shot it.

Pick a side.

Most people read this and forget it by Friday.

The other kind builds the thing that week. They stop needing free guides, because they are too busy running actual systems.

Free guides stay free. The room is where the builds happen.