How To Make AI Images Look Real: Prompt, Model, Settings, References
Four things decide whether an image reads as a photo or as a render: the prompt, the model, the settings, and the references. Here is how I run all four together in Higgsfield with Nano Banana Pro and GPT Image 2, and the checklist I use before I let anything go out.
Most AI images fail in the first half second. Not because the idea was bad. Because the skin is plastic, the light comes from nowhere, and the background is too clean for the situation. Your patient, your prospect, your feed all clock it instantly, and once they clock it, they stop trusting the words next to it.
The fix is not a magic prompt. I learned this the expensive way. An image is four decisions made together: what you ask for, which engine you ask, the format and quality you set, and what you show it as a reference. Change one without the others and you are back to guessing.
The idea that image work comes down to those four things comes from a course I took. What follows is my version of it, with the tools I actually run: Higgsfield as the studio, Nano Banana Pro for anything with a real person or a reference, GPT Image 2 for anything with text or design, and Claude to write the prompt before I touch a generate button.
This guide is the day to day routine. Set the format first, write the prompt in six parts, pick the engine by the job, generate more than one, and run the realism check before you keep anything.
The three levels
You type one line into whatever image tool is open, get something plastic, tweak the words six times, and post the least bad one.
Every image goes through the same routine: format locked, six part prompt written by your agent, model chosen by the job, four outputs, realism check, keep one. Twenty minutes, repeatable.
Your agent runs the routine from the command line on a batch: it writes the prompts, submits them to Higgsfield, returns the URLs, and flags anything that fails the realism check before you look.
The mental model
Find the perfect prompt and reuse it everywhere.
Four decisions made together, every time: prompt, model, settings, references. When an image looks fake, one of the four is wrong. Find which.
| Role | Talks to you | Job |
|---|---|---|
| Prompt | First | Six parts: subject, environment, camera, lighting, mood, style. Written by your agent, edited by you |
| Model | Second | Nano Banana Pro for people, references and edits. GPT Image 2 for text, design and product labels |
| Settings | Third | One aspect ratio per project. 1k while exploring, 2k or 4k when the direction is right. Four outputs |
| References | Always | Your real photos, with the prompt saying exactly what to keep from each one |
The routine, step by step
1. Lock the format before anything else
Decide where the image lives and set the aspect ratio once. 9:16 for Reels and Stories, 4:5 for the Instagram feed and carousels, 1:1 for profile and static ads, 16:9 for YouTube and slides. The same prompt looks completely different on a different canvas, so never generate half a project in one ratio and switch. In Higgsfield that is the aspect_ratio flag; Nano Banana Pro accepts 1:1, 3:2, 2:3, 4:3, 3:4, 4:5, 5:4, 9:16, 16:9 and 21:9.
2. Write the prompt in six parts, with your agent
Do not write the prompt yourself in the box. Give Claude the rough idea and the six part shape: who or what is in frame, where it is, how the camera sees it, where the light comes from, one or two mood words, and a style only if it helps. Read what comes back and cut anything you did not ask for. Words like masterpiece, hyper detailed and perfect lighting push an image toward fake. Shot on an iPhone, candid, natural framing pull it back.
3. Pick the engine by the job
Nano Banana Pro is my default for anything with a person, a reference photo or an edit; it takes up to 14 reference images. GPT Image 2 is my default when there is text on the image, a label, a layout or a design element; add the word photorealism at the end when realism matters. Soul 2.0 or Soul Cinematic when I want editorial or film stills of a trained identity. If one engine will not give it to you after two tries, switching engines is a legitimate fix, not a failure.
4. Set quality by stage
Explore at 1k with four outputs. Once the direction is right, rerun the winner at 2k, or 4k for anything that will be printed or cropped. Credits go to the final, not the exploration. Higgsfield defaults to 2k on both Nano Banana Pro and GPT Image 2, so set 1k on purpose while testing.
5. Reference with instructions, never bare
Upload your own photo, your own product, your own room. Then tell the model what the reference is for. Keep the face, hair and skin tone but change the location. Use this only as a pose reference, not an identity reference. Keep the packaging, logo, shape and colour exactly. A reference without instructions gets the wrong thing copied.
6. Generate four, keep one, then run the realism check
One output should never judge an idea. Generate four, put them side by side, and score each against the giveaways: airbrushed skin with zero texture, studio light in a place that would not have it, a background too tidy for the moment, hands and eyes that are not quite right, text that is almost right. Anything that fails goes back to step two with the failing part changed, not the whole prompt.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
I need an image that does not look AI generated. Job: [what it is for and where it will be posted]. Aspect ratio: [9:16, 4:5, 1:1 or 16:9]. Rough idea: [one sentence]. Write me a prompt in six parts: subject (who or what, specific, including realistic skin texture and small imperfections if a person), environment (a real place, time of day, what is around), camera (shot size, angle, lens, and shot on an iPhone if it is social content), lighting (where the light comes from and what it does to the shadow side; it must match the place), mood (one or two words), style (only if it helps; no words like masterpiece, hyper detailed or perfect lighting). I will attach reference images: [list each one and what to keep from it]. Then tell me which model to run it on, Nano Banana Pro for people and references or GPT Image 2 for text and design, and give me the settings: resolution 1k to explore, 2k to finish, four outputs.
Score these four images against this list and tell me which to keep: airbrushed or plastic skin with no texture; studio quality light in a place that would not have it; a background too clean or too staged for the moment; hands, eyes or teeth that are wrong; text that is almost right; anything that looks like a render instead of a photo. For each failure, tell me which of the six prompt parts to change, not a whole new prompt.
Using the same six part structure and the same references, write me [N] prompts for [the content type], one per scene, all in [aspect ratio], keeping the lighting logic consistent across them. Return them as a numbered list I can paste one at a time, each with the model to run it on.
Chat window vs the Higgsfield CLI
Layer 1 is the Gemini app or ChatGPT: paste the prompt, attach the references, take what you get. It works and it is how I still test a quick idea. Layer 2 is Higgsfield from the command line, which is how the agent runs it. These are the settings that actually change.
Model ID, not model name
Nano Banana Pro is nano_banana_pro. Nano Banana 2 is nano_banana_flash. GPT Image 2 is gpt_image_2. Soul 2.0 is text2image_soul_v2. Run higgsfield model list if a name has moved.
Resolution is explicit
Both Nano Banana Pro and GPT Image 2 default to 2k and accept 1k, 2k and 4k. Pass resolution 1k while exploring. GPT Image 2 also has quality low, medium, high (default high) and a background flag with a transparent option for cutouts.
References are repeated flags
Each reference is its own image flag, path or upload id. Nano Banana Pro takes up to 14. Soul 2.0 and Soul Cinematic take at most one, because the identity comes from the trained Soul, not the photo.
Wait for the URL
Add wait to every create so the command blocks and prints the media URL. Without it you are polling by hand.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
The agent path is the Higgsfield CLI driven by Claude Code, with the skill below as the instructions. Higgsfield also exposes an MCP, which is the route if you prefer the agent to talk to it as a connector.
- Install the CLI (command below), then run higgsfield auth login once in a terminal and confirm with higgsfield account status.
- Save the skill file from the bottom of this page into your skills folder, or paste it into the agent's instructions.
- Put your reference photos in a folder the agent can read. Chat pasted images are not reachable from the command line.
- Say the trigger with the job: "make me an image that does not look AI: me at the front desk of the clinic, 4:5, for the feed." The agent writes the six part prompt, submits, and returns four URLs with a pass or fail each.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh higgsfield auth login higgsfield model list higgsfield model get nano_banana_pro # aspect ratios, resolution, reference limits
ChatGPT (a project or a custom GPT)
ChatGPT has no connector into Higgsfield, but it has GPT Image built in, and that is a fine Layer 1 for anything with text or design. For people and references, use the Gemini app for Nano Banana in the same way.
- Create a ChatGPT Project. Paste the skill as the project instructions.
- Upload the reference photos to the project. Say what each one is for.
- Ask for the six part prompt first, edit it, then say generate, with the aspect ratio and the word photorealism at the end for real scenes.
- Ask for four variations, then ask it to score each one against the realism check in the skill before you download anything.
Anything with an API (a token and a curl call)
If your agent only speaks a shell, the CLI is the API. One command per image, references as paths, the result URL on stdout.
- Explore at 1k with the prompt your agent wrote and your references as repeated image flags.
- Rerun the winner at 2k or 4k with the same prompt and seed of references.
- For text on the image, switch the model to gpt_image_2 and keep everything else.
# explore: people + references higgsfield generate create nano_banana_pro \ --prompt "$(cat prompt.txt)" \ --image ./refs/me-front.jpg --image ./refs/me-3q.jpg --image ./refs/clinic-desk.jpg \ --aspect_ratio 4:5 --resolution 1k --wait # finish: same prompt, higher resolution higgsfield generate create nano_banana_pro --prompt "$(cat prompt.txt)" \ --image ./refs/me-front.jpg --image ./refs/me-3q.jpg --image ./refs/clinic-desk.jpg \ --aspect_ratio 4:5 --resolution 2k --wait # text or design on the image higgsfield generate create gpt_image_2 --prompt "$(cat prompt.txt) photorealism" \ --aspect_ratio 4:5 --resolution 2k --quality high --wait
Failure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| Skin looks like a mannequin | Add texture, pores, a stray hair, asymmetry to the subject line; switch to Nano Banana Pro |
| Ring light glow on a beach | Name the real light source and let the shadow side prove it |
| Background looks like a hotel brochure | Add shot on an iPhone, candid, natural framing; remove studio and perfect from the prompt |
| Same prompt, three different canvases | Lock the aspect ratio per project before the first generation |
| Reference copied the pose instead of the face | Tell the model what the reference is for, every time |
| Text on the image is gibberish | Move the job to GPT Image 2 and put the exact text in the prompt |
| Judged the idea from one output | Four outputs minimum; the idea is fine, the roll was not |
| Spent credits at 4k exploring | 1k to find it, 2k or 4k only for the keeper |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Higgsfield | The studio. Every image model I use in one place, web and command line. | no link, just use it |
| Nano Banana Pro (Gemini) | My default for people, references and edits. Up to 14 references per generation. | no link, just use it |
| GPT Image 2 (ChatGPT) | My default for anything with text, a label or a designed layout. | no link, just use it |
| Claude | Writes the six part prompt and runs the routine as an agent. | Open |
The free skill
It is the routine as instructions for an agent: lock the aspect ratio for the project, write the prompt in six parts (subject, environment, camera, lighting, mood, style), choose Nano Banana Pro for people and references or GPT Image 2 for text and design, set resolution by stage, generate four, then score each output against the realism check and return only what passes.
--- name: ai-images-that-dont-look-ai description: Produces AI images that read as photographs, not renders, by making four decisions together: a six part prompt, the right model for the job, explicit settings, and instructed reference images. Runs on Higgsfield (Nano Banana Pro, GPT Image 2) from the CLI or by hand in a chat window. Trigger on "make me an image that does not look AI", "this looks fake", "generate a photo of me at", "make a realistic image of". --- # AI Images That Don't Look AI You are producing an image the viewer will accept as a photograph. A fake looking image has one of four decisions wrong: prompt, model, settings, references. Your job is to make all four on purpose, generate several, and return only what passes the realism check. ## Before you start - Confirm where the image will live. Set the aspect ratio from that and keep it for the whole project: 9:16 Reels and Stories, 4:5 Instagram feed and carousels, 1:1 profile and static ads, 16:9 YouTube and slides. - Ask for reference images as files on disk (not pasted into chat) and ask what each one is for: identity, pose, outfit, product, location, or style. - If the Higgsfield CLI is available, check `higgsfield account status`. If not logged in, ask the user to run `higgsfield auth login`. If the CLI is not available, produce the prompt and settings for the user to paste into the Gemini app or ChatGPT. ## Step 1: Write the prompt in six parts Write it as one paragraph, in this order. Every part present, nothing decorative. 1. **Subject**: who or what, specific. For a person: age range, hair, skin, clothing, expression, posture, what they are doing. Include realistic skin texture and small imperfections. Never a generic default figure when a real person is referenced. 2. **Environment**: a real place, time of day, what is around, textures, weather. 3. **Camera**: shot size, angle, lens or phone. For social content, "shot on an iPhone, candid, natural framing" unless the brief is a studio shot. 4. **Lighting**: where the light comes from, hard or soft, which side, and what the shadow side does. The light must be possible in that place. 5. **Mood**: one or two words. 6. **Style**: only if it helps. Never "masterpiece", "hyper detailed", "ultra glossy", "perfect lighting"; these push the output toward a render. For each reference, add one sentence saying what to keep and what to change. ## Step 2: Choose the model by the job - Person, reference photo, or edit of an existing image: **Nano Banana Pro** (`nano_banana_pro`, up to 14 references). - Text on the image, label, layout, design element: **GPT Image 2** (`gpt_image_2`); put the exact text in the prompt and add "photorealism" at the end for real scenes. - Editorial or film still of a trained identity: Soul 2.0 (`text2image_soul_v2`) or Soul Cinematic (`soul_cinematic`), one reference maximum. - Two failed tries on one engine: switch engines before rewriting the whole prompt. ## Step 3: Settings - Aspect ratio from the brief, every generation. - Resolution `1k` while exploring, `2k` to finish, `4k` for print or heavy crops. - Four outputs per prompt while exploring. One output never judges an idea. ```bash higgsfield generate create nano_banana_pro \ --prompt "<six part prompt>" \ --image ./refs/a.jpg --image ./refs/b.jpg \ --aspect_ratio 4:5 --resolution 1k --wait ``` If a flag is rejected, run `higgsfield model get <model>` and use what it lists. ## Step 4: The realism check Score every output. Fail on any of these: - Airbrushed or plastic skin, no texture, no asymmetry. - Studio quality light in a place that would not have it. - Background too clean or too staged for the moment. - Hands, eyes or teeth wrong. - Text almost right. - Overall "render" feel: every surface reacting to light the same way. For each failure name the prompt part to change (subject, environment, camera, lighting, mood, style) or the model to switch to. Do not rewrite the whole prompt. ## Step 5: Deliver Return only outputs that pass, as URLs or file paths, with one line each on why it passed. Then the exact prompt and settings used, so the winner can be rerun at a higher resolution with the same references. ## House rules - Real photos of the real person when a real identifiable person is needed. - Never a fake patient or a fake clinician presented as real. - Never invent a logo; composite the real file afterwards.
What done looks like at thirty days
- A reference folder of your own face, products and spaces that the agent can read by path
- Every image you post went through format, six part prompt, engine, four outputs, realism check
- You can say which of the four decisions was wrong on any fake looking image
- The agent runs the routine from the command line and returns scored candidates
- You have not posted plastic skin in a month
Want the whole image system built around your brand?
Inside the AI CEO Lab, the Content Engine module sets up your reference library, your locked prompts and the agent that runs them, so every image you post looks like you shot it.