AI CEO Lab← All free guides
On this page
Images · Products

AI Product Photography Without A Photoshoot

One clean photo of the product, the label text typed into the prompt, and Higgsfield's product photoshoot modes on GPT Image 2. Studio shots, lifestyle scenes, a person holding it, a hero banner, an ad pack. No photographer, no props, no location.

AI Product Photography Without A Photoshoot

A clinic sells a serum, a supplement, a device, a package. The photo of it is a phone snap on a counter under fluorescent light, and that snap is on the landing page, in the ads and in the carousel. Nobody books a photographer for a bottle. So the bottle looks cheap and the price on the page looks wrong next to it.

The whole thing can be done from one clean reference photo. The model already knows what a studio shot, a lifestyle scene and a person holding a product look like. What it does not know is your label. Leave the label to chance and you get gibberish text and a bottle that is the wrong size in a hand.

I run product images through Higgsfield's product photoshoot command, which picks a mode, assembles the photography prompt in the backend and submits it to GPT Image 2 at 2k. I do not freehand the prompt; the mode templates are better than mine. My job is the reference, the label text, the size context and the choice of mode. The rule that the label text and product details go into the prompt, and that a person holding the product needs a locked person first, comes from a course I took. The rest is my routine.

Ten modes, one command, three to five outputs per job. Here is the order I run them in.

The three levels

Level 1 · Manual

A phone snap of the product on the counter, in every ad, forever.

Level 2 · AI + connections

One clean reference photo and the label text, run through product photoshoot modes: studio, lifestyle, in hand, hero, ad pack. A full set in an afternoon.

Level 3 · Agents on cadence

Your agent keeps the product registered with its references and label text, and produces a fresh set for any campaign, season or platform on request, with the label checked before you see it.

Connections for this guide: A Higgsfield account with the CLI logged in. One clean photo of each product (two or three angles is better), the exact label text, and a size reference if scale matters. Claude to run the command and check the output. The product skill below.

The mental model

Wrong

Upload the product photo and type make it look professional.

Right

The model is a production assistant, not a magic button. Give it the product, the person, the label text, the size context and the job of the image. Pick the mode; let the backend write the photography.

RoleTalks to youJob
The referenceFirstOne clean product photo, two or three angles if you can, sharp, well lit, nothing else in frame
The label textEvery promptEvery word on the packaging, typed exactly. Otherwise the model guesses
The modePer jobStudio, lifestyle, in hand, pin, hero, carousel, ad pack, virtual model, conceptual, restyle
The checkBefore deliveryLabel readable, scale right, hands believable, angle usable

From one photo to a full set

1. Take one clean reference, properly

Phone is fine. Plain background, daylight from a window, product filling the frame, sharp, no glare on the label. Front, a three quarter angle, and the back if the label wraps. If size matters (a small vial, a large tub), one extra photo next to a hand or a familiar object. This is the photoshoot. Everything else is the model.

2. Type the label, word for word

Brand name, product name, the line under it, the volume, anything printed. Put it in the prompt every time. The model reproduces packaging well from a reference, but text is where it drifts, and a nearly right label is worse than none. If the product has a pattern or a texture that matters, describe that too.

3. Pick the mode by intent, not by keyword

product_shot for a clean catalog or landing page image. lifestyle_scene for the product in use, on a counter, in a bag, in a gym. closeup_product_with_person for hands applying or holding. hero_banner for the wide website or email header. social_carousel for three to ten connected slides. ad_creative_pack for a coordinated set of static ads. moodboard_pin for a vertical 2:3. virtual_model_tryout for something worn. conceptual_product for floating, splash or sculptural. restyle to change the season or aesthetic of a shot you already like. When two fit, the more specific wins.

4. Run the command, do not write the photography

higgsfield product-photoshoot create with the mode, a short intent line, the reference images and a count of three to five. The backend holds the photography vocabulary per mode and submits to GPT Image 2 at 2k. Calling GPT Image 2 directly with your own prompt skips that and the output is visibly worse. The backend also varies lighting, angle and palette across the count, so the variants are not copies.

5. The in hand shot needs a locked person first

A believable person holding the product toward the camera is the most useful product image there is, and the one that fails most. Type make the man hold the bottle and you get a blurred label and the wrong size. Generate or choose the person first (see the avatar guide), then pass the person and the product together with the label text and the size context. Check the hand, the scale and the label before you keep it.

6. Check, then iterate the winner

Label readable and correct. Proportions right against the size reference. Hands with the right number of fingers. Angle the platform can use. If it is close, iterate from the winner: add an accessory, widen the frame, change the angle, restyle for the season. Do not start from zero when a small change would finish it.

Starter prompts

Paste these as written. They are short on purpose, because the long ones drift.

The prompt: make the set

I need product photos without a photoshoot. Product: [name]. Attached: [front photo, three quarter photo, size reference if any]. Label text, exactly as printed: [every word]. Use: [landing page / feed / ad pack / email header / carousel]. Platform: [where]. Pick the right mode for each use from: product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle. For each, write the short intent line for the command, include the label text and the product proportions, 2k, three variants. If you need a person, use [my Soul id / the attached photo of me]; keep face, skin tone and build exactly. Then check every output for label accuracy, scale, hands and angle and tell me which to keep.

The in hand prompt

Write a detailed image prompt where the attached person holds the attached product toward the camera. The label must be sharp, readable and match this text exactly: [label text]. The product is [size; e.g. palm sized, 30 ml]. Keep the person's face, skin tone and hands exactly as the reference. [Location], [light source and side], candid phone photo, realistic skin texture. Tell me what to check on the output before I use it.

The inspiration prompt

Here is a product image I like and here is my product. Describe what makes the reference work, then write a prompt that rebuilds the idea with my product instead, my label text exactly as printed: [text], and change [the colours / the background / the props] so the result is original, not a copy.

Layer 2

Chat window vs the product photoshoot command

Layer 1 is ChatGPT or the Gemini app: upload the product and a person, ask the chat to write a detailed prompt including the label text, then generate with both images attached. It works and it is a good way to learn what a strong product prompt contains. Layer 2 is the Higgsfield command, where the mode templates are already written.

One command, one mode

higgsfield product-photoshoot create with mode, a short intent prompt, repeated image flags and count 1 to 10. Modes: product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle.

Model and resolution are fixed

Always GPT Image 2, always 2k. Do not swap the model for product work; the enhancer is built for it.

Aspect ratio comes from the mode

The backend picks a sensible default per mode. Override only when the platform demands it: 1:1, 4:5, 5:4, 3:4, 4:3, 2:3, 3:2, 9:16, 16:9.

Packs lock the visual system

For social_carousel and ad_creative_pack the count is the number of slides or variants, and the backend keeps the look consistent across them.

Give it to your agent, three ways

Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.

An agent with connectors (Claude Desktop, Claude Code, Grok Bot)

Claude Code driving the Higgsfield CLI, with the skill below, is how I run it. The agent asks at most four short questions, picks the mode, runs the command and checks the label before it shows you anything.

  1. Install and log in to the CLI (commands below).
  2. Save the skill file. Put product photos in a folder and the label text in a file next to them.
  3. Say "make product photos" with the product and the use: "the recovery serum, for the landing page and a Meta ad pack." The agent picks product_shot and ad_creative_pack, runs both, checks labels and hands, and returns URLs.
  4. For an in hand shot, name the person: your Soul, your sheet, or a real photo of you.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
higgsfield product-photoshoot create --help

ChatGPT (a project or a custom GPT)

ChatGPT has GPT Image built in, so Layer 1 works entirely inside it. What you lose is the mode templates, so the chat has to write the photography itself.

  1. Create a Project. Paste the skill as instructions. Upload the product photos and the label text file.
  2. Ask for a detailed image prompt for the job (studio, lifestyle, in hand, hero), including every word on the label, the product proportions and a size reference. Edit it.
  3. Generate with the product photo attached, four variations, photorealism at the end.
  4. Ask it to check label, scale and hands on each before you download. For in hand shots, upload the person's photo too and say what to keep from it.

Anything with an API (a token and a curl call)

The CLI is the API. One command per job, references as paths, URLs on stdout. Count controls the variants.

  1. Studio, lifestyle and hero for the flagship, three to five each.
  2. In hand with a locked person: pass the person and the product together with the label text in the prompt.
  3. Restyle a winner for the season instead of regenerating.
LABEL="$(cat products/serum/label.txt)"

higgsfield product-photoshoot create --mode product_shot \
  --prompt "clean catalog shot of the serum bottle, label reads exactly: $LABEL" \
  --image products/serum/front.jpg --image products/serum/3q.jpg --count 3

higgsfield product-photoshoot create --mode lifestyle_scene \
  --prompt "the serum on a bathroom counter at morning, window light, label reads exactly: $LABEL" \
  --image products/serum/front.jpg --count 3

higgsfield product-photoshoot create --mode closeup_product_with_person \
  --prompt "the attached person holding the serum toward the camera, label sharp and readable: $LABEL, bottle is palm sized" \
  --image products/serum/front.jpg --image products/serum/in-hand-size.jpg --image refs/me-front.jpg --count 3

higgsfield product-photoshoot create --mode restyle \
  --prompt "same shot, winter promotion, quiet luxury" --image outputs/winner.jpg

Failure modes

Every one of these has happened to me or to someone I set this up for.

FailureFix
Label reads like a different languageType the label text into the prompt; keep GPT Image 2
Bottle is the size of a thermos in the handAdd a size reference photo and state the size in the prompt
Hand has six fingersRegenerate the variant; person first, product second; check hands before keeping
Called GPT Image 2 directly and it looked flatUse the product photoshoot command; the mode templates are the difference
Ad pack slides look like five different brandsUse ad_creative_pack with a count; the backend locks the visual system
Copied a competitor's campaign imageRebuild the idea with your product, colours and details
Reference photo had glare on the labelReshoot the reference in window light; it takes a minute
Winter promo regenerated from scratch, lost the winnerrestyle the winner instead

The tools I use for this

ToolWhat it is for here
HiggsfieldThe product photoshoot command and its ten mode templates.no link, just use it
GPT Image 2 (ChatGPT)The model under every product mode. Best with label text in the prompt.no link, just use it
ClaudeRuns the routine, picks the mode, checks the label.Open
Meta Ads ManagerWhere the ad pack goes. Test the variants against each other.no link, just use it
Some links are affiliate links. I only recommend tools I run in my own accounts.

The free skill

It is the product routine as instructions for an agent: get the clean reference and the exact label text, pick the mode by intent (studio, lifestyle, closeup with person, pin, hero, carousel, ad pack, virtual model, conceptual, restyle), run the product photoshoot command on GPT Image 2 at 2k with three to five variants, then check label, scale, hands and angle before returning URLs.

How to use it: copy the whole thing, paste it into your bot (or save it as a skill file if you use Claude Code), and say “make product photos”. It walks you through the rest. Works with any agent that can read your files.
ai-product-photography.md
---
name: product-photos-no-photoshoot
description: Produces brand quality product images from one clean reference photo using Higgsfield's product photoshoot modes on GPT Image 2: studio, lifestyle, closeup with person, pin, hero banner, carousel, ad pack, virtual model, conceptual, restyle. Enforces the label text rule, the size reference rule and person first for in hand shots, and checks label, scale and hands before delivery. Trigger on "make product photos", "product shot", "put my product in a scene", "someone holding my product", "ad images for my product".
---

# Product Photos Without A Photoshoot

You are turning a clean product reference into finished product visuals. The model is
a production assistant. Give it the product, the label text, the size context, the
person if there is one, and the job. Pick the mode; let the backend write the
photography.

## Before you start

- Check `higgsfield account status`; ask for `higgsfield auth login` if needed.
- Ask for the product photo as a file (front, three quarter, back if the label wraps).
  If the reference has glare, blur or clutter, ask for a reshoot in window light.
- Ask for the label text exactly as printed and store it in a file.
- If scale matters, ask for one photo of the product next to a hand or a familiar object.
- Ask at most four short questions, with labelled options. Skip anything obvious.

## Pick the mode by intent

| Intent | Mode |
| --- | --- |
| Clean, catalog, white, landing page | `product_shot` |
| In use, on a counter, in a gym, atmosphere | `lifestyle_scene` |
| Hands applying, holding, demonstrating | `closeup_product_with_person` |
| Vertical Pinterest pin | `moodboard_pin` |
| Wide web, email or campaign header | `hero_banner` |
| 3 to 10 connected slides | `social_carousel` |
| Coordinated static ad variants | `ad_creative_pack` |
| Worn or used by a model | `virtual_model_tryout` |
| Floating, splash, surreal, sculptural | `conceptual_product` |
| Change season or aesthetic of an existing shot | `restyle` |

When two fit, the more specific wins. Pinterest wins on platform, banner wins on
format, carousel wins on multi slide, closeup wins on genre.

## Run it

```bash
higgsfield product-photoshoot create \
  --mode <mode> \
  --prompt "<short intent line, including: label reads exactly: <label text>; product size>" \
  --image <reference> [--image <more>] \
  --count 3
```

- Always GPT Image 2 under the hood, always 2k. Never call `gpt_image_2` directly for
  product work and never write the photography prompt yourself; the mode enhancer is
  the quality.
- Count 3 to 5 for single shots. For `social_carousel` and `ad_creative_pack`, count is
  the number of slides or variants; the backend locks the visual system across them.
- Override aspect ratio only when the platform demands it (1:1, 4:5, 5:4, 3:4, 4:3,
  2:3, 3:2, 9:16, 16:9).

## In hand shots

Person first. Use the user's Soul, character sheet, or a real photo of them. Pass the
person and the product together, label text in the prompt, size stated. "Make the
person hold the bottle" alone produces a blurred label and the wrong scale.

## Check before delivery

Fail any output on: label text not matching the file, wrong proportions against the
size reference, wrong hands, an angle the platform cannot use, plastic skin on a
person, light that does not match the scene. Iterate from the best near miss (add,
widen, re angle, restyle) instead of starting over.

## Deliver

URLs only, one per line, grouped by mode. No prompt text, no ids. Say which failed and
why in one line each.

## House rules

- Inspiration images are rebuilt with the user's product, colours and details, never copied.
- Real logos come from real files, composited afterwards if the model gets them wrong.
- Never a fake patient or clinician presented as real.

What done looks like at thirty days

  • Every product has clean references and a label text file the agent can read
  • The flagship has studio, lifestyle, hero and in hand images that pass the label and scale check
  • An ad pack of variants is live and being tested against each other
  • Seasonal versions come from restyle, not from scratch
  • You have not used the counter snap in a month

Want the whole product image library built for your offer?

Inside the AI CEO Lab, the Demand Engine module builds the product set, the ad pack and the landing page images together, so the offer looks like it costs what it costs.

Pick a side.

Most people read this and forget it by Friday.

The other kind builds the thing that week. They stop needing free guides, because they are too busy running actual systems.

Free guides stay free. The room is where the builds happen.