AI Product Photography Without A Photoshoot
One clean photo of the product, the label text typed into the prompt, and Higgsfield's product photoshoot modes on GPT Image 2. Studio shots, lifestyle scenes, a person holding it, a hero banner, an ad pack. No photographer, no props, no location.
A clinic sells a serum, a supplement, a device, a package. The photo of it is a phone snap on a counter under fluorescent light, and that snap is on the landing page, in the ads and in the carousel. Nobody books a photographer for a bottle. So the bottle looks cheap and the price on the page looks wrong next to it.
The whole thing can be done from one clean reference photo. The model already knows what a studio shot, a lifestyle scene and a person holding a product look like. What it does not know is your label. Leave the label to chance and you get gibberish text and a bottle that is the wrong size in a hand.
I run product images through Higgsfield's product photoshoot command, which picks a mode, assembles the photography prompt in the backend and submits it to GPT Image 2 at 2k. I do not freehand the prompt; the mode templates are better than mine. My job is the reference, the label text, the size context and the choice of mode. The rule that the label text and product details go into the prompt, and that a person holding the product needs a locked person first, comes from a course I took. The rest is my routine.
Ten modes, one command, three to five outputs per job. Here is the order I run them in.
The three levels
A phone snap of the product on the counter, in every ad, forever.
One clean reference photo and the label text, run through product photoshoot modes: studio, lifestyle, in hand, hero, ad pack. A full set in an afternoon.
Your agent keeps the product registered with its references and label text, and produces a fresh set for any campaign, season or platform on request, with the label checked before you see it.
The mental model
Upload the product photo and type make it look professional.
The model is a production assistant, not a magic button. Give it the product, the person, the label text, the size context and the job of the image. Pick the mode; let the backend write the photography.
| Role | Talks to you | Job |
|---|---|---|
| The reference | First | One clean product photo, two or three angles if you can, sharp, well lit, nothing else in frame |
| The label text | Every prompt | Every word on the packaging, typed exactly. Otherwise the model guesses |
| The mode | Per job | Studio, lifestyle, in hand, pin, hero, carousel, ad pack, virtual model, conceptual, restyle |
| The check | Before delivery | Label readable, scale right, hands believable, angle usable |
From one photo to a full set
1. Take one clean reference, properly
Phone is fine. Plain background, daylight from a window, product filling the frame, sharp, no glare on the label. Front, a three quarter angle, and the back if the label wraps. If size matters (a small vial, a large tub), one extra photo next to a hand or a familiar object. This is the photoshoot. Everything else is the model.
2. Type the label, word for word
Brand name, product name, the line under it, the volume, anything printed. Put it in the prompt every time. The model reproduces packaging well from a reference, but text is where it drifts, and a nearly right label is worse than none. If the product has a pattern or a texture that matters, describe that too.
3. Pick the mode by intent, not by keyword
product_shot for a clean catalog or landing page image. lifestyle_scene for the product in use, on a counter, in a bag, in a gym. closeup_product_with_person for hands applying or holding. hero_banner for the wide website or email header. social_carousel for three to ten connected slides. ad_creative_pack for a coordinated set of static ads. moodboard_pin for a vertical 2:3. virtual_model_tryout for something worn. conceptual_product for floating, splash or sculptural. restyle to change the season or aesthetic of a shot you already like. When two fit, the more specific wins.
4. Run the command, do not write the photography
higgsfield product-photoshoot create with the mode, a short intent line, the reference images and a count of three to five. The backend holds the photography vocabulary per mode and submits to GPT Image 2 at 2k. Calling GPT Image 2 directly with your own prompt skips that and the output is visibly worse. The backend also varies lighting, angle and palette across the count, so the variants are not copies.
5. The in hand shot needs a locked person first
A believable person holding the product toward the camera is the most useful product image there is, and the one that fails most. Type make the man hold the bottle and you get a blurred label and the wrong size. Generate or choose the person first (see the avatar guide), then pass the person and the product together with the label text and the size context. Check the hand, the scale and the label before you keep it.
6. Check, then iterate the winner
Label readable and correct. Proportions right against the size reference. Hands with the right number of fingers. Angle the platform can use. If it is close, iterate from the winner: add an accessory, widen the frame, change the angle, restyle for the season. Do not start from zero when a small change would finish it.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
I need product photos without a photoshoot. Product: [name]. Attached: [front photo, three quarter photo, size reference if any]. Label text, exactly as printed: [every word]. Use: [landing page / feed / ad pack / email header / carousel]. Platform: [where]. Pick the right mode for each use from: product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle. For each, write the short intent line for the command, include the label text and the product proportions, 2k, three variants. If you need a person, use [my Soul id / the attached photo of me]; keep face, skin tone and build exactly. Then check every output for label accuracy, scale, hands and angle and tell me which to keep.
Write a detailed image prompt where the attached person holds the attached product toward the camera. The label must be sharp, readable and match this text exactly: [label text]. The product is [size; e.g. palm sized, 30 ml]. Keep the person's face, skin tone and hands exactly as the reference. [Location], [light source and side], candid phone photo, realistic skin texture. Tell me what to check on the output before I use it.
Here is a product image I like and here is my product. Describe what makes the reference work, then write a prompt that rebuilds the idea with my product instead, my label text exactly as printed: [text], and change [the colours / the background / the props] so the result is original, not a copy.
Chat window vs the product photoshoot command
Layer 1 is ChatGPT or the Gemini app: upload the product and a person, ask the chat to write a detailed prompt including the label text, then generate with both images attached. It works and it is a good way to learn what a strong product prompt contains. Layer 2 is the Higgsfield command, where the mode templates are already written.
One command, one mode
higgsfield product-photoshoot create with mode, a short intent prompt, repeated image flags and count 1 to 10. Modes: product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle.
Model and resolution are fixed
Always GPT Image 2, always 2k. Do not swap the model for product work; the enhancer is built for it.
Aspect ratio comes from the mode
The backend picks a sensible default per mode. Override only when the platform demands it: 1:1, 4:5, 5:4, 3:4, 4:3, 2:3, 3:2, 9:16, 16:9.
Packs lock the visual system
For social_carousel and ad_creative_pack the count is the number of slides or variants, and the backend keeps the look consistent across them.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
Claude Code driving the Higgsfield CLI, with the skill below, is how I run it. The agent asks at most four short questions, picks the mode, runs the command and checks the label before it shows you anything.
- Install and log in to the CLI (commands below).
- Save the skill file. Put product photos in a folder and the label text in a file next to them.
- Say "make product photos" with the product and the use: "the recovery serum, for the landing page and a Meta ad pack." The agent picks product_shot and ad_creative_pack, runs both, checks labels and hands, and returns URLs.
- For an in hand shot, name the person: your Soul, your sheet, or a real photo of you.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh higgsfield auth login higgsfield product-photoshoot create --help
ChatGPT (a project or a custom GPT)
ChatGPT has GPT Image built in, so Layer 1 works entirely inside it. What you lose is the mode templates, so the chat has to write the photography itself.
- Create a Project. Paste the skill as instructions. Upload the product photos and the label text file.
- Ask for a detailed image prompt for the job (studio, lifestyle, in hand, hero), including every word on the label, the product proportions and a size reference. Edit it.
- Generate with the product photo attached, four variations, photorealism at the end.
- Ask it to check label, scale and hands on each before you download. For in hand shots, upload the person's photo too and say what to keep from it.
Anything with an API (a token and a curl call)
The CLI is the API. One command per job, references as paths, URLs on stdout. Count controls the variants.
- Studio, lifestyle and hero for the flagship, three to five each.
- In hand with a locked person: pass the person and the product together with the label text in the prompt.
- Restyle a winner for the season instead of regenerating.
LABEL="$(cat products/serum/label.txt)" higgsfield product-photoshoot create --mode product_shot \ --prompt "clean catalog shot of the serum bottle, label reads exactly: $LABEL" \ --image products/serum/front.jpg --image products/serum/3q.jpg --count 3 higgsfield product-photoshoot create --mode lifestyle_scene \ --prompt "the serum on a bathroom counter at morning, window light, label reads exactly: $LABEL" \ --image products/serum/front.jpg --count 3 higgsfield product-photoshoot create --mode closeup_product_with_person \ --prompt "the attached person holding the serum toward the camera, label sharp and readable: $LABEL, bottle is palm sized" \ --image products/serum/front.jpg --image products/serum/in-hand-size.jpg --image refs/me-front.jpg --count 3 higgsfield product-photoshoot create --mode restyle \ --prompt "same shot, winter promotion, quiet luxury" --image outputs/winner.jpg
Failure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| Label reads like a different language | Type the label text into the prompt; keep GPT Image 2 |
| Bottle is the size of a thermos in the hand | Add a size reference photo and state the size in the prompt |
| Hand has six fingers | Regenerate the variant; person first, product second; check hands before keeping |
| Called GPT Image 2 directly and it looked flat | Use the product photoshoot command; the mode templates are the difference |
| Ad pack slides look like five different brands | Use ad_creative_pack with a count; the backend locks the visual system |
| Copied a competitor's campaign image | Rebuild the idea with your product, colours and details |
| Reference photo had glare on the label | Reshoot the reference in window light; it takes a minute |
| Winter promo regenerated from scratch, lost the winner | restyle the winner instead |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Higgsfield | The product photoshoot command and its ten mode templates. | no link, just use it |
| GPT Image 2 (ChatGPT) | The model under every product mode. Best with label text in the prompt. | no link, just use it |
| Claude | Runs the routine, picks the mode, checks the label. | Open |
| Meta Ads Manager | Where the ad pack goes. Test the variants against each other. | no link, just use it |
The free skill
It is the product routine as instructions for an agent: get the clean reference and the exact label text, pick the mode by intent (studio, lifestyle, closeup with person, pin, hero, carousel, ad pack, virtual model, conceptual, restyle), run the product photoshoot command on GPT Image 2 at 2k with three to five variants, then check label, scale, hands and angle before returning URLs.
--- name: product-photos-no-photoshoot description: Produces brand quality product images from one clean reference photo using Higgsfield's product photoshoot modes on GPT Image 2: studio, lifestyle, closeup with person, pin, hero banner, carousel, ad pack, virtual model, conceptual, restyle. Enforces the label text rule, the size reference rule and person first for in hand shots, and checks label, scale and hands before delivery. Trigger on "make product photos", "product shot", "put my product in a scene", "someone holding my product", "ad images for my product". --- # Product Photos Without A Photoshoot You are turning a clean product reference into finished product visuals. The model is a production assistant. Give it the product, the label text, the size context, the person if there is one, and the job. Pick the mode; let the backend write the photography. ## Before you start - Check `higgsfield account status`; ask for `higgsfield auth login` if needed. - Ask for the product photo as a file (front, three quarter, back if the label wraps). If the reference has glare, blur or clutter, ask for a reshoot in window light. - Ask for the label text exactly as printed and store it in a file. - If scale matters, ask for one photo of the product next to a hand or a familiar object. - Ask at most four short questions, with labelled options. Skip anything obvious. ## Pick the mode by intent | Intent | Mode | | --- | --- | | Clean, catalog, white, landing page | `product_shot` | | In use, on a counter, in a gym, atmosphere | `lifestyle_scene` | | Hands applying, holding, demonstrating | `closeup_product_with_person` | | Vertical Pinterest pin | `moodboard_pin` | | Wide web, email or campaign header | `hero_banner` | | 3 to 10 connected slides | `social_carousel` | | Coordinated static ad variants | `ad_creative_pack` | | Worn or used by a model | `virtual_model_tryout` | | Floating, splash, surreal, sculptural | `conceptual_product` | | Change season or aesthetic of an existing shot | `restyle` | When two fit, the more specific wins. Pinterest wins on platform, banner wins on format, carousel wins on multi slide, closeup wins on genre. ## Run it ```bash higgsfield product-photoshoot create \ --mode <mode> \ --prompt "<short intent line, including: label reads exactly: <label text>; product size>" \ --image <reference> [--image <more>] \ --count 3 ``` - Always GPT Image 2 under the hood, always 2k. Never call `gpt_image_2` directly for product work and never write the photography prompt yourself; the mode enhancer is the quality. - Count 3 to 5 for single shots. For `social_carousel` and `ad_creative_pack`, count is the number of slides or variants; the backend locks the visual system across them. - Override aspect ratio only when the platform demands it (1:1, 4:5, 5:4, 3:4, 4:3, 2:3, 3:2, 9:16, 16:9). ## In hand shots Person first. Use the user's Soul, character sheet, or a real photo of them. Pass the person and the product together, label text in the prompt, size stated. "Make the person hold the bottle" alone produces a blurred label and the wrong scale. ## Check before delivery Fail any output on: label text not matching the file, wrong proportions against the size reference, wrong hands, an angle the platform cannot use, plastic skin on a person, light that does not match the scene. Iterate from the best near miss (add, widen, re angle, restyle) instead of starting over. ## Deliver URLs only, one per line, grouped by mode. No prompt text, no ids. Say which failed and why in one line each. ## House rules - Inspiration images are rebuilt with the user's product, colours and details, never copied. - Real logos come from real files, composited afterwards if the model gets them wrong. - Never a fake patient or clinician presented as real.
What done looks like at thirty days
- Every product has clean references and a label text file the agent can read
- The flagship has studio, lifestyle, hero and in hand images that pass the label and scale check
- An ad pack of variants is live and being tested against each other
- Seasonal versions come from restyle, not from scratch
- You have not used the counter snap in a month
Want the whole product image library built for your offer?
Inside the AI CEO Lab, the Demand Engine module builds the product set, the ad pack and the landing page images together, so the offer looks like it costs what it costs.