AI CEO Lab← All free guides
On this page
Images · Identity

Consistent AI Characters: The Avatar I Reuse Forever

One face that shows up the same in every image: yours, trained as a Higgsfield Soul, or an invented character locked with a sheet and a written descriptor. How I build it once, prove it holds across scenes, and never re-upload the same ten photos again.

Consistent AI Characters: The Avatar I Reuse Forever

The first AI image of you is easy. The fiftieth is the problem. By then the jaw has drifted, the beard has changed twice, the skin tone is a shade off, and your audience feels a stranger wearing your name. Consistency is the whole game with a character, and it is exactly the thing a single reference photo cannot give you.

I run one recurring figure through my carousels and covers, and every prompt that contains a person carries the same written descriptor word for word, backed by real photos of me as ground truth. That descriptor exists because the models kept inventing hair on a bald head the moment I left it out. Every line of it was a failed generation.

There are two honest ways to lock an identity, and I use both. For my own face: train a Higgsfield Soul once on real photos and reference it by id forever. For an invented character: generate a base face, build a character sheet, write the descriptor, and attach both every time. The idea of building the avatar in steps rather than one perfect roll comes from a course I took; the workflow below is mine.

One rule before you start. A real identifiable person is a real photo or a Soul trained on their real photos, with their consent. Never a fake patient, never a fake clinician presented as real.

The three levels

Level 1 · Manual

You upload the same selfie every time, get a cousin of yourself back, and pick the least wrong one.

Level 2 · AI + connections

Your identity is trained once as a Soul (or locked once as a sheet plus descriptor). Every new scene references it by id, in the aspect ratio of the project, and the face holds.

Level 3 · Agents on cadence

Your agent keeps the identity, the descriptor and the reference folder, generates scene batches on request, checks likeness against the ground truth photos, and only shows you what matched.

Connections for this guide: A Higgsfield account on a paid plan (Soul training needs one), the CLI logged in, five to twenty real photos of the person on disk, and Claude to write the descriptor and scene prompts. The avatar skill below.

The mental model

Wrong

A good reference photo is enough. Upload it every time and describe the scene.

Right

Identity is a trained model or a sheet plus a written descriptor, attached to every prompt. References tell the model what to keep. The descriptor tells it what it will otherwise invent.

RoleTalks to youJob
Ground truthFirstFive to twenty real photos, varied angles and light, on disk. The thing every output is checked against
The Soul or the sheetOnceTrained identity by id (your face) or a multi angle character sheet (invented face)
The descriptorEvery promptOne paragraph, verbatim, naming the features the model drifts on
The scene promptEvery imageSame person plus location, outfit, camera, light, expression, realism

Build it once

1. Decide: your real face or an invented character

Your face means a Higgsfield Soul trained on your real photos, referenced by id. An invented character means a base face you generate, then lock with a sheet and a descriptor. Do not mix them: a Soul trained on an invented face is a Soul of a picture, and it drifts. Pick one per character and write it down.

2. Gather ground truth (real face)

Five to twenty photos, eight to twelve is the sweet spot. Front, three quarter left and right, slightly above and below. Indoor and outdoor light. Neutral, smiling, talking. Headshot, head and shoulders, full body. Sharp, eyes visible, one person per photo, no sunglasses, no hats, no heavy filters, nothing you would not normally wear. Put them in one folder. Every future check happens against this folder.

3. Train the Soul

One command with a one word name and the photo paths. Pick the soul-2 variant for images, soul-cinematic for film stills and video work. Training takes minutes; wait for it, capture the reference id, and store the id with the descriptor. From then on every Soul model call takes the id instead of the photos. Soul 2.0 and Soul Cinematic accept at most one extra reference image, because the identity is already in the Soul.

4. Or build the character sheet (invented face)

Give Claude the rough idea, get a six part prompt, generate a batch on Nano Banana Pro or GPT Image 2 with photorealism at the end, and pick the one face that could exist on a phone camera: texture, asymmetry, light that matches the room. Then a sheet from that image: front, side profile, three quarter, full body, same face and hair in every panel, 16:9 on GPT Image 2. The sheet is the reference you attach forever.

5. Write the descriptor, verbatim, and keep it

One paragraph that names every feature the model drifts on: skin tone, head (bald or not, say it), beard, build, glasses, the two or three things you always wear. Mine is included word for word in every prompt with a person in it, and it says completely bald shaved head because the moment I left that out the model gave me hair. Store it next to the Soul id. Your agent pastes it; you never retype it.

6. Prove it across five scenes

Same person plus a new location, outfit, camera, light and expression, in the project's aspect ratio. Five scenes, four outputs each. Put every keeper next to the ground truth folder and ask one question: would someone who knows this person recognise them at a glance. Close ups fail first. If a close up drifts, add references or fix the descriptor before you generate anything else.

Starter prompts

Paste these as written. They are short on purpose, because the long ones drift.

The prompt: lock the identity

I am building a reusable AI avatar. It is [my real face / an invented character]. [If real: here is a folder of my photos; check them against this spec: 5 to 20 photos, 8 to 12 ideal, front and three quarter angles, indoor and outdoor light, neutral and smiling, headshot to full body, sharp, one person, no sunglasses or hats, no heavy filters. Tell me what is missing.] [If invented: here is the rough idea in one sentence: [idea]; write a six part prompt for a believable base face with real skin texture and imperfections, shot on a phone, and a second prompt for a character sheet with front, profile, three quarter and full body panels in 16:9.] Then write my identity descriptor: one paragraph, under 60 words, naming every feature a model would otherwise invent (skin tone, head and hair or bald, facial hair, build, glasses, signature clothing), written so it can be pasted verbatim at the start of every prompt. Tell me the training command for a Higgsfield Soul, or the references to attach if I am not training.

The scene prompt

[Descriptor verbatim.] Same person. Scene: [location], wearing [outfit], [shot size and angle, phone or lens], [where the light comes from and which side is in shadow], [expression], realistic skin texture, no studio lighting. Aspect ratio [X]. Generate four. Then compare each to the ground truth photos and tell me which one drifted and where.

The library prompt

Using the same descriptor and identity, write ten scene prompts for a month of content: three at work, three lifestyle, two with [product], two cover shots with strong negative space for a headline. All in [aspect ratio], consistent lighting logic within each set. Number them so I can run one at a time.

Layer 2

Chat window vs a trained Soul

Layer 1 is attaching the sheet and the descriptor to every prompt in the Gemini app or ChatGPT. It works and it is how an invented character runs. Layer 2 is a trained Soul in Higgsfield, which is how my own face runs. The settings that change:

Training is its own command

higgsfield soul-id create with a name, the variant flag (soul-2 for images, soul-cinematic for film and video) and five to twenty image flags. Then soul-id wait on the returned id. Paid plan required.

Generation uses the id, not photos

text2image_soul_v2 for stills, soul_cinematic for cinematic frames, each with the soul-id flag. Both default to 2k quality and accept at most one extra image reference.

Aspect ratios differ by model

Soul 2.0 accepts 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3. Soul Cinematic adds 21:9. Neither offers 4:5, so feed carousels are generated 3:4 and cropped, or run on Nano Banana Pro with the sheet.

Invented characters stay on Nano Banana Pro

Up to 14 references: the sheet, the base face, the best previous keepers. The descriptor goes in the prompt every time.

Give it to your agent, three ways

Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.

An agent with connectors (Claude Desktop, Claude Code, Grok Bot)

Claude Code with the Higgsfield CLI is the tested path. The agent keeps the Soul id, the descriptor and the reference folder, and generates scenes on request.

  1. Install and log in to the CLI (commands below). Confirm the plan allows Soul training with higgsfield account status.
  2. Save the skill file from the bottom of this page. Add the Soul id, the descriptor and the folder path to it once trained.
  3. Say "build my avatar" with the photo folder path. The agent checks the photos against the spec, trains, waits, stores the id, and writes the descriptor for you to edit.
  4. Then say a scene: "me at the whiteboard in the clinic, 3:4, warm window light from the left." It generates four, compares to ground truth, and returns what matched.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
higgsfield soul-id create --name "emeka" --soul-2 \
  --image ./refs/me-01.jpg --image ./refs/me-02.jpg --image ./refs/me-03.jpg   # 5 to 20 photos
higgsfield soul-id wait <reference_id>
higgsfield soul-id list

ChatGPT (a project or a custom GPT)

ChatGPT cannot train a Soul, so this is the invented character path, or a real face carried by references and the descriptor. It holds well enough for feed content; it drifts on close ups sooner than a Soul does.

  1. Create a Project. Paste the skill as instructions. Upload the character sheet, the base face and five real photos if it is a real person.
  2. Ask it to write the descriptor from the photos, naming every feature that could drift. Edit it. Save it in the project instructions.
  3. For every scene: the descriptor verbatim, then same person plus location, outfit, camera, light, expression, realism, then the aspect ratio. Four variations.
  4. Ask it to compare each output to the uploaded photos and say which one drifted, before you download.

Anything with an API (a token and a curl call)

The CLI is the API. Train once, then every scene is one command with the id. For an invented character, the same command on Nano Banana Pro with the sheet as references.

  1. Store SOUL_ID and the descriptor file once. Your agent reads both.
  2. Stills on text2image_soul_v2; film frames on soul_cinematic. Same prompt shape either way.
  3. Invented character: nano_banana_pro with the sheet and best keepers as repeated image flags, descriptor at the top of the prompt.
export SOUL_ID=<reference_id>
DESC="$(cat descriptor.txt)"

# your face, still
higgsfield generate create text2image_soul_v2 --soul-id $SOUL_ID \
  --prompt "$DESC standing at the front desk of a small clinic, candid iPhone photo, soft window light from the left, calm, realistic skin texture" \
  --aspect_ratio 3:4 --quality 2k --wait

# your face, cinematic frame
higgsfield generate create soul_cinematic --soul-id $SOUL_ID \
  --prompt "$DESC in a dark machine room, hard rim light from behind, smoke, low angle medium shot, 35mm" \
  --aspect_ratio 16:9 --quality 2k --wait

# invented character, sheet as reference
higgsfield generate create nano_banana_pro \
  --prompt "$DESC Use the attached sheet as the identity reference; keep face, hair, skin tone and build exactly. Scene: ..." \
  --image ./char/sheet.png --image ./char/base.png --aspect_ratio 4:5 --resolution 2k --wait

Failure modes

Every one of these has happened to me or to someone I set this up for.

FailureFix
The face is a cousin of youMore varied training photos; close up references; descriptor names the drifting feature
Hair appeared on a bald headSay it in the descriptor, verbatim, every prompt
Soul training failedFewer than 5 usable faces, sunglasses, group photos or filters; swap photos and retrain
Wide shots fine, close ups wrongTest close ups first; add head and shoulders references
Same prompt, different beard every timeThe descriptor is missing or paraphrased; paste it, do not retype it
Trained a Soul on a generated faceInvented characters use the sheet and descriptor on Nano Banana Pro, not a Soul
Face fine, skin plasticAdd texture and imperfection to the scene prompt; light must match the room
Fake patient in a testimonial style imageDo not. Real people, real consent, or no person

The tools I use for this

ToolWhat it is for here
HiggsfieldSoul training and every image model I use, web and command line.no link, just use it
Nano Banana Pro (Gemini)Invented characters and reference heavy scenes, up to 14 references.no link, just use it
GPT Image 2 (ChatGPT)Character sheets in 16:9 and base faces with photorealism.no link, just use it
ClaudeWrites the descriptor and the scene prompts, runs the checks as an agent.Open
Some links are affiliate links. I only recommend tools I run in my own accounts.

The free skill

It is the identity lock as instructions for an agent: decide real face or invented character, gather or generate the references, train the Soul or build the sheet, write the verbatim descriptor, then generate every scene with the identity attached and check the result against ground truth before returning it.

How to use it: copy the whole thing, paste it into your bot (or save it as a skill file if you use Claude Code), and say “build my avatar”. It walks you through the rest. Works with any agent that can read your files.
consistent-ai-characters.md
---
name: reusable-ai-avatar
description: Builds and operates a reusable AI identity: a Higgsfield Soul trained on a real person's photos, or an invented character locked with a multi angle sheet and a verbatim descriptor. Generates scenes with the identity attached and checks every output against ground truth photos before returning it. Trigger on "build my avatar", "train my face", "make my character consistent", "put me in a scene", "why does my avatar keep changing".
---

# Reusable AI Avatar

You are locking one identity so it appears the same across every image. Identity is
never re-described from scratch. It is a trained Soul (real face) or a sheet plus a
descriptor (invented face), attached to every prompt, and every output is checked
against ground truth.

## House rules

- A real identifiable person means their real photos, with their consent, trained as
  a Soul or attached as references. Never a fake patient or a fake clinician presented
  as real. Never a real public figure.
- The descriptor is pasted verbatim. Never paraphrase it.

## Step 1: Decide the path

Ask once: is this the user's (or a consenting person's) real face, or an invented
character? Real face goes to Soul training. Invented face goes to sheet plus descriptor
on Nano Banana Pro. Do not train a Soul on a generated face.

## Step 2: Ground truth (real face)

Check the photo folder against this spec and report gaps:
- 5 to 20 photos, 8 to 12 ideal. Files on disk, not pasted into chat.
- Angles: front, three quarter left, three quarter right, slightly above and below.
- Light: indoor and outdoor, soft and harsh. Expressions: neutral, smiling, talking.
- Distances: headshot, head and shoulders, full body.
- Sharp, eyes visible, one person per photo, no sunglasses, no hats, no heavy filters,
  no costumes. Around 1024px or larger.

## Step 3: Train the Soul (real face)

Requires a paid Higgsfield plan. If `higgsfield account status` shows free, say so.

```bash
higgsfield soul-id create --name "<oneword>" --soul-2 --image ./refs/a.jpg --image ./refs/b.jpg ...
higgsfield soul-id wait <reference_id>
```

Use `--soul-cinematic` instead when the downstream use is film stills or video.
Store the returned id with the descriptor. Training failures usually mean too few or
too uniform faces, occlusion, or group photos; fix the photos and retrain.

## Step 4: Base face and sheet (invented character)

1. Six part prompt for the base face (subject, environment, camera, lighting, mood,
   style), phone camera realism, skin texture, small imperfections. Generate a batch
   on `nano_banana_pro` or `gpt_image_2` (add "photorealism"). Choose the one face that
   could exist on a phone.
2. Character sheet from that image: front, side profile, three quarter, full body,
   identical face and hair in every panel, `gpt_image_2`, 16:9.
3. Keep base and sheet in one folder. They are attached to every future prompt.

## Step 5: Write the descriptor

One paragraph, under 60 words, naming every feature a model drifts on: skin tone,
head and hair (say "completely bald" explicitly if so), facial hair, build, glasses,
signature clothing. Written to be pasted at the start of every prompt. Store it as a
file. If a feature keeps drifting in outputs, add it to the descriptor, not to the
scene prompt.

## Step 6: Generate a scene

Prompt shape: descriptor verbatim, then "same person", then location, outfit, shot size
and angle, light source and shadow side, expression, realism details. Aspect ratio from
the project. Four outputs.

```bash
# real face
higgsfield generate create text2image_soul_v2 --soul-id <id> --prompt "<descriptor> same person ..." --aspect_ratio 3:4 --quality 2k --wait
higgsfield generate create soul_cinematic     --soul-id <id> --prompt "..." --aspect_ratio 16:9 --quality 2k --wait
# invented character
higgsfield generate create nano_banana_pro --prompt "<descriptor> Use the attached sheet as identity reference; keep face, hair, skin tone, build exactly. Scene: ..." --image ./char/sheet.png --image ./char/base.png --aspect_ratio 4:5 --resolution 2k --wait
```

Soul models accept at most one extra image reference and do not offer 4:5; use 3:4 and
crop, or run feed sizes on Nano Banana Pro with the sheet.

## Step 7: Likeness check

Compare every output to the ground truth folder (or the sheet). Test close ups first;
they drift before wide shots. Fail on: different jaw or nose, changed facial hair,
invented hair, wrong skin tone, wrong build, plastic skin, light that does not match
the scene. Return only what passes, with the prompt and settings used. Add the best
keepers to the reference folder for future prompts.

What done looks like at thirty days

  • One Soul id (or one sheet) and one verbatim descriptor stored with the reference folder
  • Twenty plus scenes exist and the face matches ground truth in all of them, close ups included
  • Your agent generates a scene from a sentence and checks likeness before you see it
  • You have not uploaded the same selfie to a prompt box in a month
  • Nobody in your audience has asked why you look different this week

Want your character built and wired into your content engine?

Inside the AI CEO Lab, the Content Engine module trains your identity, writes your descriptor and sets up the agent that turns one face into a month of carousels and covers.

Pick a side.

Most people read this and forget it by Friday.

The other kind builds the thing that week. They stop needing free guides, because they are too busy running actual systems.

Free guides stay free. The room is where the builds happen.