Consistent AI Characters: The Avatar I Reuse Forever
One face that shows up the same in every image: yours, trained as a Higgsfield Soul, or an invented character locked with a sheet and a written descriptor. How I build it once, prove it holds across scenes, and never re-upload the same ten photos again.
The first AI image of you is easy. The fiftieth is the problem. By then the jaw has drifted, the beard has changed twice, the skin tone is a shade off, and your audience feels a stranger wearing your name. Consistency is the whole game with a character, and it is exactly the thing a single reference photo cannot give you.
I run one recurring figure through my carousels and covers, and every prompt that contains a person carries the same written descriptor word for word, backed by real photos of me as ground truth. That descriptor exists because the models kept inventing hair on a bald head the moment I left it out. Every line of it was a failed generation.
There are two honest ways to lock an identity, and I use both. For my own face: train a Higgsfield Soul once on real photos and reference it by id forever. For an invented character: generate a base face, build a character sheet, write the descriptor, and attach both every time. The idea of building the avatar in steps rather than one perfect roll comes from a course I took; the workflow below is mine.
One rule before you start. A real identifiable person is a real photo or a Soul trained on their real photos, with their consent. Never a fake patient, never a fake clinician presented as real.
The three levels
You upload the same selfie every time, get a cousin of yourself back, and pick the least wrong one.
Your identity is trained once as a Soul (or locked once as a sheet plus descriptor). Every new scene references it by id, in the aspect ratio of the project, and the face holds.
Your agent keeps the identity, the descriptor and the reference folder, generates scene batches on request, checks likeness against the ground truth photos, and only shows you what matched.
The mental model
A good reference photo is enough. Upload it every time and describe the scene.
Identity is a trained model or a sheet plus a written descriptor, attached to every prompt. References tell the model what to keep. The descriptor tells it what it will otherwise invent.
| Role | Talks to you | Job |
|---|---|---|
| Ground truth | First | Five to twenty real photos, varied angles and light, on disk. The thing every output is checked against |
| The Soul or the sheet | Once | Trained identity by id (your face) or a multi angle character sheet (invented face) |
| The descriptor | Every prompt | One paragraph, verbatim, naming the features the model drifts on |
| The scene prompt | Every image | Same person plus location, outfit, camera, light, expression, realism |
Build it once
1. Decide: your real face or an invented character
Your face means a Higgsfield Soul trained on your real photos, referenced by id. An invented character means a base face you generate, then lock with a sheet and a descriptor. Do not mix them: a Soul trained on an invented face is a Soul of a picture, and it drifts. Pick one per character and write it down.
2. Gather ground truth (real face)
Five to twenty photos, eight to twelve is the sweet spot. Front, three quarter left and right, slightly above and below. Indoor and outdoor light. Neutral, smiling, talking. Headshot, head and shoulders, full body. Sharp, eyes visible, one person per photo, no sunglasses, no hats, no heavy filters, nothing you would not normally wear. Put them in one folder. Every future check happens against this folder.
3. Train the Soul
One command with a one word name and the photo paths. Pick the soul-2 variant for images, soul-cinematic for film stills and video work. Training takes minutes; wait for it, capture the reference id, and store the id with the descriptor. From then on every Soul model call takes the id instead of the photos. Soul 2.0 and Soul Cinematic accept at most one extra reference image, because the identity is already in the Soul.
4. Or build the character sheet (invented face)
Give Claude the rough idea, get a six part prompt, generate a batch on Nano Banana Pro or GPT Image 2 with photorealism at the end, and pick the one face that could exist on a phone camera: texture, asymmetry, light that matches the room. Then a sheet from that image: front, side profile, three quarter, full body, same face and hair in every panel, 16:9 on GPT Image 2. The sheet is the reference you attach forever.
5. Write the descriptor, verbatim, and keep it
One paragraph that names every feature the model drifts on: skin tone, head (bald or not, say it), beard, build, glasses, the two or three things you always wear. Mine is included word for word in every prompt with a person in it, and it says completely bald shaved head because the moment I left that out the model gave me hair. Store it next to the Soul id. Your agent pastes it; you never retype it.
6. Prove it across five scenes
Same person plus a new location, outfit, camera, light and expression, in the project's aspect ratio. Five scenes, four outputs each. Put every keeper next to the ground truth folder and ask one question: would someone who knows this person recognise them at a glance. Close ups fail first. If a close up drifts, add references or fix the descriptor before you generate anything else.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
I am building a reusable AI avatar. It is [my real face / an invented character]. [If real: here is a folder of my photos; check them against this spec: 5 to 20 photos, 8 to 12 ideal, front and three quarter angles, indoor and outdoor light, neutral and smiling, headshot to full body, sharp, one person, no sunglasses or hats, no heavy filters. Tell me what is missing.] [If invented: here is the rough idea in one sentence: [idea]; write a six part prompt for a believable base face with real skin texture and imperfections, shot on a phone, and a second prompt for a character sheet with front, profile, three quarter and full body panels in 16:9.] Then write my identity descriptor: one paragraph, under 60 words, naming every feature a model would otherwise invent (skin tone, head and hair or bald, facial hair, build, glasses, signature clothing), written so it can be pasted verbatim at the start of every prompt. Tell me the training command for a Higgsfield Soul, or the references to attach if I am not training.
[Descriptor verbatim.] Same person. Scene: [location], wearing [outfit], [shot size and angle, phone or lens], [where the light comes from and which side is in shadow], [expression], realistic skin texture, no studio lighting. Aspect ratio [X]. Generate four. Then compare each to the ground truth photos and tell me which one drifted and where.
Using the same descriptor and identity, write ten scene prompts for a month of content: three at work, three lifestyle, two with [product], two cover shots with strong negative space for a headline. All in [aspect ratio], consistent lighting logic within each set. Number them so I can run one at a time.
Chat window vs a trained Soul
Layer 1 is attaching the sheet and the descriptor to every prompt in the Gemini app or ChatGPT. It works and it is how an invented character runs. Layer 2 is a trained Soul in Higgsfield, which is how my own face runs. The settings that change:
Training is its own command
higgsfield soul-id create with a name, the variant flag (soul-2 for images, soul-cinematic for film and video) and five to twenty image flags. Then soul-id wait on the returned id. Paid plan required.
Generation uses the id, not photos
text2image_soul_v2 for stills, soul_cinematic for cinematic frames, each with the soul-id flag. Both default to 2k quality and accept at most one extra image reference.
Aspect ratios differ by model
Soul 2.0 accepts 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3. Soul Cinematic adds 21:9. Neither offers 4:5, so feed carousels are generated 3:4 and cropped, or run on Nano Banana Pro with the sheet.
Invented characters stay on Nano Banana Pro
Up to 14 references: the sheet, the base face, the best previous keepers. The descriptor goes in the prompt every time.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
Claude Code with the Higgsfield CLI is the tested path. The agent keeps the Soul id, the descriptor and the reference folder, and generates scenes on request.
- Install and log in to the CLI (commands below). Confirm the plan allows Soul training with higgsfield account status.
- Save the skill file from the bottom of this page. Add the Soul id, the descriptor and the folder path to it once trained.
- Say "build my avatar" with the photo folder path. The agent checks the photos against the spec, trains, waits, stores the id, and writes the descriptor for you to edit.
- Then say a scene: "me at the whiteboard in the clinic, 3:4, warm window light from the left." It generates four, compares to ground truth, and returns what matched.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh higgsfield auth login higgsfield soul-id create --name "emeka" --soul-2 \ --image ./refs/me-01.jpg --image ./refs/me-02.jpg --image ./refs/me-03.jpg # 5 to 20 photos higgsfield soul-id wait <reference_id> higgsfield soul-id list
ChatGPT (a project or a custom GPT)
ChatGPT cannot train a Soul, so this is the invented character path, or a real face carried by references and the descriptor. It holds well enough for feed content; it drifts on close ups sooner than a Soul does.
- Create a Project. Paste the skill as instructions. Upload the character sheet, the base face and five real photos if it is a real person.
- Ask it to write the descriptor from the photos, naming every feature that could drift. Edit it. Save it in the project instructions.
- For every scene: the descriptor verbatim, then same person plus location, outfit, camera, light, expression, realism, then the aspect ratio. Four variations.
- Ask it to compare each output to the uploaded photos and say which one drifted, before you download.
Anything with an API (a token and a curl call)
The CLI is the API. Train once, then every scene is one command with the id. For an invented character, the same command on Nano Banana Pro with the sheet as references.
- Store SOUL_ID and the descriptor file once. Your agent reads both.
- Stills on text2image_soul_v2; film frames on soul_cinematic. Same prompt shape either way.
- Invented character: nano_banana_pro with the sheet and best keepers as repeated image flags, descriptor at the top of the prompt.
export SOUL_ID=<reference_id> DESC="$(cat descriptor.txt)" # your face, still higgsfield generate create text2image_soul_v2 --soul-id $SOUL_ID \ --prompt "$DESC standing at the front desk of a small clinic, candid iPhone photo, soft window light from the left, calm, realistic skin texture" \ --aspect_ratio 3:4 --quality 2k --wait # your face, cinematic frame higgsfield generate create soul_cinematic --soul-id $SOUL_ID \ --prompt "$DESC in a dark machine room, hard rim light from behind, smoke, low angle medium shot, 35mm" \ --aspect_ratio 16:9 --quality 2k --wait # invented character, sheet as reference higgsfield generate create nano_banana_pro \ --prompt "$DESC Use the attached sheet as the identity reference; keep face, hair, skin tone and build exactly. Scene: ..." \ --image ./char/sheet.png --image ./char/base.png --aspect_ratio 4:5 --resolution 2k --wait
Failure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| The face is a cousin of you | More varied training photos; close up references; descriptor names the drifting feature |
| Hair appeared on a bald head | Say it in the descriptor, verbatim, every prompt |
| Soul training failed | Fewer than 5 usable faces, sunglasses, group photos or filters; swap photos and retrain |
| Wide shots fine, close ups wrong | Test close ups first; add head and shoulders references |
| Same prompt, different beard every time | The descriptor is missing or paraphrased; paste it, do not retype it |
| Trained a Soul on a generated face | Invented characters use the sheet and descriptor on Nano Banana Pro, not a Soul |
| Face fine, skin plastic | Add texture and imperfection to the scene prompt; light must match the room |
| Fake patient in a testimonial style image | Do not. Real people, real consent, or no person |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Higgsfield | Soul training and every image model I use, web and command line. | no link, just use it |
| Nano Banana Pro (Gemini) | Invented characters and reference heavy scenes, up to 14 references. | no link, just use it |
| GPT Image 2 (ChatGPT) | Character sheets in 16:9 and base faces with photorealism. | no link, just use it |
| Claude | Writes the descriptor and the scene prompts, runs the checks as an agent. | Open |
The free skill
It is the identity lock as instructions for an agent: decide real face or invented character, gather or generate the references, train the Soul or build the sheet, write the verbatim descriptor, then generate every scene with the identity attached and check the result against ground truth before returning it.
--- name: reusable-ai-avatar description: Builds and operates a reusable AI identity: a Higgsfield Soul trained on a real person's photos, or an invented character locked with a multi angle sheet and a verbatim descriptor. Generates scenes with the identity attached and checks every output against ground truth photos before returning it. Trigger on "build my avatar", "train my face", "make my character consistent", "put me in a scene", "why does my avatar keep changing". --- # Reusable AI Avatar You are locking one identity so it appears the same across every image. Identity is never re-described from scratch. It is a trained Soul (real face) or a sheet plus a descriptor (invented face), attached to every prompt, and every output is checked against ground truth. ## House rules - A real identifiable person means their real photos, with their consent, trained as a Soul or attached as references. Never a fake patient or a fake clinician presented as real. Never a real public figure. - The descriptor is pasted verbatim. Never paraphrase it. ## Step 1: Decide the path Ask once: is this the user's (or a consenting person's) real face, or an invented character? Real face goes to Soul training. Invented face goes to sheet plus descriptor on Nano Banana Pro. Do not train a Soul on a generated face. ## Step 2: Ground truth (real face) Check the photo folder against this spec and report gaps: - 5 to 20 photos, 8 to 12 ideal. Files on disk, not pasted into chat. - Angles: front, three quarter left, three quarter right, slightly above and below. - Light: indoor and outdoor, soft and harsh. Expressions: neutral, smiling, talking. - Distances: headshot, head and shoulders, full body. - Sharp, eyes visible, one person per photo, no sunglasses, no hats, no heavy filters, no costumes. Around 1024px or larger. ## Step 3: Train the Soul (real face) Requires a paid Higgsfield plan. If `higgsfield account status` shows free, say so. ```bash higgsfield soul-id create --name "<oneword>" --soul-2 --image ./refs/a.jpg --image ./refs/b.jpg ... higgsfield soul-id wait <reference_id> ``` Use `--soul-cinematic` instead when the downstream use is film stills or video. Store the returned id with the descriptor. Training failures usually mean too few or too uniform faces, occlusion, or group photos; fix the photos and retrain. ## Step 4: Base face and sheet (invented character) 1. Six part prompt for the base face (subject, environment, camera, lighting, mood, style), phone camera realism, skin texture, small imperfections. Generate a batch on `nano_banana_pro` or `gpt_image_2` (add "photorealism"). Choose the one face that could exist on a phone. 2. Character sheet from that image: front, side profile, three quarter, full body, identical face and hair in every panel, `gpt_image_2`, 16:9. 3. Keep base and sheet in one folder. They are attached to every future prompt. ## Step 5: Write the descriptor One paragraph, under 60 words, naming every feature a model drifts on: skin tone, head and hair (say "completely bald" explicitly if so), facial hair, build, glasses, signature clothing. Written to be pasted at the start of every prompt. Store it as a file. If a feature keeps drifting in outputs, add it to the descriptor, not to the scene prompt. ## Step 6: Generate a scene Prompt shape: descriptor verbatim, then "same person", then location, outfit, shot size and angle, light source and shadow side, expression, realism details. Aspect ratio from the project. Four outputs. ```bash # real face higgsfield generate create text2image_soul_v2 --soul-id <id> --prompt "<descriptor> same person ..." --aspect_ratio 3:4 --quality 2k --wait higgsfield generate create soul_cinematic --soul-id <id> --prompt "..." --aspect_ratio 16:9 --quality 2k --wait # invented character higgsfield generate create nano_banana_pro --prompt "<descriptor> Use the attached sheet as identity reference; keep face, hair, skin tone, build exactly. Scene: ..." --image ./char/sheet.png --image ./char/base.png --aspect_ratio 4:5 --resolution 2k --wait ``` Soul models accept at most one extra image reference and do not offer 4:5; use 3:4 and crop, or run feed sizes on Nano Banana Pro with the sheet. ## Step 7: Likeness check Compare every output to the ground truth folder (or the sheet). Test close ups first; they drift before wide shots. Fail on: different jaw or nose, changed facial hair, invented hair, wrong skin tone, wrong build, plastic skin, light that does not match the scene. Return only what passes, with the prompt and settings used. Add the best keepers to the reference folder for future prompts.
What done looks like at thirty days
- One Soul id (or one sheet) and one verbatim descriptor stored with the reference folder
- Twenty plus scenes exist and the face matches ground truth in all of them, close ups included
- Your agent generates a scene from a sentence and checks likeness before you see it
- You have not uploaded the same selfie to a prompt box in a month
- Nobody in your audience has asked why you look different this week
Want your character built and wired into your content engine?
Inside the AI CEO Lab, the Content Engine module trains your identity, writes your descriptor and sets up the agent that turns one face into a month of carousels and covers.