AI CEO Lab← All free guides
On this page
Ads · Avatar Video

AI UGC Ads: The Avatar Ads I Run For Clinics

The doctor will not film. Fine. Script it the way I script every talking head, put a presenter on it who reads human, give it a voice, generate the hook, the body and the call to action as separate clips, and let the variations come from swapping parts. The rules that keep it legal in a medical practice are half this page.

AI UGC Ads: The Avatar Ads I Run For Clinics

The best performing ad in most clinic accounts is a person talking to the camera for forty seconds. And the most common reason a clinic has no such ad is that the person will not film. Too busy, too self conscious, three retakes and a cancelled session. I have lost months of a launch to this.

Avatar video fixed the bottleneck, and then created a new one. The tools are now good enough that a generated presenter can deliver a script with real lip sync and hands that behave. The failure moved from can we make it to should this one ship. A presenter with plastic skin, a fake patient telling a made up story, a monthly payment with no terms. Any one of those does more damage than no ad at all.

The idea of a factory that takes a product, a presenter and a script and hands back a finished ad comes from a video ads course I took. What follows is the clinic version, on the tools I actually use, with the parts a medical practice cannot skip.

Three things carry the whole system. The script is written before any avatar exists, in the same hook, body, call to action structure I use for every filmed ad. The presenter is either a real member of staff with consent and an identity lock, or a clearly generic presenter, never someone pretending to be a patient. And every clip gets the same human read test as an image: if it looks AI, it does not ship.

The three levels

Level 1 · Manual

You ask the doctor to film, she says next week, next week becomes next quarter, and the account runs on statics only.

Level 2 · AI + connections

A written script, a locked presenter, a chosen voice, and a folder of hook, body and CTA clips you assemble into ads. A new hook is a new clip, not a new shoot.

Level 3 · Agents on cadence

Your agent writes the script set from your offer, generates the clips in batches, assembles the permutations, and queues the next hooks from what won. You review a contact sheet and approve the slate.

Connections for this guide: An avatar video tool that takes a reference image and a script (I use Higgsfield). An image model for the presenter's base image (Nano Banana Pro or the one inside ChatGPT). A voice, either a stock one or a clone made from a real staff member's own recording with consent. A caption tool. An editor for the stitch. The script skill below.

The mental model

Wrong

Type a script into an avatar tool, export one video, run it, and wonder why it looks like a deepfake selling a serum.

Right

Script in parts. Lock a presenter that reads human, real staff with consent or clearly generic. Choose a voice. Generate hook, body and CTA as separate clips, sized to their length. Stitch, caption, gate, ship. New ads come from swapping parts.

RoleTalks to youJob
The scriptFirstFive hooks, one body under 160 words, three CTAs; words counted per clip
The presenterSecondReal staff via identity lock with consent, or a generic presenter; never a patient
The voiceThirdStock, or a consented clone of the real person; one voice per presenter
The clipsFourthOne generation per part, batch, contact sheet, human read test
The stitchLastHook plus body plus CTA, captions, the gates, naming, then permutations

Script, presenter, voice, clips, stitch

1. Write the script before any avatar exists

Same structure as a filmed ad. Five hooks of three to five seconds, each a different shape: the buried truth, the before you do this, the contrarian, the identity call out, the first timer. One body under 160 words that does not repeat any hook, in first person, talking to one person, with one honest admission (it is not for everyone, results vary, there is downtime) and one concrete mechanism. Three calls to action: book the consult with the location named, the free resource for the people not ready to book, and a one word comment trigger only if the automation behind it is already live. That is fifteen combinations before a single clip is generated.

2. Size every line to its clip

Avatar models generate a fixed length per clip. Too many words for the length and the presenter speeds up or gets cut off. Too few and she fills the gap with sounds that are not words. So every part of the script carries a word count and a target length. A five second hook is a short sentence. A body under 160 words is three to four clips, and each clip's words are counted. Write the script for the tool, not the other way round.

3. Lock the presenter

Two honest options. A real member of staff, with written consent, from a clean portrait or a set of photos, with the identity lock instruction in every prompt: preserve the exact face, hair, skin tone, build and clothing, re-light and re-compose only, never alter the face. Or a generic presenter you build from a base image: stated age, average realistic build, real skin micro texture, natural asymmetry, no sheen, candid framing, natural light. Generate several, pick the one that could exist, and keep that image as the reference for every future clip. What is not an option: a generated person presented as a patient, telling a story that did not happen. In a medical practice that is a fabricated testimonial, and it is the fastest way to lose an ad account and a licence.

4. Give it a voice

A stock voice from the avatar tool is fine for a generic presenter. For real staff, either their own recorded delivery uploaded as the audio track, or a clone made from their own recording with their consent, used only for scripts they have approved. One voice per presenter, forever. The moment the doctor's face has three voices the account is a museum of uncanny valley.

5. Generate in parts, on a contact sheet

One generation per hook, per body clip, per CTA. Reference image attached, script line pasted, product or device reference attached if it appears, highest quality setting. Batch them, wait, download, and lay them out where you can see them side by side. The human read test on every clip: skin, hands, eyes, the mouth on consonants. At most one refinement per clip; if the second try still looks AI, reword the line or reshoot the reference, do not roll the dice a fifth time. Trim the one second where the hand went wrong; do not keep it because the rest was good.

6. Stitch, caption, gate, name

Hook plus body plus CTA in the editor. Captions burned in. A two line on-screen title over the first three seconds that withholds the payoff rather than labelling the topic. Then the gates: no claim that is not on the page, no cure or guarantee words, no drug brand names, any monthly figure with example, term and APR, zero em dashes. Where a platform asks you to disclose digitally created or altered realistic people, do it; the tick box costs nothing and an undisclosed synthetic doctor costs everything. Name each file by its parts so the account can report which hook and which CTA won.

7. Autopilot is the second month

The first month is done by hand so you learn what the tool does to your presenter and which hooks she delivers well. Then the agent takes the script set, generates the batch, and hands you a contact sheet. You approve. New ads are new hooks against the same body, a new CTA against the winning hook, a second presenter against the winning combo. The body gets rewritten only when the offer changes.

Starter prompts

Paste these as written. They are short on purpose, because the long ones drift.

The prompt: make me an avatar ad

You are writing an avatar video ad for my clinic, in parts. My offer: [procedure, price shape, city, the page it sends to, the free resource if one exists]. The presenter is [a real staff member, name and role, consent on file / a generic presenter]. Clip lengths my tool generates: [5, 10, 15] seconds. Write: five hooks of three to five seconds, each a different shape (buried truth, before you do this, contrarian, identity call out, first timer), each with a word count that fits its clip; one body under 160 words, first person, to one person, with one honest admission and one concrete mechanism, that pairs with every hook and repeats none of them, split into clips with the word count for each; three calls to action (book with the location named, the free resource, a one word comment trigger only if I confirm the automation is live). Then write the presenter description for the base image: stated age, average realistic build, real skin micro texture, natural asymmetry, no sheen, candid framing, natural light; for real staff write the identity lock line instead. Rules: the presenter is the clinic's spokesperson and never speaks as a patient or tells a story that did not happen; nothing that is not on my page; no cure, guarantee, or drug brand names; any monthly figure carries example, term and APR; zero em dashes; no phrase over five words borrowed from anywhere; no brochure voice. Return an assembly sheet: the first three hook plus body plus CTA combinations to generate, and a file name for each in the form id-format-hook-presenter-offer-date.

The read test prompt

Here are [N] generated clips of my presenter: [attach frames or describe]. For each, check skin texture, hand geometry, eye symmetry, the mouth on hard consonants, and whether the lighting could exist in that room. Mark each ship, refine once, or reword. For any refine, give me the one line to change in the prompt. For any reword, tell me why the line itself is the problem.

The next batch prompt

Last week's results by hook: [paste hook click rates and cost per lead by CTA from my CRM]. Keep the body. Write three new hooks in the shape of the winner without repeating its words, sized to a [5] second clip, and tell me which CTA to pair them with and why. Zero em dashes.

Layer 2

From a chat script to the avatar tool

The script and presenter description come out of any chat. The clips come out of an avatar video tool, and these are the settings that decide whether the result reads human.

Reference image, not a text description

Always attach the presenter's reference image. Left to chance, the tool invents a less real person every time. For real staff, attach the portrait and keep the identity lock line in the prompt. If a reference has been through a failed real face job, re-upload the photo fresh; a rejected reference keeps failing.

Clip length and quality

Pick the clip length first, then paste only the words that fit it. Highest quality setting for anything that ships; the cheaper tier is for testing whether a hook delivers. If the product or device appears, attach its photo too, and a photo of what is inside if the packaging hides it.

Voice

Stock voice from the tool for a generic presenter. For real staff, upload their own recording as the audio track or use a consented clone. Never let the tool pick a new voice per clip.

After the winner

Once a combination wins, a multiplier tool that takes one finished video and returns versions with a different background, pacing, crop, or on-screen text is the cheapest next ten ads you will ever make. The face and the words stay; everything around them moves.

Give it to your agent, three ways

Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.

An agent with connectors (Claude Desktop, Claude Code, Grok Bot)

Higgsfield publishes an MCP, so an agent can write the script, build the presenter, and submit the clip generations in one session. Claude is where I run it.

  1. Claude Desktop or claude.ai: add Higgsfield as a connector from its integrations page (it is a remote MCP you sign into with your Higgsfield account), then create a Project with the skill from the bottom of this page as the instructions.
  2. Upload the presenter reference (or the staff portrait with the consent note) and the offer page text. Say "make me an avatar ad for [procedure]". It writes the script set with word counts, then asks which parts to generate first.
  3. It submits the hook, body and CTA generations as a batch, waits, and shows you the results together. You run the read test; it refines at most once per clip.
  4. Claude Code: install the Higgsfield CLI and the agent drives it with the same skill (commands in the API block).

ChatGPT (a project or a custom GPT)

ChatGPT has no avatar video connector, so the honest split is: ChatGPT writes the script and generates the presenter's base image; the clips are made in the avatar tool by hand.

  1. Create a Project with the skill as the instructions. Upload the offer page text. Say "make me an avatar ad for [procedure]" and get the script set with word counts and the presenter description.
  2. Ask it to generate the presenter base image with its image model from that description (stated age, average build, real skin texture, natural light, candid framing). Generate several, keep the one that could exist. For real staff, skip this step: use their real portrait.
  3. Open the avatar tool, attach the reference, paste one part of the script per clip at the matching length, generate, and bring the clips back to ChatGPT only for the assembly sheet and the naming.

Anything with an API (a token and a curl call)

If the agent only speaks a terminal, the Higgsfield CLI does the whole thing: upload the reference, create a reusable avatar, and submit a talking clip with the script as the prompt. Tokens live in the CLI's own login, never in a prompt.

  1. Install the CLI and sign in once. Upload the presenter image and register it as a custom avatar so every future clip references the same identity.
  2. Submit one generation per script part with the avatar attached and the line as the prompt. For a real staff member's own voice, pass their recording as the reference audio.
  3. Download the results, lay them out, run the read test, then stitch with your editor or an ffmpeg concat.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
# reusable presenter identity
higgsfield upload create ./presenter.png            # returns an upload id
higgsfield marketing-studio avatars create --name "clinic-presenter" --image <upload_id>
# one clip per script part
printf '[{"id":"<avatar_id>","type":"custom"}]' > avatars.json
higgsfield generate create marketing_studio_video --avatars @avatars.json \
  --mode ugc --duration 5 --aspect_ratio 9:16 --resolution 720p \
  --prompt "Presenter speaks to camera, eye level, clinic treatment room, natural light. She says: 'Before you book a laser, give me sixty seconds.'" --wait
# check flags for your CLI version with: higgsfield model get marketing_studio_video --json

Failure modes

Every one of these has happened to me or to someone I set this up for.

FailureFix
Presenter speeds through the line or trails into gibberishCount words per clip length; rewrite the line to fit, not the length to fit the line
Skin looks like plastic, eyes too symmetricalReference image with real texture and natural light; stated age and average build in the prompt
Generated person introduced as a patient with a storyDelete it; the presenter is the clinic's spokesperson, never a patient
Doctor's face with a different voice every weekOne voice per presenter, their own recording or a consented clone
Ad ran undisclosed and the platform flagged itTick the digitally created disclosure wherever the platform asks
Kept a clip because most of it was fineTrim the bad second or regenerate that part alone
Same reference fails every jobRe-upload the photo fresh; a rejected reference stays rejected
Body repeats the hook so half the combinations sound brokenThe body must pair with any hook; rewrite it without the hook's line
Comment keyword with nothing behind itBuild the automation first or use the direct CTA

The tools I use for this

ToolWhat it is for here
HiggsfieldWhere the presenter lives and the clips are generated, by MCP, by CLI, or by hand.no link, just use it
Nano Banana ProGoogle's image model, for the presenter's base image and the device or product reference.no link, just use it
ChatGPTWrites the script set and can generate a base image when you have no tool connected.no link, just use it
ClaudeRuns the skill, drives the batch, builds the assembly sheet.Open
SubmagicCaptions on every finished ad.Get it
DaVinci Resolve or CapCutThe stitch: hook plus body plus CTA, trim the bad second, export.no link, just use it
GoHighLevelWhere the comment keyword and the free resource actually live, so the CTA is never a fake funnel.Get it
Some links are affiliate links. I only recommend tools I run in my own accounts.

The free skill

It writes a clinic avatar ad as parts, not as one video. Five hooks, one body under 160 words that pairs with any hook, three calls to action, each sized to the clip length it will be generated at so the presenter never rushes or drifts into gibberish. It writes the presenter description so the base image reads human, refuses patient impersonation and testimonials, runs every line through the anti-slop and compliance gates, and returns an assembly sheet: which hook, body and CTA combos to generate first and how to name them.

How to use it: copy the whole thing, paste it into your bot (or save it as a skill file if you use Claude Code), and say “make me an avatar ad”. It walks you through the rest. Works with any agent that can read your files.
ai-ugc-ads.md
---
name: avatar-ads-on-autopilot
description: Writes a clinic avatar video ad as parts (five hooks, one body under 160 words, three CTAs) with a word count sized to each clip length, writes the presenter description or the identity lock line for real staff, refuses patient impersonation and testimonials, runs the anti-slop and compliance gates on every spoken line, and returns an assembly sheet of hook plus body plus CTA combinations with file names. Trigger on "make me an avatar ad", "talking head ad without filming", "AI presenter ad", "script for my avatar".
---

# Avatar Ads On Autopilot

You are producing a talking head ad for a clinic, med spa, IV bar, or aesthetics
practice where nobody will film. The presenter is generated. The script is not. The
script is written first, in parts, and everything downstream is assembly.

Never invent a fact, price, review, patient, or result. Missing facts stay in brackets.

## Step 1: Intake (ask only what is missing)
- Offer: procedure, price shape, city, destination page, free resource (must exist).
- Presenter: a real staff member (name, role, consent on file, portrait available) or a
  generic presenter.
- Clip lengths the avatar tool generates (for example 5, 10, 15 seconds).
- Compliance sensitivities: claims the clinic will not make, before and after policy.

## Step 2: The script set, with word counts
Write parts, not a video:
- **Five hooks**, three to five seconds each, five different shapes: buried truth,
  before you do this, contrarian, identity call out, first timer. Each carries the
  word count and the clip length it fits.
- **One body under 160 words**, first person, spoken to one person, one honest
  admission (not for everyone, results vary, there is downtime), one concrete
  mechanism, zero menu dumps. It must pair with every hook and repeat none of them.
  Split it into clips and put the word count on each clip.
- **Three CTAs**: direct (book, location named), indirect (the free resource, only if
  it exists), comment trigger (one word, only if the human confirms the automation is
  live).

Words per clip rule: too many words and the presenter rushes or gets cut off; too few
and she fills the gap with non-words. Rewrite the line to fit the clip. Never stretch
the clip to fit the line.

## Step 3: The presenter
Generic presenter: write the base image description. Stated age, average realistic
build, real skin micro texture, natural asymmetry, no retouching sheen, candid framing,
natural light, eyes open, dignified. Tell the human to generate several and keep the one
that could exist on a phone camera.

Real staff: do not describe a face. Write the identity lock line for every prompt:
"Preserve the exact facial likeness, hairstyle, skin tone, build, and clothing as
photographed. Do not alter the face. Re-light and re-compose only." Require consent on
file before proceeding. One voice per presenter: their own recording as the audio
track, or a clone made with their consent for approved scripts only.

Hard rule: the presenter is the clinic's spokesperson. A generated person never speaks
as a patient, never says "I had this done", never tells a story that did not happen.
That is a fabricated testimonial. Refuse it and say why.

## Step 4: The gates (every spoken line and every on-screen word)
Anti-slop, two failures is a rewrite: no phrase over five words borrowed from any
source; no brochure voice; no banned words (game-changer, unlock, elevate, seamless,
holistic, journey, empower, transformative, cutting-edge, revolutionary); zero em
dashes; no claim without a number, a mechanism, or an honest scope; the anybody's-ad
test; the body must not repeat a hook.

Compliance: nothing that is not on the clinic's page or in a verified source; no cure,
treat, guarantee, prevent-illness language; no drug brand names; no FDA wording beyond
the exact clearance; any monthly figure carries example, term, APR, and "subject to
eligibility"; remind the human to tick the platform's digitally created or altered
disclosure when the presenter is synthetic.

## Step 5: Generation notes for the human (or for you, if a tool is connected)
- Always attach the presenter reference image. Never let the tool invent the person.
- Attach the device or product photo if it appears, and what is inside if packaging
  hides it.
- Highest quality tier for anything that ships. Cheaper tier only to test whether a
  hook delivers.
- One generation per part. Batch, then review all results together.
- Read test on every clip: skin, hands, eyes, the mouth on hard consonants, lighting
  that could exist in that room. At most one refinement per clip; then reword the line
  or re-shoot the reference. Trim a bad second rather than keeping it.
- A reference that has been through a failed real face job keeps failing. Re-upload it
  fresh.

## Step 6: The assembly sheet
Return a table of the first three combinations to generate and stitch:
hook, body, CTA, total length, file name in the form
`id-format-hook-presenter-offer-MMDD`. Then the next three, for after the first read.
Ad name equals file name. utm_content equals the ad name.

## Weekly mode
Given hook click rates and cost per lead by CTA from the CRM: keep the body, write
three new hooks in the winning shape without repeating its words, sized to the clip,
pair each with a CTA and say why. Recommend handing the winning finished ad to a
multiplier tool for background, pacing, crop and on-screen text variants; the face and
the words stay.

What done looks like at thirty days

  • A script set with word counts per clip lives in the campaign folder
  • One presenter, locked, with consent on file or a generic base image that passed the read test
  • Fifteen possible ads from parts; at least three live
  • Every clip passed the read test and every ad passed the compliance gate
  • Synthetic presenters disclosed where the platform asks
  • Next week's hooks are queued from this week's hook click rates

Want the script set, the presenter and the first fifteen clips made with you?

Inside the AI CEO Lab the Demand Engine module walks the script, the presenter lock, and the batch on your own offer, with the compliance lines built into the prompts.

Pick a side.

Most people read this and forget it by Friday.

The other kind builds the thing that week. They stop needing free guides, because they are too busy running actual systems.

Free guides stay free. The room is where the builds happen.