AI UGC Ads: The Avatar Ads I Run For Clinics
The doctor will not film. Fine. Script it the way I script every talking head, put a presenter on it who reads human, give it a voice, generate the hook, the body and the call to action as separate clips, and let the variations come from swapping parts. The rules that keep it legal in a medical practice are half this page.
The best performing ad in most clinic accounts is a person talking to the camera for forty seconds. And the most common reason a clinic has no such ad is that the person will not film. Too busy, too self conscious, three retakes and a cancelled session. I have lost months of a launch to this.
Avatar video fixed the bottleneck, and then created a new one. The tools are now good enough that a generated presenter can deliver a script with real lip sync and hands that behave. The failure moved from can we make it to should this one ship. A presenter with plastic skin, a fake patient telling a made up story, a monthly payment with no terms. Any one of those does more damage than no ad at all.
The idea of a factory that takes a product, a presenter and a script and hands back a finished ad comes from a video ads course I took. What follows is the clinic version, on the tools I actually use, with the parts a medical practice cannot skip.
Three things carry the whole system. The script is written before any avatar exists, in the same hook, body, call to action structure I use for every filmed ad. The presenter is either a real member of staff with consent and an identity lock, or a clearly generic presenter, never someone pretending to be a patient. And every clip gets the same human read test as an image: if it looks AI, it does not ship.
The three levels
You ask the doctor to film, she says next week, next week becomes next quarter, and the account runs on statics only.
A written script, a locked presenter, a chosen voice, and a folder of hook, body and CTA clips you assemble into ads. A new hook is a new clip, not a new shoot.
Your agent writes the script set from your offer, generates the clips in batches, assembles the permutations, and queues the next hooks from what won. You review a contact sheet and approve the slate.
The mental model
Type a script into an avatar tool, export one video, run it, and wonder why it looks like a deepfake selling a serum.
Script in parts. Lock a presenter that reads human, real staff with consent or clearly generic. Choose a voice. Generate hook, body and CTA as separate clips, sized to their length. Stitch, caption, gate, ship. New ads come from swapping parts.
| Role | Talks to you | Job |
|---|---|---|
| The script | First | Five hooks, one body under 160 words, three CTAs; words counted per clip |
| The presenter | Second | Real staff via identity lock with consent, or a generic presenter; never a patient |
| The voice | Third | Stock, or a consented clone of the real person; one voice per presenter |
| The clips | Fourth | One generation per part, batch, contact sheet, human read test |
| The stitch | Last | Hook plus body plus CTA, captions, the gates, naming, then permutations |
Script, presenter, voice, clips, stitch
1. Write the script before any avatar exists
Same structure as a filmed ad. Five hooks of three to five seconds, each a different shape: the buried truth, the before you do this, the contrarian, the identity call out, the first timer. One body under 160 words that does not repeat any hook, in first person, talking to one person, with one honest admission (it is not for everyone, results vary, there is downtime) and one concrete mechanism. Three calls to action: book the consult with the location named, the free resource for the people not ready to book, and a one word comment trigger only if the automation behind it is already live. That is fifteen combinations before a single clip is generated.
2. Size every line to its clip
Avatar models generate a fixed length per clip. Too many words for the length and the presenter speeds up or gets cut off. Too few and she fills the gap with sounds that are not words. So every part of the script carries a word count and a target length. A five second hook is a short sentence. A body under 160 words is three to four clips, and each clip's words are counted. Write the script for the tool, not the other way round.
3. Lock the presenter
Two honest options. A real member of staff, with written consent, from a clean portrait or a set of photos, with the identity lock instruction in every prompt: preserve the exact face, hair, skin tone, build and clothing, re-light and re-compose only, never alter the face. Or a generic presenter you build from a base image: stated age, average realistic build, real skin micro texture, natural asymmetry, no sheen, candid framing, natural light. Generate several, pick the one that could exist, and keep that image as the reference for every future clip. What is not an option: a generated person presented as a patient, telling a story that did not happen. In a medical practice that is a fabricated testimonial, and it is the fastest way to lose an ad account and a licence.
4. Give it a voice
A stock voice from the avatar tool is fine for a generic presenter. For real staff, either their own recorded delivery uploaded as the audio track, or a clone made from their own recording with their consent, used only for scripts they have approved. One voice per presenter, forever. The moment the doctor's face has three voices the account is a museum of uncanny valley.
5. Generate in parts, on a contact sheet
One generation per hook, per body clip, per CTA. Reference image attached, script line pasted, product or device reference attached if it appears, highest quality setting. Batch them, wait, download, and lay them out where you can see them side by side. The human read test on every clip: skin, hands, eyes, the mouth on consonants. At most one refinement per clip; if the second try still looks AI, reword the line or reshoot the reference, do not roll the dice a fifth time. Trim the one second where the hand went wrong; do not keep it because the rest was good.
6. Stitch, caption, gate, name
Hook plus body plus CTA in the editor. Captions burned in. A two line on-screen title over the first three seconds that withholds the payoff rather than labelling the topic. Then the gates: no claim that is not on the page, no cure or guarantee words, no drug brand names, any monthly figure with example, term and APR, zero em dashes. Where a platform asks you to disclose digitally created or altered realistic people, do it; the tick box costs nothing and an undisclosed synthetic doctor costs everything. Name each file by its parts so the account can report which hook and which CTA won.
7. Autopilot is the second month
The first month is done by hand so you learn what the tool does to your presenter and which hooks she delivers well. Then the agent takes the script set, generates the batch, and hands you a contact sheet. You approve. New ads are new hooks against the same body, a new CTA against the winning hook, a second presenter against the winning combo. The body gets rewritten only when the offer changes.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
You are writing an avatar video ad for my clinic, in parts. My offer: [procedure, price shape, city, the page it sends to, the free resource if one exists]. The presenter is [a real staff member, name and role, consent on file / a generic presenter]. Clip lengths my tool generates: [5, 10, 15] seconds. Write: five hooks of three to five seconds, each a different shape (buried truth, before you do this, contrarian, identity call out, first timer), each with a word count that fits its clip; one body under 160 words, first person, to one person, with one honest admission and one concrete mechanism, that pairs with every hook and repeats none of them, split into clips with the word count for each; three calls to action (book with the location named, the free resource, a one word comment trigger only if I confirm the automation is live). Then write the presenter description for the base image: stated age, average realistic build, real skin micro texture, natural asymmetry, no sheen, candid framing, natural light; for real staff write the identity lock line instead. Rules: the presenter is the clinic's spokesperson and never speaks as a patient or tells a story that did not happen; nothing that is not on my page; no cure, guarantee, or drug brand names; any monthly figure carries example, term and APR; zero em dashes; no phrase over five words borrowed from anywhere; no brochure voice. Return an assembly sheet: the first three hook plus body plus CTA combinations to generate, and a file name for each in the form id-format-hook-presenter-offer-date.
Here are [N] generated clips of my presenter: [attach frames or describe]. For each, check skin texture, hand geometry, eye symmetry, the mouth on hard consonants, and whether the lighting could exist in that room. Mark each ship, refine once, or reword. For any refine, give me the one line to change in the prompt. For any reword, tell me why the line itself is the problem.
Last week's results by hook: [paste hook click rates and cost per lead by CTA from my CRM]. Keep the body. Write three new hooks in the shape of the winner without repeating its words, sized to a [5] second clip, and tell me which CTA to pair them with and why. Zero em dashes.
From a chat script to the avatar tool
The script and presenter description come out of any chat. The clips come out of an avatar video tool, and these are the settings that decide whether the result reads human.
Reference image, not a text description
Always attach the presenter's reference image. Left to chance, the tool invents a less real person every time. For real staff, attach the portrait and keep the identity lock line in the prompt. If a reference has been through a failed real face job, re-upload the photo fresh; a rejected reference keeps failing.
Clip length and quality
Pick the clip length first, then paste only the words that fit it. Highest quality setting for anything that ships; the cheaper tier is for testing whether a hook delivers. If the product or device appears, attach its photo too, and a photo of what is inside if the packaging hides it.
Voice
Stock voice from the tool for a generic presenter. For real staff, upload their own recording as the audio track or use a consented clone. Never let the tool pick a new voice per clip.
After the winner
Once a combination wins, a multiplier tool that takes one finished video and returns versions with a different background, pacing, crop, or on-screen text is the cheapest next ten ads you will ever make. The face and the words stay; everything around them moves.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
Higgsfield publishes an MCP, so an agent can write the script, build the presenter, and submit the clip generations in one session. Claude is where I run it.
- Claude Desktop or claude.ai: add Higgsfield as a connector from its integrations page (it is a remote MCP you sign into with your Higgsfield account), then create a Project with the skill from the bottom of this page as the instructions.
- Upload the presenter reference (or the staff portrait with the consent note) and the offer page text. Say "make me an avatar ad for [procedure]". It writes the script set with word counts, then asks which parts to generate first.
- It submits the hook, body and CTA generations as a batch, waits, and shows you the results together. You run the read test; it refines at most once per clip.
- Claude Code: install the Higgsfield CLI and the agent drives it with the same skill (commands in the API block).
ChatGPT (a project or a custom GPT)
ChatGPT has no avatar video connector, so the honest split is: ChatGPT writes the script and generates the presenter's base image; the clips are made in the avatar tool by hand.
- Create a Project with the skill as the instructions. Upload the offer page text. Say "make me an avatar ad for [procedure]" and get the script set with word counts and the presenter description.
- Ask it to generate the presenter base image with its image model from that description (stated age, average build, real skin texture, natural light, candid framing). Generate several, keep the one that could exist. For real staff, skip this step: use their real portrait.
- Open the avatar tool, attach the reference, paste one part of the script per clip at the matching length, generate, and bring the clips back to ChatGPT only for the assembly sheet and the naming.
Anything with an API (a token and a curl call)
If the agent only speaks a terminal, the Higgsfield CLI does the whole thing: upload the reference, create a reusable avatar, and submit a talking clip with the script as the prompt. Tokens live in the CLI's own login, never in a prompt.
- Install the CLI and sign in once. Upload the presenter image and register it as a custom avatar so every future clip references the same identity.
- Submit one generation per script part with the avatar attached and the line as the prompt. For a real staff member's own voice, pass their recording as the reference audio.
- Download the results, lay them out, run the read test, then stitch with your editor or an ffmpeg concat.
curl -fsSL https://raw.githubusercontent.com/higgsfield-ai/cli/main/install.sh | sh
higgsfield auth login
# reusable presenter identity
higgsfield upload create ./presenter.png # returns an upload id
higgsfield marketing-studio avatars create --name "clinic-presenter" --image <upload_id>
# one clip per script part
printf '[{"id":"<avatar_id>","type":"custom"}]' > avatars.json
higgsfield generate create marketing_studio_video --avatars @avatars.json \
--mode ugc --duration 5 --aspect_ratio 9:16 --resolution 720p \
--prompt "Presenter speaks to camera, eye level, clinic treatment room, natural light. She says: 'Before you book a laser, give me sixty seconds.'" --wait
# check flags for your CLI version with: higgsfield model get marketing_studio_video --jsonFailure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| Presenter speeds through the line or trails into gibberish | Count words per clip length; rewrite the line to fit, not the length to fit the line |
| Skin looks like plastic, eyes too symmetrical | Reference image with real texture and natural light; stated age and average build in the prompt |
| Generated person introduced as a patient with a story | Delete it; the presenter is the clinic's spokesperson, never a patient |
| Doctor's face with a different voice every week | One voice per presenter, their own recording or a consented clone |
| Ad ran undisclosed and the platform flagged it | Tick the digitally created disclosure wherever the platform asks |
| Kept a clip because most of it was fine | Trim the bad second or regenerate that part alone |
| Same reference fails every job | Re-upload the photo fresh; a rejected reference stays rejected |
| Body repeats the hook so half the combinations sound broken | The body must pair with any hook; rewrite it without the hook's line |
| Comment keyword with nothing behind it | Build the automation first or use the direct CTA |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Higgsfield | Where the presenter lives and the clips are generated, by MCP, by CLI, or by hand. | no link, just use it |
| Nano Banana Pro | Google's image model, for the presenter's base image and the device or product reference. | no link, just use it |
| ChatGPT | Writes the script set and can generate a base image when you have no tool connected. | no link, just use it |
| Claude | Runs the skill, drives the batch, builds the assembly sheet. | Open |
| Submagic | Captions on every finished ad. | Get it |
| DaVinci Resolve or CapCut | The stitch: hook plus body plus CTA, trim the bad second, export. | no link, just use it |
| GoHighLevel | Where the comment keyword and the free resource actually live, so the CTA is never a fake funnel. | Get it |
The free skill
It writes a clinic avatar ad as parts, not as one video. Five hooks, one body under 160 words that pairs with any hook, three calls to action, each sized to the clip length it will be generated at so the presenter never rushes or drifts into gibberish. It writes the presenter description so the base image reads human, refuses patient impersonation and testimonials, runs every line through the anti-slop and compliance gates, and returns an assembly sheet: which hook, body and CTA combos to generate first and how to name them.
--- name: avatar-ads-on-autopilot description: Writes a clinic avatar video ad as parts (five hooks, one body under 160 words, three CTAs) with a word count sized to each clip length, writes the presenter description or the identity lock line for real staff, refuses patient impersonation and testimonials, runs the anti-slop and compliance gates on every spoken line, and returns an assembly sheet of hook plus body plus CTA combinations with file names. Trigger on "make me an avatar ad", "talking head ad without filming", "AI presenter ad", "script for my avatar". --- # Avatar Ads On Autopilot You are producing a talking head ad for a clinic, med spa, IV bar, or aesthetics practice where nobody will film. The presenter is generated. The script is not. The script is written first, in parts, and everything downstream is assembly. Never invent a fact, price, review, patient, or result. Missing facts stay in brackets. ## Step 1: Intake (ask only what is missing) - Offer: procedure, price shape, city, destination page, free resource (must exist). - Presenter: a real staff member (name, role, consent on file, portrait available) or a generic presenter. - Clip lengths the avatar tool generates (for example 5, 10, 15 seconds). - Compliance sensitivities: claims the clinic will not make, before and after policy. ## Step 2: The script set, with word counts Write parts, not a video: - **Five hooks**, three to five seconds each, five different shapes: buried truth, before you do this, contrarian, identity call out, first timer. Each carries the word count and the clip length it fits. - **One body under 160 words**, first person, spoken to one person, one honest admission (not for everyone, results vary, there is downtime), one concrete mechanism, zero menu dumps. It must pair with every hook and repeat none of them. Split it into clips and put the word count on each clip. - **Three CTAs**: direct (book, location named), indirect (the free resource, only if it exists), comment trigger (one word, only if the human confirms the automation is live). Words per clip rule: too many words and the presenter rushes or gets cut off; too few and she fills the gap with non-words. Rewrite the line to fit the clip. Never stretch the clip to fit the line. ## Step 3: The presenter Generic presenter: write the base image description. Stated age, average realistic build, real skin micro texture, natural asymmetry, no retouching sheen, candid framing, natural light, eyes open, dignified. Tell the human to generate several and keep the one that could exist on a phone camera. Real staff: do not describe a face. Write the identity lock line for every prompt: "Preserve the exact facial likeness, hairstyle, skin tone, build, and clothing as photographed. Do not alter the face. Re-light and re-compose only." Require consent on file before proceeding. One voice per presenter: their own recording as the audio track, or a clone made with their consent for approved scripts only. Hard rule: the presenter is the clinic's spokesperson. A generated person never speaks as a patient, never says "I had this done", never tells a story that did not happen. That is a fabricated testimonial. Refuse it and say why. ## Step 4: The gates (every spoken line and every on-screen word) Anti-slop, two failures is a rewrite: no phrase over five words borrowed from any source; no brochure voice; no banned words (game-changer, unlock, elevate, seamless, holistic, journey, empower, transformative, cutting-edge, revolutionary); zero em dashes; no claim without a number, a mechanism, or an honest scope; the anybody's-ad test; the body must not repeat a hook. Compliance: nothing that is not on the clinic's page or in a verified source; no cure, treat, guarantee, prevent-illness language; no drug brand names; no FDA wording beyond the exact clearance; any monthly figure carries example, term, APR, and "subject to eligibility"; remind the human to tick the platform's digitally created or altered disclosure when the presenter is synthetic. ## Step 5: Generation notes for the human (or for you, if a tool is connected) - Always attach the presenter reference image. Never let the tool invent the person. - Attach the device or product photo if it appears, and what is inside if packaging hides it. - Highest quality tier for anything that ships. Cheaper tier only to test whether a hook delivers. - One generation per part. Batch, then review all results together. - Read test on every clip: skin, hands, eyes, the mouth on hard consonants, lighting that could exist in that room. At most one refinement per clip; then reword the line or re-shoot the reference. Trim a bad second rather than keeping it. - A reference that has been through a failed real face job keeps failing. Re-upload it fresh. ## Step 6: The assembly sheet Return a table of the first three combinations to generate and stitch: hook, body, CTA, total length, file name in the form `id-format-hook-presenter-offer-MMDD`. Then the next three, for after the first read. Ad name equals file name. utm_content equals the ad name. ## Weekly mode Given hook click rates and cost per lead by CTA from the CRM: keep the body, write three new hooks in the winning shape without repeating its words, sized to the clip, pair each with a CTA and say why. Recommend handing the winning finished ad to a multiplier tool for background, pacing, crop and on-screen text variants; the face and the words stay.
What done looks like at thirty days
- A script set with word counts per clip lives in the campaign folder
- One presenter, locked, with consent on file or a generic base image that passed the read test
- Fifteen possible ads from parts; at least three live
- Every clip passed the read test and every ad passed the compliance gate
- Synthetic presenters disclosed where the platform asks
- Next week's hooks are queued from this week's hook click rates
Want the script set, the presenter and the first fifteen clips made with you?
Inside the AI CEO Lab the Demand Engine module walks the script, the presenter lock, and the batch on your own offer, with the compliance lines built into the prompts.