Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Originally published by @EXM7777 on X. Tech Twitter preserves the original source alongside this readable edition.
Before you dive in
• This matters because it reveals how to systematically produce AI-generated user-created content (UGC) ads at scale by treating ad creation as an engineering problem—starting with customer research and copywriting before any generation, then using Claude Code to orchestrate multiple AI models (image, video, audio) through a parallel pipeline that treats identity consistency and script accuracy as measurable constraints
• For builders and founders, this directly addresses the gap between raw AI capability and commercially viable output: the framework shows that winning ads are determined by vault-sourced persuasion patterns and word-counted scripts rather than generation quality, meaning someone without design skills can compete with human-created ads using proper workflow architecture and platform orchestration through Higgsfield's supercomputer interface.
Best for builders who want practical takeaways. 21 min read.
I will teach you how to go from complete beginner in AI UGC to owning an agent that produces finished ads on autopilot. I'm going to hand you the exact prompts, the tools to use, the skills to build, and the full pipeline from research to finished media... plus how to wire all of it into higgsfield supercomputer for a loop that runs without you
UGC meaning the ads that look like a customer filmed them on their phone, because that's the format people actually watch... and every video in this piece came out of this exact pipeline: invented brands, generated presenters, real category language
you'll also get the generation laws that only show up after real failed renders: the sixth finger, the duplicate tube, the laces that tie themselves, the label that melts into noise
and hold onto one thing before any tool gets named: the words sell the ad, the render only has to stay out of their way
consumer testing run this year put human ads and AI ads head to head on the same strategy briefs, and the human ads still won
only about a quarter of viewers could even tell the AI ads were AI... so what separates the two is craft, and the craft lives in the words
which means the factory starts at customer language and ends at the render, never the other way around
here's what's inside:
the engine: one platform, two doors
the research layer: finding UGC that actually sells
the vault: hundreds of winning scripts in one obsidian folder
script math: the copywriting layer that converts
the avatar: one presenter that survives every clip
the generation laws: what the model can and cannot do
how i prompt Seedance, clip by clip
the factory run: hook first, then everything in parallel
assembly, and the gate that catches what you miss
if you want the full kit behind this piece, the prompt scaffolds, the vault template and the ready-to-run briefs, that's what the real time AI ops community at weeklyaiops.com is for
We doomscroll, you upskill
Get the 10 tweets shaping how builders think today.
Newsletter
We doomscroll, you upskill
Get the 10 tweets shaping how builders think today.
a single UGC ad touches an image model for the product shot, another for the character sheet, a video model for the clips, and audio on every one of them, which normally means four subscriptions and four render queues
higgsfield puts all of it on one surface behind one login, and the platform is built for agentic use, which is what a factory needs
you have two options: supercomputer, the easy way, an agent that runs on the platform itself...
or the build-it-yourself way, where every higgsfield model is addressable straight from Claude Code and your own agent holds the pipeline, what gets generated, when, with which model, in what order
supercomputer deserves a longer look, because UGC is exactly the workload it was built for
a UGC ad needs completely different kinds of agent work in one pipeline: researching the product and the niche, writing the script, generating the product image and the character, turning stills into video clips, then cutting, clipping and assembling the montage...
normally that's five tools and five handoffs, and every handoff is a place to lose quality
supercomputer does all of it in one place without you setting anything up:
you describe the product, it runs the research, writes, generates the image, renders the clips, does the montage, and refines the result until the UGC is coherent... and planning costs nothing until you render, so exploring is free
the rest of this article is the build-it-yourself way, built as a factory
but the factory doesn't start with a render
the research layer: find what sells before you generate anything
the first job is mining ads that already work, and i mean the selling argument, never the look
i got lucky enough to try the product from @eptwts, a research agent for creatives that isn't announced yet... you ask it for the top performing ads in a niche and it returns the videos, the creators, the exact hook lines and the script beats, and it carried a big part of my research for this build
without it, the manual route still works:
search the product category on TikTok Shop and sort by what's moving, then watch the top creators' recent ads
pull the ad libraries (TikTok's creative center, Meta's library) knowing their honest limit: they show what runs, not what converts, so everything you find is a candidate, never a proven winner
save every ad with real traction into the vault (next section) with its link... the link is the receipt, and a swipe file without receipts rots
the patterns below showed up in every niche i swept, and they're worth more than any prompt in this piece:
a named enemy: the winning sunscreen ads all attack white cast by name, the winning bra ads attack wires digging by six pm... the hook calls out the category's failure, and the viewer feels seen
live physical proof: the biggest bra ad proves the claim with a bend test on camera, the sunscreen winners blend the cream into skin in real time... showing does what telling can't
authority from exhaustion: “the only one i can tolerate” outsells expert credentials, because the presenter earned the verdict by suffering through the alternatives
one detail doing the believing: the silicone strip, the wide soft band, the front clasp... every winner anchors its claim to one physical detail the viewer can check
and the ordering rule that decides everything downstream: customer complaints first, competitor ads last
competitor research finds candidates and gaps... the direction comes from what buyers complain about in their own words, so mine reviews and comments before you ever open an ad library
when you do tear a winning ad down, split the teardown into two records
the persuasion record: hook family, beat timing, when the product enters, the proof device, which objection gets answered, the CTA
the capture record: device, framing, light, cuts
you inherit the persuasion record and re-decide every word... a teardown that only analyzes aesthetics, the wobble and the lighting and the mic, is measuring the half that doesn't sell
every record the teardowns produce needs one home, and the home is a folder of markdown files
the vault: hundreds of winning scripts in one folder
this section is the asset that grows with every ad you save, and it comes before any generation on purpose
open an obsidian vault (a folder of markdown files) and give every winning ad one note:
the link to the live ad
the transcript, word for word
the hook line, tagged with its family
the beat map with timing: hook, pain, mechanism, proof, CTA
the proof device it uses
the CTA wording
one line on why it wins
then three bank files the notes feed:
hooks.md, every hook you've captured, grouped by family
ctas.md, every close that felt natural
one map-of-content note per niche linking its ads
the reason this comes early: your scripts eat this vault
a script written from ten captured winners outships the blank page, because it starts where the niche already landed, and the wikilinks show you that the white cast callout and the wires-digging callout are the same move wearing different categories, so you can carry a proven hook shape into a niche that hasn't seen it yet
that transfer is why you store hundreds of these... and the vault only pays out when the words get measured
script math: the copywriting layer
the words are the pacing, because the video model stretches the spoken line to fill whatever clip length you request
so the math is fixed before writing:
talking head speech runs about 3.5 words per second
a 30 second ad is roughly 105 words, give or take ten percent... counted, never guessed, because the count IS the delivery speed
each generated clip stays under 9 seconds, so a 30 second ad is four or five clips stitched
each clip's line is its duration times 3.5... a 6 second clip carries about 21 words
overshoot the budget and the delivery rushes, undershoot and it drags... count the words before you render anything
hook families, from the sweeps and the vault: open with an offer, a confession, or a call-out
“i own six push-up bras and i hate five of them” is a confession hook from my own build, written from real category complaints... a question hook in the same slot reads weaker because it delays the claim
the remaining script laws:
one message per ad... a script carrying two beliefs carries none
every ad ends on the CTA, nothing after it
if the voice garbles the brand name, respell it in the script the way it should sound
write like the presenter talks: contractions, connective tissue, sentences you'd say to a friend... read every line out loud before it renders, because the model will speak exactly what you wrote
the script decides what gets said... now we build the person who says it
the avatar: one presenter that survives every clip
a text description is not identity memory
describe the same woman in two prompts and you get two different women, because each generation starts from scratch... the fix is a character sheet the model can see
my process, and then the exact prompts:
generate a headshot with a photoreal portrait model: age, skin, hair, one closed wardrobe line, phone-photo look
generate the fullbody FROM that headshot as a reference, same outfit
those two exact images ride every single clip as references, never re-cropped, never regenerated
here's a headshot prompt that produces a presenter you can hold for a whole campaign, word for word:
casual iPhone-quality selfie-style headshot portrait of a woman in her late 20s, warm chestnut brown hair in a loose low bun with face-framing strands, hazel eyes, light olive skin with a few freckles, friendly wry expression, minimal natural makeup. wearing a heather-grey ribbed tank top, small gold stud earrings, no other jewellery. soft indoor daylight, plain bedroom wall behind, slightly imperfect framing, authentic phone-photo look, not studio, not glossy
every phrase in there has a job, and this is where realistic characters are actually won:
“iPhone-quality selfie-style” sets the sensor... the model knows what phone photos look like, so name the device class, never “photorealistic”
the freckles are an identity anchor... small fixed imperfections survive across generations and make drift instantly visible, so give every character one (freckles, a beauty mark, a specific brow shape)
“face-framing strands” is the hair rule... perfect hair reads rendered, loose strands read human, and the video model animates them for you later
“slightly imperfect framing” kills the studio look in one phrase
“minimal natural makeup” because skin needs texture... glossy flawless skin is the single loudest AI tell in UGC
one closed wardrobe line, because this outfit now exists in every clip
then the fullbody, generated with Nano Banana using the headshot as its reference:
full-body casual photo of the SAME woman as the reference image, identical face, hair and outfit: late 20s, chestnut brown hair in a loose low bun, heather-grey ribbed tank top, relaxed black lounge joggers, barefoot, small gold stud earrings, no other jewellery. standing relaxed in a cozy bedroom, soft daylight, phone-photo realism, full figure visible head to feet, natural posture
“the SAME woman as the reference image, identical face” is the line doing the work... and keep the outfit identical between the two images, because the video model averages whatever it sees
Nano Banana is also how you make character variants when the story needs two states of the same person, a before and an after... generate state one first, then feed it back with “the SAME character as the reference image, identical face, eyes, nose and shirt, but now with a FULL head of thick wavy hair”... same character, two states, and the cut between them is the whole ad
a sheet built this way holds one presenter through five separate generations across five different rooms... same face, same freckles, same tank top
two casting lessons the build forced:
casting is proof: for the sunscreen ad the claim was “you can't even see it on skin”, so the presenter has deep skin, because white cast reads worst there... who is on camera proves what the prompt can't
faces lock both ways: an avatar's face styling comes from its reference photo, and no prompt overrides it... you can't add makeup to a bare-faced sheet and you can't remove it from a made-up one, so if the face needs a specific look, build the sheet with that look already in it
the product gets the same treatment: its own still, generated once, attached to every clip where it appears, described in the same words in every prompt, with its real size stated
a product prompt that renders a clean label on the first try:
product photography of a matte soft-white plastic sunscreen tube standing upright, flip cap down. label design: warm orange wordmark 'NOON' in clean modern sans-serif, below it smaller text 'daily SPF 50' and 'invisible finish · 50 ml'. minimal skincare aesthetic, soft daylight studio lighting, pale warm background, gentle shadow. centered, whole tube visible
the two rules hiding in it: name the exact label text in quotes with its font style, or the model invents typography... and describe the object in material words (matte soft-white plastic, flip cap, gentle shadow), never in brand adjectives
and print the brand name LARGE on the product... text is the most fragile thing in any frame... a small fabric tag will garble four different ways across a campaign while a bold tube label reads clean in every clip
the sheets keep everyone the same from clip to clip... the laws below keep the scene from breaking around them
the generation laws
everything in this section came out of a failed render first, which is why it's worth something
the model can hold anything and change nothing
it holds a tube flawlessly for fifteen seconds... ask it to pick the tube up and you get a duplicate tube, a wrong-side hand, a sixth finger, or the object teleporting between frames
it cannot fasten a button, tie a lace, uncap a bottle, or push a foot into a shoe... push it and you get a sleeve that dresses itself, a knot that ties itself between frames, a cap that never comes off, or a third disembodied foot
and it cannot show a transformation either: a cream swatch being rubbed in either sits frozen like a sticker or vanishes between two frames
the law that survives all of it: every state change happens across a cut
shot one, the swatch sits on her hand... cut... shot two, half blended... cut... shot three, clean skin
each clip is one held state, no clip contains the transition, and the ad still shows the full demo, because the viewer's brain supplies the change the way it has in every edited film since editing existed
the supporting laws, each from a real failure:
objects start in hand: never script a grasp, the reach-and-lift is where anatomy breaks
count objects scene-wide: “one tube” isn't enough, the system will happily stage a second branded tube on a background shelf... write “exactly one tube in the entire scene, no duplicate on any surface, no reflection copies”
close the wardrobe: whatever you leave unstated gets invented and locked... leave it open and the system adds a silver ring nobody asked for, so the lock reads “small gold studs, no other jewellery of any kind, no rings, no watch”
one action per clip: dense ordered action lists break everything, text first... one scene, one continuous action, one emotion
the viewer is the camera: name the filming device and the model renders a phone into the shot... describe the framing instead
faces anchor, products drift: identity holds across clips with a good sheet, product details morph between them... repeat the product description verbatim in every prompt like you repeat the character line
one last layer sits under all of these: the technical register
end every clip prompt with a camera-and-sound block: “iPhone front-facing 23mm equivalent, gentle handheld drift, built-in mic, no music”... that block is the difference between footage that reads like a propped phone and footage that reads like a produced video, and it's the layer almost everyone forgets to write
with the look and the laws handled, a clip comes down to a hundred words of direction
how i prompt Seedance, clip by clip
Seedance is the video engine i run every clip through, because it takes image references for identity, an audio reference for the voice, and generates the speech natively in the same pass
the clip prompt is a fixed skeleton, under 120 words... here's the shape, filled in for a bra ad:
a realistic, authentic UGC ad, handheld front-camera selfie video, 9:16. keep the character consistent with the reference images. her voice matches the reference audio exactly
standing in front of her open closet with a mirror, holding the blush-pink bra in one hand. delivery: fed up, venting to a friend. she says exactly: “they all do the same trade. great shape for the first hour, and by six pm the wires are digging and you're counting minutes till home”
the bra stays in her hand as an object the whole clip, never worn. she is fully dressed in her grey ribbed tank top. no on-screen text or captions
the skeleton is always the same five parts: one register line, one scene line, one delivery word, the exact quote, the hard rules... the reference images carry the look, so a long prompt only fights them
the details inside that skeleton decide whether the clip feels alive:
“she says exactly:” with the line in quotes... the model speaks it word for word with lip sync, so the script you wrote is the script you hear
one delivery word, not an adjective stack... “delivery: fed up” produces a performance, “natural authentic relatable energy” produces nothing
give the eyes a target... eyes that hold the lens for eight seconds read dead, so write the glance in: “around second 8 she pauses, glances down at her hand, gives a small half-laugh, then looks back to camera”... a look away and back is the single cheapest realism upgrade i know
write the micro-behaviors, never the mood... a pre-speech breath, a half-smile that settles, fingers adjusting grip... small involuntary movement is what reads human, and you get it by naming it, never by asking for “realistic”
one physics cue per clip... necklaces swaying with the motion, strands moving as she turns, and the frame stops feeling still
whole seconds only, 4 to 9, and i live at 5 to 8... the model paces the line to the duration, so the word count and the seconds have to agree
“holds the final pose to the last frame” on the closing clip... it gives the stitch a clean out-point instead of a mid-gesture chop
lived-in sets... “slightly cluttered counter, everyday products in the background, not staged”... tidy rooms read like renders, mess reads like a tuesday
for voices: native audio with the anchor flow gives you one consistent voice across every clip for nothing extra... if you want to design the voice instead, generate the hook line in ElevenLabs first and that file becomes the anchor, same flow, one extra step
one clip prompted well is a sample... the run below turns it into an ad
the factory run: hook first, then everything in parallel
this is where loop engineering earns its name... a clip is never one generation, it's a loop: generate, watch, change one variable, regenerate, until the clip earns its place, and the factory is that loop nested inside a bigger one that walks the whole shot list while you do something else
and the pipeline itself is a graph, not a line... the character sheet feeds every clip, the hook's audio feeds every later clip, the vault feeds every script, so draw your system as nodes and edges and you know exactly what one re-roll invalidates and what survives it... that's graph engineering doing the quiet work under the factory
here is the wiring for one 30 second ad:
render the HOOK clip alone, before anything else
watch it and judge the face and the voice... if either feels off, re-roll now, because everything else inherits this clip, and a bad identity caught here costs one clip instead of the whole batch
extract the approved hook's audio and attach it as the audio reference on every remaining clip, so one voice carries the ad... if you'd rather design the voice first, generate the line with ElevenLabs and use that file as the anchor from clip one instead
fire the remaining clips in parallel, every one carrying the same character sheet, the same product still, the same voice anchor
give every beat a different room or framing... adjacent clips in one location read as a broken jump, a scene change per beat reads as the native multi-location UGC style
respect the concurrency cap: the video backend runs a few jobs at once and rejects the rest, so the factory submits when a slot frees, which is correct... the cap is what makes an unattended run predictable
Claude Code drives every generation and stitches the result with ffmpeg, the free command line video tool it already knows how to use: concatenate the clips in order, normalize the loudness, done... the edit is one command you can rerun in seconds
and if you read this list and want none of the driving, this is the loop supercomputer runs for you on the platform: same research, same anchors, same refine cycle, hands off
either way the render is now holding up its end... whether the ad sells was decided back in the vault, which is the thesis of this whole piece working in both directions
assembly and the gate
before your eye judges anything, a machine check catches what eyes skip... i call mine the frame check
my check: pull sixteen evenly spaced frames from the finished ad and hand them to a vision model with a fixed checklist... limbs and finger counts, object permanence, label text read frame by frame, background continuity, filming device leaks
a check like this catches what full-speed watching misses: a prop merged into the product, a background that changes mid-take, a label that reads fine at speed but garbles on every third frame
two rules for reading its verdicts:
the frame check is cut-blind... it flags every intentional cut as a continuity break, because it judges neighboring frames, so read its report with your shot list open and dismiss the flags that are your own edits
the frame check measures broken, never good... the clips it scores worst are sometimes the best ads, because clean takes are often clean because the action got stripped out of them... the machine protects you from defects, your eye is still the only judge of whether it sells
one more thing before you publish anywhere: platforms require AI-generated ads to carry their AI label, TikTok rejects significantly generated ad media without it... disclosure is a checkbox in the upload flow... tick it and move on
the playbook
mine winners into the vault: link, transcript, hook family, beat map, CTA
write from complaints: named enemy hook, one message, 3.5 words per second, end on the CTA
build the sheets: headshot, fullbody from it, product still with a big label
render the hook alone, approve it, anchor the rest with its audio
one held state per clip... every change across a cut
parallel render, ffmpeg stitch, the frame check, your eye ships it
the first full run fits in an evening, and most of that is render wait
what the factory is actually for
nothing in this pipeline knows what product it's selling... the same loop covers skincare, supplements, grooming, apparel, jewelry, and even claymation or podcast formats, because the factory is the asset and each ad is just a run
and if you don't have time for any of this... if you just want good ads out of the batch without spending hours on research and tweaking, upload your product into Higgsfield Supercomputer and it does the whole thing for you: the research, the script, the images, the clips, the montage, refined until it's coherent
the models in this piece are available to everyone reading it
the vault is the part you build by hand, one winning ad at a time