Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Originally published by @EXM7777 on X. Tech Twitter preserves the original source alongside this readable edition.
Before you dive in
• Seedance 2.5's doubling of clip length to 30 seconds and tripling of reference capacity (up to 30 images, 10 video clips, 10 audio clips per generation) enables builders to produce full commercial content in single passes, but requires disciplined prompt structure across four timed beats and explicit reference role assignments to avoid quality degradation
• The Higgsfield CLI integration transforms this into an agentic workflow where agents can spawn and stitch multiple 30-second jobs into complete ads or films autonomously, making AI video production programmable rather than manual, which fundamentally changes how founders can scale content creation and how AI practitioners can build reliable video generation pipelines.
Best for builders who want practical takeaways. 13 min read.
Seedance 2.5 is coming soon on Higgsfield and I'm going to show you exactly how to master this model to produce ads, movies, vlogs, UGC... anything you could ever think of.
this new release can one-shot an entire 30 second video from any prompt you throw at it
that's up from 15 seconds on the last version, and now the audio and the video generate together in the same pass instead of getting stitched on afterward
it also carries a much bigger reference budget into that one shot, up to 30 images, 10 video clips, and 10 audio clips, each with its own job, so a single generation can hold the cast, the location, a key prop, the camera motion, a voice, the ambience, and the music, all at once
if you want to learn how to get the most out of ai video and turn it into revenue, that's covered inside my community: weeklyaiops.com
here's what you're getting in this piece:
examples of 5 visual formats you can achieve
step one: build your reference bible in obsidian
step two: review your bible after every session
the one idea the rest of this guide is built on
step three: write every shot with the same six details
step four: break your 30 seconds into four beats
step five: tell every reference exactly what it's for
step six: build your reference images
your final build
the platform where your workflow runs
higgsfield already puts every ai video model on one surface, one login, one credit pool, built for agentic use, not just clicking a generate button by hand
the cli is what turns it agentic: it makes every model callable straight from your own agent, so claude code or whatever harness you run can hold the entire pipeline, what gets generated, in what order, with which model
three commands put the cli in your terminal:
`npm install -g @higgsfield/cli`
We doomscroll, you upskill
Get the 10 tweets shaping how builders think today.
Newsletter
We doomscroll, you upskill
Get the 10 tweets shaping how builders think today.
the same cli that runs seedance 2.0 today is what seedance 2.5 lands on the moment access opens
so nothing about your setup changes, one agent that already knows how to submit and retry a job just starts spawning a fleet of 30 second jobs, one per shot, stitched into a full ad, a launch video, or a short film
[visual: 1-higgsfield-doors-gpt.png: the cli as the agentic door, install commands to a fleet of 30-second jobs]
old habits from 2.0 cost you time here, so build the reference bible below first
try it for a KPOP music video
aim for tight editing on the beat, hard cuts landing exactly on the snare or the drop, never soft crossfades, a cut lands on the beat or it doesn't land at all
build the whole scene around symmetrical, center-weighted framing, recurring archways or circular portals the camera can dolly backward through while staying eye-level with the lead performer
color grade toward one bold, saturated palette locked across every single scene, block colors instead of naturalistic tones, and keep skin tones neutral-to-warm so the saturation pops without casting on faces
lip sync holds up best on medium close-ups with clear, simple phoneme shapes, never whispered or mumbled delivery, so write the vocal delivery directly into the prompt: crisp diction, open vowels, staccato phrasing
pair one dance move to one cut, not a continuous take, since the sync between motion and the beat is what makes the whole clip read as choreographed instead of accidental
try it for a vlog
skip tripod language entirely and describe the shake instead: handheld micro-jitter, arm's-length selfie framing, natural body sway as she walks
light it with whatever's already there, golden hour backlight outdoors, cool overhead fluorescent underground, never a described studio setup, and let the light quality change with the location
cut on movement, not on a fixed rhythm, a turn, a step off a curb, a hand reaching for a door, and drop in quick b-roll cutaways between the talking segments
grade warm and soft outdoors, filmic contrast with real shadow detail, and let interiors go cooler and flatter for contrast against the exterior shots
the giveaway details sell the format harder than any camera move: a visible mic cable, hair blowing across the face, a direct look at the lens like she's talking to a friend, not performing for a crowd
try it for a product shot with 3d elements
describe materials by their physical property, not their name, glossy polycarbonate with internal light scatter reads completely differently to the model than “shiny plastic,” and matte PBT-style plastic wants high diffuse roughness called out directly
separate the camera moves: a slow 360-degree orbit for texture and material read, a straight pull-back dolly for an exploded-parts reveal, never blend both moves into the same shot
light it soft and bright, large diffused softbox sources so shadows stay minimal, with a thin rim light tracing every translucent edge and a gradient specular sweep crossing the surface as the camera moves
keep the background an infinite studio backdrop, a soft pastel gradient with gentle ambient occlusion under anything floating, no visible edges, no real-world environment competing for attention
grade toward pastel, lifted blacks, no blown-out bright spots, the clean, modern look that reads as lifestyle-hardware branding rather than a generic render
try it for realistic lighting
name the light's direction and its color temperature separately, a warm key light against a cool sky-fill is what reads as real, a single flat light never does
use hard directional light for crisp specular hits and sharp shadow edges up close, then let atmospheric haze soften everything at distance, that shift from sharp to soft with distance is doing a lot of the realism work
put the weight on edges: a rim light catching flyaway hair, a lens flare blooming exactly when the camera tilts toward the sun, bright spots that roll off smoothly instead of clipping into a flat white
let shadows stay imperfect, picking up bounced color from whatever's nearby, cool blue from open sky, warm tones from stone or earth, rather than one flat, uniform fill tone
and let background layers lose contrast and saturation the farther they sit from camera, that gradual fade is atmospheric perspective, and skipping it is the fastest way to make a real-looking scene look synthetic again
try it for animation
pick one reference era and commit fully, never blend three animation styles into a single prompt, name the exact one: cel-shaded, painterly-over-3d, or classic 2d
give different elements different frame rates, on purpose: the character animates on twos for a snappy, hand-drawn feel, while the vehicle or the camera itself moves in smooth, continuous motion, that contrast is what reads as an intentional style choice instead of a technical error
mix your camera language deliberately: extreme close-ups on hands and small mechanical details, smooth tracking shots following the action, dramatic low-angle shots for whatever moment needs to feel largest
shade it painterly, visible brushstroke texture over the 3d geometry, hand-painted bright spots instead of a photoreal shader, that's what pushes it from “3d render” to “illustrated world”
commit to one color-palette family across every single scene, neon-on-midnight, sun-bleached pastel, whatever it is, and let that one choice carry the whole clip's mood instead of resetting it scene to scene
step one: build your reference bible in obsidian
think of it as a folder, not a notebook
for every world or character you build, save one page each for: the idea behind it, the colors and style you locked in, the character sheets, the reference images, the exact prompts that worked
keep them all in one place, linked together, the same way you'd keep client files or a project folder organized
one generation can pull in 30 images, 10 video clips, and 10 audio clips
that's too many to remember, so it has to be something you look up, and a simple folder with one page per item and one index page pointing to all of them is the only setup that still works once a project runs past a week
step two: review your bible after every session
do this or the folder turns into a pile of dead files nobody opens again
after each time you generate something: ask Claude to write one short note on what worked, update your index page to point to it
that's it
skip this step and you lose the whole point of building the bible in the first place
do it, and every project after your first one moves faster, because half the work is already done and sitting there ready to reuse
the one idea the rest of this guide is built on
seedance 2.5 got better in two specific ways: the single clip length doubled, from 15 seconds to 30, and the number of references you can use more than tripled
here's the part that trips people up: more room to work with does not mean you can be lazier
it means there are more places to mess up without noticing
bytedance's own guide says the range that actually holds steady is 1 to 8 images, and 1 to 5 video or audio clips, not the bigger numbers on the box
why? because the more references you throw in, the less stable the result gets
so the real number to remember is small, not big
here's what happens when people ignore this...
they expect the model to just refuse a 30 second request if it can't handle it
instead, it tries anyway
stretch an old 15 second prompt to 30 seconds without breaking it into parts, and you get actions that make no sense, props that appear out of nowhere, and characters that slowly stop looking like themselves, only visible after you've already waited for the render
so remember these two things together
the model can now hold a longer story together on its own
but you still have to give it the same level of detail per shot that you always did, just four times over instead of once
everything below is how you actually do that
step three: write every shot with the same six details
the basic prompt formula from the older model still works here
for every shot, state: what's in it, what it's doing, where it is, how the camera moves, what style it's in, and any rules to follow
what changed is not this list
what changed is how many times you have to use it
step four: break your 30 seconds into four beats
one prompt now covers a 30 second story, so split it into four timed chunks
here's the map bytedance itself uses:
seconds 0 to 6: set the scene
seconds 6 to 14: build it out
seconds 14 to 24: the turn, the big moment
seconds 24 to 30: how it ends
write out the six details above for each of those four chunks, inside the same prompt, with the timestamps included
skip this step and paste one long paragraph for all 30 seconds instead, and the model has no idea where your story is supposed to turn
step five: tell every reference exactly what it's for
the old tagging system still works: @image `, @video 1, @audio 1, now with extras like @images 6 to 10 for groups, and @clay render 1 for a new kind of reference
but tags alone aren't enough anymore
each one now needs two things stated: what it controls, and what it should not touch
for example, just saying “@video 1 defines motion, camera movement, and pacing” works, but it's not enough
add one more line: “do not use the person's identity, clothing, or scene from the video”
that second line is what stops one reference's details from leaking into a shot they were never meant to touch
one more new tool worth knowing: “@clay render 1”
it's a bare 3d shape with no texture, used only to lock in camera movement and blocking, while a separate image reference handles the actual look, the materials, the lighting
each one does one job
editing works the same way
“edit @video 1, keep the characters and visual style unchanged, adjust only the camera movement over 6 to 12 seconds” is a real prompt pattern that works today, not a guess
step six: build your reference images first, by hand
here's the part almost nobody's written down yet
the 30 image slots in your seedance prompt don't come from seedance at all
they come from a separate image model, made and picked by hand, before you ever write your seedance prompt
pick the tool by what the shot needs: midjourney is where you go for cinema shots, anime shots, anything that needs creativity and a distinct art style, nanobanana pro or gpt images 2 is where you go when you need realism, a real face, real materials, real light
here's exactly how, using midjourney as the example
first, lock one look for your whole project
either tune midjourney's style creator until you've made 20 to 22 picks and grab the `--sref` code it gives you, or roll `--sref random` a few times until 2 or 3 codes look right
second, generate the categories your project actually needs under that same locked style: your wide establishing shot, your environment, your main character, a key prop, your closing shot
each one gets its own short prompt
third, for anything that has to show up more than once, a character, a creature, a product, take your best result and turn it into a full reference sheet: front, side, back, blank background
this is what lets the same character or object drop into new scenes later without being redrawn from scratch
fourth, once you've locked your reference images, don't touch them again
if the motion looks wrong later, fix the motion prompt
don't regenerate the image, because a new image undoes all the consistency work you just did
the same four steps apply if you're building your set in nanobanana pro or gpt images 2 instead: lock the look first, generate by category, build reference sheets for anything recurring, then leave the images alone
do all four of these steps and you'll have a finished reference bible in about an afternoon, staying well inside the 1-to-8 range that actually works, instead of the 50-slot number that just sounds impressive
save this checklist
write the six details above for every single beat, not once for the whole clip
use the four-beat map as your default: 0-6, 6-14, 14-24, 24-30
give every reference a job, and state what it should not carry over
stay inside 1-8 images and 1-5 video or audio clips, well under the 50-slot ceiling
build your reference images by hand first, midjourney for creative/stylized shots, nanobanana pro or gpt images 2 for realism, using one locked style
once your reference images are locked, fix the motion prompt instead of remaking the image
keep it all in obsidian, one page per asset, reviewed after every session
skip the reference-role step and you'll find the leak in the one shot you didn't double check