Blog

Seedance 2.5:
Full Breakdown

The model, the workflow, and a complete worked example: turning a studio talking head into a man bursting out of a field.

Part 1 · The basics

Seedance 2.5 is a reference driven video model. You hand it assets you already have, tell it what job each one is doing, and it builds the result around those anchors. That is the whole idea, and every good habit below follows from it.

30 sec

Single generation, not stitched

720p / 24fps

10 bit color depth

50 refs

Up to 30 images, 10 video, 10 audio

Multi shot

Cuts inside one prompt

  • Audio is co generated with the picture, so dialogue, effects and music land in sync instead of being layered afterward.
  • Aspect ratios, run from 21:9 down to 9:16, plus an adaptive option that follows your input.
  • Four entry points:, text to video, image to video, reference to video, and first and last frame.
  • Editing without regenerating:, timestamped replacement of a window, plus extending a clip forward or backward.

Note. Specs on a pre release model move. Treat the numbers as the current shape of the tool, not a contract.

Part 2 · The workflow

  1. 1

    Decide what you are anchoring.

    What has to stay exactly as it is? A face, a logo, a camera move, a room's geometry. That comes from a reference, never from a description. Descriptions drift, references do not. Everything else is free to change.

  2. 2

    Give each reference one scoped job.

    This asset supplies the subject. That one supplies the style. This one supplies motion and timing. Then say what each reference should not contribute. Two assets competing for the same job is the single most common cause of a muddy result.

  3. 3

    Write in four parts: subject, action, camera, style.

    Not as labeled sections, as natural direction in that order. Subject names what is on screen. Action uses verbs in playback order. Camera picks one shot size and one movement. Style describes tangible choices, not adjectives like "cinematic".

  4. 4

    Block your beats on a timeline.

    For anything past a few seconds, write the events with timecodes and name the transitions between them. If you do not name the cut, the model picks, and it picks differently every time.

  5. 5

    Tie every sound to something visible.

    Vague audio prompts produce busy mixes. Say what changes and when, and say plainly when you want no music or no dialogue.

  6. 6

    Test in passes.

    Pass 1 at low quality proves the foundation. Pass 2 adds exactly one control. Pass 3 is the final render. Three cheap generations beat six expensive ones.

  7. 7

    Edit instead of regenerating.

    When a clip is 90 percent right, swap the bad window with a timestamped replacement. The version you already have is usually closer than the next roll of the dice.

Write measurable, not moody

The model responds to things it can measure. Speeds in km/h. Sizes in centimeters. Distances in meters. Field of view in degrees. Color temperature in Kelvin. "Fast, realistic dirt" gives it nothing to work with. "Clumps 6 to 8cm across thrown 2 meters toward camera" gives it everything.

Positive phrasing only

State the target, not the prohibition. Instead of "does not get dirt on his face", write "dirt lands on the shoulders and sweater only".

Part 3 · The worked example

The goal

Target resultLow angle, subject bursting up through the turf, soil spraying outward, wide sky behind.

The source material is a plain studio talking head and an empty field. The generation has to do three jobs at once: put him in the field, invent the burst, and keep his performance intact while relighting him from studio flat to midday sun.

The assets and their roles

The location plate. Low camera, horizon at the lower third, cirrus upper left, small cumulus along the horizon.
@image1The location plate. Low camera, horizon at the lower third, cirrus upper left, small cumulus along the horizon.
The performance. Cream backdrop, flat frontal studio light, tight framing. Only the person and the performance travel forward.
@video1The performance. Cream backdrop, flat frontal studio light, tight framing. Only the person and the performance travel forward.
SlotAssetScoped role
@video1Studio talking headIdentity, face, mouth movement, head motion, gaze, voice
@image1Empty field plateLocation, horizon height, sky, cloud pattern, daylight direction

Upload in this order so the tag numbering matches. If the platform assigns different tags, rename them in the prompt to whatever it shows.

The prompt

SCENE CONTEXT
An empty green grass field under a bright blue sky at midday. A young man
bursts up out of the ground through the grass, dirt flying, then settles and
speaks directly to camera for the rest of the clip.

ACTIVE REFERENCES
@video1 is the man. It supplies his identity, face, performance, mouth
movement, head motion, gaze and voice for the entire speaking section, frame
for frame. Late teens to mid twenties, short black curly hair, dark navy
square frame glasses, thin mustache, black ribbed knit crew neck sweater.
100% matches the reference. Take only the person and the performance from
@video1. The environment, framing and lighting come from @image1 and the
LIGHTING block below.

@image1 is the location. It supplies the grass field, the horizon height, the
sky, the cloud pattern and the daylight direction. 100% matches the reference.

LOCATION MAP
Foreground: mown green grass filling the lower third of frame, individual
blades sharp and readable within 1 meter of the lens. Midground: the man's
emergence point, centered, roughly 2 meters from camera. Background: the grass
plane rolls to a low horizon at the lower third, above it an open blue sky with
thin cirrus streaks upper left and small scattered cumulus sitting along the
horizon line. Camera sits on the shadow side, low and forward on the grass
plane. Sun high and slightly behind camera left.

FIRST FRAME / BLOCKING
The first frame is the empty field exactly as in @image1: unbroken grass, low
horizon at the lower third, sky filling the upper two thirds. Three or four
birds are already gliding across the upper right sky, moving slowly camera
left. The ground is intact. Nobody is in frame yet.

FORMAT MODE
One continuous shot. The camera does not cut on its own.

OPTICS
63 degree FOV, rectilinear, held for the whole shot, no drift mid shot. Medium
shot once the man is up: head and shoulders occupying the center of frame with
wide sky readable around him. Natural motion blur on the flying soil, sharp on
the face.

CAMERA
Locked off at 25cm above the grass, tilted up 8 degrees, positioned 2 meters
back from the emergence point at eye level with the ground plane. Focus holds
on the emergence point from frame one and stays there. Wide tonal latitude,
gentle highlight roll off in the clouds, clean neutral skin rendering.

ACTION
0.0s to 1.0s: the empty field holds. Grass sways lightly. Birds glide in the
upper right sky.
1.0s to 1.4s: the turf splits and the man drives upward through the surface,
head and shoulders clearing the grass in one fast motion, rising at roughly
12 km/h. Clumps of soil and torn grass fly outward and upward in a spray, the
largest pieces 6 to 8cm across, thrown 1.5 to 2 meters up and out toward
camera.
1.4s to 2.6s: the soil arcs, tumbles and rains back down, the heaviest clumps
landing first, fine dirt drifting behind. It settles into a raised ring of
loose earth around his chest and shoulders. His body comes to rest, chest
rising once from the effort.
2.6s to end: he holds his position in the ground and speaks to camera,
delivering the performance from @video1 exactly. The camera stays locked. Birds
continue drifting across the sky behind him for the full duration.

PERFORMANCE
Mouth shapes, jaw movement, blink timing, eyebrow lifts, head tilts and eye
line follow @video1 frame for frame. Pore level skin detail, faint sheen on the
forehead and nose, live catch lights in both eyes, small capillary flush across
the cheeks. Micro pauses between phrases read as natural speech.

PHYSICS
Soil behaves with real mass: clumps hold together in flight, break on landing,
bounce once and stop. Fine dirt separates from heavy clumps mid air and falls
slower. Grass blades bend away from the emergence point and stay bent. Contact
shadows land under every clump on the grass. The raised earth around his
shoulders holds its shape once settled. The knit sweater carries loose soil in
its ribbing and a fine dust layer across the shoulders.

LIGHTING
Open midday daylight, 5600K, sun high at roughly 70 degrees elevation and
slightly behind camera left. Soft top frontal key on the face with a short soft
shadow under the jaw and the nose. Sky acts as a broad fill, lifting the shadow
side of his face. Bright open exposure, highlights holding detail in the white
cumulus, green grass reading saturated but natural under direct sun.

WARDROBE
Black ribbed knit crew neck sweater, dry, dusted with loose soil across the
shoulders and collar after the burst.

AUDIO
A muffled earth crack and a burst of soil at the moment he breaks through,
followed by clumps pattering onto grass. Light open field wind underneath. His
speaking voice from @video1, clean and forward in the mix. No music.

STYLE
Photoreal live action, sharp detailed footage, natural daylight color, fine
grain, real soil and grass physics, energetic in the burst and steady in the
delivery.

OUTPUT SETTINGS
720p, 24fps, real time throughout, 16:9.

POSITIVE LOCKS
The first frame is the intact empty field from @image1.
His face, glasses, hair and mustache stay identical to @video1 for the entire
clip.
His face and glasses stay clean of soil. Dirt lands on the shoulders and
sweater only.
His mouth movement matches @video1 exactly for every word.
Horizon height, cloud pattern and grass texture stay identical to @image1 from
first frame to last.
The camera stays locked at 63 degrees for the full duration.
Birds are visible in the sky in every second of the clip.
Daylight direction and 5600K white balance stay constant throughout.

Why each block is there

BlockThe failure it prevents
ACTIVE REFERENCESThe cream studio backdrop and tight framing bleeding in from the talking head clip. Scoping @video1 to the person only is the single highest impact line in the prompt.
LOCATION MAPHorizon height wandering between generations. Stating foreground, midground and background separately keeps the plate readable.
FIRST FRAMEThe model opening on him already half out of the ground, which kills the reveal.
OPTICS + CAMERAFraming drift and focus hunting. A stated FOV, height and tilt make the shot repeatable across passes.
ACTIONA rushed or mushy burst. Timecodes give the dirt somewhere to travel and time to land.
PERFORMANCELip sync sliding off, or the face going plastic once it is composited into daylight.
PHYSICSDirt that floats or dissolves instead of falling with weight.
LIGHTINGHim staying lit like a studio subject while standing in a sunlit field, which is what makes a composite read as fake.
POSITIVE LOCKSEverything above quietly drifting somewhere in the back half of a 30 second clip.

What changed from the first draft

Every reference got one scoped role

The original prompt described the event well but left the assets unscoped, so the model had to guess whether the cream backdrop, the flat lighting and the tight framing should carry over. That is the most likely source of drift. Now @video1 is performance only and @image1 owns the world.

The camera got locked

63 degree FOV, 25cm height, 8 degree up tilt, 2 meters back. Without this the horizon wanders between generations and the framing will not match the plate.

Lighting actively re keys the subject

The source clip is flat cream studio light. The LIGHTING block moves him to 5600K midday sun at 70 degrees elevation, slightly behind camera left, so his face matches the field instead of looking pasted on.

A face stays clean lock got added

Dirt flying prompts routinely end up putting soil across the glasses and face. The lock keeps it on the shoulders and sweater, which is also what actually happens when someone comes up through turf.

Physics got real numbers

Clump size in centimeters, throw distance in meters, rise speed in km/h. The model responds to measurable values, not to "fast" or "realistic".

The birds got consolidated

They now appear in the first frame and drift for the full duration, stated once in blocking and once in the locks rather than twice in the body. Saying a thing twice in the same register does not reinforce it, it just competes.

Bird variant to test on pass 2

To swap the calm few birds for a bigger flock, replace the two bird references with:

A flock of roughly twenty birds scatters upward off the grass at 1.0s as he breaks through, then settles into slow drifting flight across the sky for the rest of the clip.

This reads as more impactful but adds a second thing competing for attention at the exact moment of the burst. Worth testing as a single pass 2 change rather than baking it into the first generation.

Test order

  1. 1

    Pass 1, lowest quality.

    Check only that the burst happens at the right moment, the framing matches the plate, and his face holds. Ignore how it looks.

  2. 2

    Pass 2, one change.

    The bird variant, or burst timing, or soil volume. One only, so you learn what the model responded to.

  3. 3

    Pass 3, final settings.

    Full duration, final resolution.

If a generation comes back 90 percent right, use timestamped replacement on the bad window instead of regenerating the whole clip.