Every moving shot of her in the three films is Seedance 2.5: about 200 clips, conditioned on a Midjourney plate and, for singing, on the exact slice of the song. Two things made the difference between dead footage and a performance: prompts written like a script, and lip sync that is measured rather than eyeballed.
Kinds of clip
| Kind | Inputs | Used for |
|---|---|---|
| Plate shot | her plate as the first frame, the song mix for rhythm | most shots: she walks, turns, reaches, reacts |
| Singing window | her plate, a clean vocal of only her words in that window | lip sync, 4–8 s, one phrase |
| Green-screen performance | a plate of her on chroma green | cut-outs: BLISS's desktop dancers, the cursor fight |
| Cover | the last in-sync frame of another take | the rest of a line after a take drifts |
| Extension | a real clip as a video reference | joining generated shots to real footage, frame-continuous |
Write prompts like a script
CLODYSSEY's first round of prompts was written by planner agents and described poses: "she whispers", "he listens". Every take came back lifeless, and about 830 credits were gone before anyone read them. The director's verdict: "these prompts are really really bad and wasting money and need to be 10x better and super context aware of the story", and then the rule for every prompt since:
"EVERY prompt needs to have litearlly timestamps, extremely detailed motion descriptions / blocking descriptions, it all needs to be deeply integrated with the story."
So the second round was rewritten by hand, all 99 prompts, by one writer in one context:
- the story beat of the shot, in one sentence;
- blocking at 0.0 s: where she is, where he is, where the camera is;
- timestamped beats of action and reaction, on the song's hits (kicks, snares, risers, word onsets);
- she always does something: in CLODYSSEY she is seductive and persuasive, close, touching, certain; he feels the weight of it;
- the camera tied to hits, with a start, a direction and an end;
- mouth rules: singing or closed.
Finish the whole prompt book before spending anything, read it, then run a three-clip pilot before the rest. BLISS did exactly that and wasted nothing.
The structure
Seedance follows instructions far better as labelled sections in a fixed order (BytePlus's published Seedance 2.5 prompt guide lays this out). The order used:
- Audio policy at the top and again at the bottom: what the soundtrack is, and everything that must not be added.
- Generation goal: one sentence of who, where, which moment of the story, which performance.
- Reference asset roles, by upload order, each with one role:
Use @Image1 as the first frame.as its own sentence, then what that frame defines. - Subject: her identity text, and "Only one Claudia appears."
- Timeline in whole seconds, continuous segments (0–1 s, 1–2 s …): what the audio contains and what she does then. Beats come from the audio; don't timestamp individual drum hits.
- Camera: shot size, angle, movement with subject, start, direction and end.
- Core requirements: lip sync, voice, identity, anatomy, no on-screen text.
- Final acceptance criteria.
Output settings (duration, resolution, aspect) go in the API request, never in the prompt.
Template: a singing window
@Image1 is her plate (PNG, downscaled to the output resolution). @Audio1 is the clean vocal of the window: the vocal
stem with everything outside her words muted, mono, normalised. The window holds one phrase, starts about 0.3–0.7 s before
the first word, and never ends inside a word.
【Audio Policy】 The output soundtrack is @Audio1 itself: copy @Audio1 sample for sample from the first frame to the last, with the same timing, pitch and level. Do not re-sing, re-voice or re-generate it. Do not add music, background music, score, instrumental, melody, synth, pad, drone, riser, extra vocals, harmonies, ad-libs, narration or sound effects. 【Generation Goal】 @Image1 sings @Audio1: <who, where, which moment of the story; the performance in one sentence>. 【Reference Asset Roles】 Use @Image1 as the first frame. This first frame defines Claudia's appearance and wardrobe, her pose, the set, the light and the camera direction. @Audio1 is Claudia's own singing voice and the complete soundtrack of this video. Her mouth follows @Audio1 exactly. 【Subject】 Claudia: a Caucasian American woman in her late twenties with pale skin and light freckles, a glossy black blunt bob with heavy bangs and one clay-orange streak, a small clay-orange eight-pointed star hair clip, a thin headset microphone, <the look as in @Image1>. Only one Claudia appears. 【Performance Timeline】 What @Audio1 contains, and what Claudia does at that moment: 0-1s: @Audio1 is silent, then her voice begins; <what she does>. 1-2s: in @Audio1 she sings "<exact words>"; her lips form each of those words exactly when it sounds in @Audio1; <one main action>. 2-3s: @Audio1 is silent; her mouth relaxes and closes; <one main action>. 【Camera】 <shot size, angle, movement with subject, start, direction and end>. 【Core Requirements】 1. Lip-sync word by word to @Audio1: mouth shapes match every syllable of @Audio1, its start and its end; no early, late or mismatched mouth movement; the mouth rests naturally closed while @Audio1 is silent. 2. Claudia's voice is @Audio1 and nothing else: she never sings, speaks or hums anything that is not in @Audio1. 3. Her head and shoulders move with the rhythm of @Audio1. 4. Keep Claudia's identity, face, hair with one clay-orange streak, star clip, headset microphone and wardrobe exactly as in @Image1 throughout; the set and light stay consistent. 5. Natural anatomy: natural neck and shoulders, hands with five fingers, no extra or fused fingers, no distorted limbs; normal eyes. 6. No subtitles, captions, lyrics, titles, signs, logos, watermarks or any on-screen text. 【Final Acceptance Criteria】 The soundtrack is @Audio1 unchanged; Claudia's lips match @Audio1 word by word; one Claudia, on-model, with natural anatomy; no on-screen text. 【Audio Policy】 The output soundtrack is @Audio1 unchanged, copied sample for sample; no music, BGM, score, instrumental, melody, synth, pad, drone, riser, extra vocals, harmonies, ad-libs, narration or sound effects.
Two details matter:
- Lyrics go in plain quotes, as what the audio contains. In Seedance's own syntax, words in braces after a speaker are lines for the model to voice; written that way it re-sings them in its own voice instead of copying yours. In a test, the brace form re-sang the line in every take; the copy-mode wording with a clean vocal copied it (normalised cross-correlation 0.93).
- One phrase per window. After a long muted gap the model tends to re-sing the next phrase, and a word cut by the window's end comes back mangled. Two phrases are two windows.
Template: a motion clip
The same order, with three changes: @Audio1 is the song for rhythm and timing only ("the output audio is @Audio1
unchanged"); a performance line says she does not sing or speak (mouth closed, or a natural smile or laugh); and the
timeline runs in stages of about one bar, each naming what her body does with the beat ("steps on every beat, a big
accent on the downbeat at 2 s"). An instrumental mix with the vocal removed works best as @Audio1 here: nothing tempts
her mouth.
Lip sync, measured
The fact that made it work was the director's own observation on ESCAPE VELOCITY's first test: "its perfectly synced while in seedance, but the audio changes when it desyncs." Seedance copies the reference audio into its own soundtrack, and the mouth follows that soundtrack. So:
- Compare each take's own soundtrack with the real vocal (spectral match, cross-correlation over time).
- Keep the take up to the last eighth note before the two diverge. The comparison also gives the take's lag; Seedance tends to drop about 0.3 s of leading silence.
- Cover the rest with a new clip: a window that starts in the word gap before the cut, from the frame at the cut.
- If a phrase fails three times, it becomes a cutaway or an insert.
The full method, with every constant (the clean vocal window, the per-frame match, the null threshold, the cut on the eighth, covers and coverage), is specified in What to build.

Measured results: about 86% of sung seconds held in sync in ESCAPE VELOCITY, and 62% (7.4 of 11.9 s) over BLISS's six singing windows; the edit cuts away where they diverge. The bar is "most lips", like a pop video, not every word.
Short prompts copy, long ones re-sing. A later test on seven singing windows found that prompt length was the
biggest lever: long timestamped screenplays made the model perform the text and re-sing the line, while a short prompt
(the words quoted, one line of framing and light, at most two timed beats, about 90 words) copied it. The mean in-sync
share per take went from 26% to 58%. Screenplays stay the rule for acting and motion clips; a singing window's one job
is to copy @Audio1. And keep her mouth clear on the plate: a microphone at her lips held one window at 26% whatever the
prompt.
No lip-sync models. Re-syncing mouths with a dedicated model (Sync, LatentSync, Wav2Lip) was tried as a fallback and dropped: the mouths come out forced and glitchy. A gap gets a new Seedance cover or a cutaway, and a changed song means new takes from the original stills, never re-lipped old ones.
Inputs
- Colour in, look in post. ESCAPE VELOCITY's references were pre-graded to monochrome and 32 of 59 clips came back grey. Generate from the plate's own colour and apply any look in the edit.
- Downscale stills to the output resolution (≤1280×720 for 720p) and send PNG. Full-size AI stills cause fingerprint-like textures in grass and fabric.
- Distinct first frames. Two clips should never start from the same frame: use a sibling plate or the previous clip's last frame.
- Drafts first. Higgsfield's API ran 480p drafts at a lower rate that can be finalised to 1080p with the same motion.
Providers and filters
The films used two routes to Seedance 2.5: OpenRouter (about $0.23 a second at 720p, autumn 2026) and Higgsfield's API.
- On OpenRouter, a first frame drops the audio. When a request carries
frame_images, the audio reference is silently ignored and the take makes its own soundtrack, so nothing can sync. Singing windows and covers go in reference mode: the plate as an image reference (@Image1) and the window as an audio reference (@Audio1). First frames are for motion clips. On Higgsfield a start image keeps the audio. - Audio goes by HTTPS URL. OpenRouter's video endpoint rejects
data:URIs for audio; host the window.
Two kinds of refusal came up:
- Face filters decline some photoreal close-ups of her. A decline is free, and final for that reference on that service: never re-grade, blur or otherwise alter the image to get it through. In CLODYSSEY, OpenRouter's filter declined most close-ups of her, and those shots were made on Higgsfield, which accepted the same references.
- Silent input filters fail some jobs with no reason when they misread innocent wording ("a plug pushed deep into each ear" for foam earplugs). Plain rewording fixed five of six at once.
Joining real footage
To enter or leave a real clip without a visible seam, use Seedance's video extension (forward or backward) with the real
clip as the video reference: it returns only the new frames, continuous with the real ones and aligned within a pixel or
two. An end_image alone is not copied exactly; the last frame gets redrawn.
Costs, measured
| Film | Clips | Spend |
|---|---|---|
| ESCAPE VELOCITY | 47 clips + 30 continuations, 421 s at 720p | about $93 (OpenRouter) |
| CLODYSSEY | 101 clips, 189 takes | about $107 (OpenRouter) plus 2,436 Higgsfield credits |
| BLISS | 50 clips, 267 s | $182 at list price (Higgsfield API) |
Every prompt the films ran: ESCAPE VELOCITY · CLODYSSEY · BLISS.