Claudia.gallery

What to build

The tools behind the films, specified so that you or an agent can rebuild them: the job of each, what it takes and returns, how it works with the constants that were measured, and the check that says it is done.

Markdown  ·  MCP: get_guide("toolkit")  ·  29 min · 10 sections

The films' tools lived in a private workspace and are not published. This page is the spec instead. For each tool: its job, what goes in and what comes out, how it works (with the numbers that were measured or tuned), and the check that says it is done. None of it is exotic. Most of these tools are a few hundred lines of Python or JavaScript around ffmpeg, an audio model or a headless browser. Build them in the order of Start here. Their files are in Data formats.

Two principles run through all of them:

  • A check a script can run beats a judgement. Every bar that held on the films was something a script measured: a word on screen at its time, a cut on a beat, a mouth on the audio, money from the bill.
  • Fixed formats with slots beat improvisation. The Suno pack, the plate prompt order, the clip prompt, the display tokens and the shot schema are templates. An agent fills them in; it doesn't reinvent them.

The project

1. The project folder and its logs

Job. Hold everything a fresh agent needs to pick the film up cold.

Out.

  • One folder per film (layout in Data formats).
  • A decisions log: the director's words, verbatim, dated, by whom. A later decision is a new entry that supersedes an older one, never an edit of it.
  • A ledger with one line per paid call.
  • A runs file with every long or delegated run's id and arguments.
  • One status file, rewritten in place: the stage, what is running, real spend, the checkpoints passed, and what needs the person.

How. Name the project on every command. An export inside a backgrounded shell chain once sent commands to the wrong project. Write state with a locked read-modify-write: parallel runs raced on a picks file and lost picks.

Done when a new session can answer "where are we, what is running, what did the director say, what did it cost" from the files alone.

The song

2. The lyric scorer: scansion and assonance

Job. Before any take is spent, predict how Suno will deliver a lyric, and report two numbers per section and for the song: scansion and assonance. Also lint the sheet.

In → out. The lyric sheet, with a section tag on each section naming its delivery ([Verse 1 - spoken, deadpan], [Chorus - sung], [Drop - chant]) → a report per line and per section. Optionally, each line drawn on the 16th-note grid.

How.

  • The delivery model. It was measured by laying ESCAPE VELOCITY's real take, forced-aligned, on its beat map, then calibrated on the director's verdicts on CLODYSSEY's drafts. Suno gives each spoken line its own two-bar slot. The line enters on the "and" of beat 2 of the first bar, its last word lands on beat 4 of the second bar, and Suno speaks it at natural rhythm stretched to fit (verse lines 0.8–1.5×, median 1.12).
  • Rates: spoken about 3.3 syllables a second, sung about 2.9, chanted about 3.0. Aim for 10–13 syllables per spoken line; 9 can pass. A sung hook runs 4–8 syllables, with common words and open vowels.
  • Syllables and stress. Split every word into syllables with stress, using a pronouncing dictionary such as CMUdict plus a fallback for unknown words. Mark the important syllables: lexical stress, names, numbers, words in caps, and the last content word of each phrase. Function words (the, a, of, to, and, is) should fall between beats.
  • Flow, for spoken sections, scores:
    • how many important syllables land on the grid;
    • stress spacing (no pile-ups, no long runs of weak syllables);
    • how far Suno must stretch the line;
    • the line's shape: at most two phrases between full stops, and no one-word fragment mid-line ("Ugh.", which the director called "the most awkward part");
    • whether couplets rhyme on the beat-4 word.
  • Groove, for chants and hooks, scores whether short clauses lock to the pulse. Never use it on verses: optimising verses for groove made one draft choppy.
  • Scansion = flow without its rhyme term, averaged with groove.
  • Assonance. For each line end, find its best partner within four lines: a rhyme, a slant rhyme, or shared stressed vowels. Add the echo of stressed vowels inside lines. Report each line's end sound and its partner.
  • Lint:
    • digits and unspaced acronyms (write "A G I" in the sheet; the screen shows "AGI");
    • brackets and symbols, and homographs;
    • lines that read as AI-written, and the project's taboo words;
    • lines borrowed from other songs;
    • a lyrics field over about 3,000 characters (Suno rushes or truncates silently).

Reference scores (scansion / assonance) on the finished songs: ESCAPE VELOCITY 75 / 39, CLODYSSEY 83 / 55, BLISS 80 / 43. Spoken, deadpan lines may stay unrhymed; every sung couplet should ring.

The rewrite loop:

  1. Write 4–8 rewrites of a flagged line that keep its meaning and its reference.
  2. Score them and keep the top three.
  3. A judge rates each for meaning, wit and the risk of reading as AI-written.
  4. Pick the best.

Never touch a line the director has locked, and never optimise past their ear. A draft tuned to a perfect score got: "perfect beat precision without attention to what words land where doesnt lead to this being great".

Done when the scorer ranks every section the director has judged in the director's own order. Keep a calibration file of those sections with their verdicts, and re-run it after every change to the scorer. If the ranking breaks, the director is right and the score is only advice.

3. The Suno pack

Job. Everything a person pastes into Suno, with each field in its own block so it copies from a phone.

Out.

  • Style: genre, BPM, key, instruments, vocal delivery and mix, as one string.
  • Exclude styles.
  • The lyrics field: spelled as sung, with section tags that carry delivery and arrangement cues; 1,800–2,400 characters for about 2:20.
  • Title and settings.
  • A takes log: each take with the director's verdict on it.

How.

  • Respell a name Suno mispronounces in the Suno field only; the screen keeps the real spelling.
  • Allow one Voice per source song (a combined spoken-and-sung Voice), and also try takes with no Voice.
  • If the film has a sound palette, put the song in its key. XP's startup and shutdown sounds are in E-flat major.
  • Syncopate hooks and stop lines.
  • Change one thing per round.
  • Never take audio from Suno's CDN or by recording the player. The person's download is also the licensing step.

Done when the person can paste every field without editing, and the takes log names the chosen take.

Listening

4. Stems

Job. Split the master into vocals, drums, bass and other.

How. htdemucs, four stems. It took about 64 s on 8 CPU threads for a 5-minute song, with a 2.7 GB peak and 84 MB of weights; the published weights carry no explicit licence. The stems feed different tools:

  • vocals: alignment, lip-sync windows and the sync check;
  • drums: the beat map and the kicks;
  • bass and other: impacts, risers and section energy.

5. Timed words and the ear check

Job. Give every sung word a start, an end and a confidence, then list the ones a person should hear.

In → out. The vocal stem and the lyric sheet → words.json (each word: index, text, line, t0, t1, conf, and its source), plus an ear-check list.

How.

  1. Forced alignment of the known lyric: a wav2vec2 CTC model (facebook/wav2vec2-large-960h-lv60-self, Apache-2.0) gives character probabilities every 20 ms, and a Viterbi pass walks the sheet's characters through them. The confidence is the mean probability over the word's frames.
    • Re-run on ESCAPE VELOCITY, it reproduced the shipped times exactly for 78.6% of words, and within 0.10 s for 92.5%. Target at least 90% within 0.10 s.
    • Singing aligns worse than speech: mean confidence 0.64 on ESCAPE VELOCITY, 0.75 on CLODYSSEY.
  2. A transcription fallback for words the aligner is unsure of. Whisper hears sung words far better (92.5% text match against 72%), but it starts words early: mean error −0.33 s. Merge per word, in this order:
    1. wav2vec2 confidence ≥ 0.5: keep it.
    2. Otherwise, if Whisper heard the word, take Whisper's end time. For words longer than 0.45 s, also move the start forward to the first 10 ms frame where vocal energy passes 0.3 × the word's peak, unless wav2vec2's start falls inside the word.
    3. Otherwise, if the word sits between two placed words of its line, give it an even share of the gap.
    4. Otherwise, leave it unplaced, for the ear check.
  3. Matching heard words to the sheet:
    • spell digits out;
    • split a heard word that equals 2–4 sheet words ("AGI" → "A G I");
    • map names and brands by hand, in a project table.
  4. The ear-check flags, most urgent first:
Flag When Usually
letters a sheet line has no letters the sheet is wrong
missing a word has no time fix it
order a word starts more than 0.05 s before the previous one an alignment slip
gap more than 1 s between two words of one line a delay throw, an echo, or the wrong repeat
verylow confidence under 0.1 the flag that matters most: 11 of the 12 words that were really off on ESCAPE VELOCITY's take carried it
length a word longer than 1.2 s it swallowed a pause or a neighbour's slot
lowconf confidence under 0.5, or a time from the fallback or interpolation usually right

Two flags were tried and dropped, because every word they flagged was correct: a word on near-silence, and a word shorter than 40 ms.

Judging without ears. Weigh the signals in this order:

  1. the flag itself;
  2. a second aligner: if the two agree within about 0.3 s, keep the word;
  3. neighbours and repeats: which occurrence of a repeated line sits inside its line's span;
  4. vocal energy, only to choose between candidates that the signals above already proposed.

Send at most five spots per song to the person, as 10–20 s previews with the words on screen: "say ok, early or late for each".

What the take really sings. Transcribe the vocal stem and compare it with the sheet line by line: ok, spelling, changed, extra, missing or not heard. Never apply a proposed change the person hasn't heard: on CLODYSSEY's take, all nine proposals were transcription errors.

Rules.

  • If the take sings something else, fix the sheet. If it sings the sheet, fix the times. Never edit the audio.
  • Manual fixes carry a note and survive re-runs.

Done when every flagged item is kept, moved or asked about, with a note, and all of it is finished before the board.

6. The beat map

Job. Every beat and every bar line of this take. Never one BPM: Suno takes drift like live music. ESCAPE VELOCITY went from 131.5 to 133.9 BPM and CLODYSSEY from 129.2 to 132.1, so a fixed grid is more than a beat off by the end.

How.

  1. Take the onset strength (spectral flux) of the drum stem.

  2. Fit local tempos on 20 s windows every 5 s: BPM in 0.05 steps, phase in 4 ms steps, inside a tempo range. Drop the weakest 20% of windows (the breakdowns).

  3. Smooth the tempo curve, step beats along it, and snap each beat to a real drum onset within ±30 ms.

  4. Find the downbeat phase (which of every four beats is beat 1) with two votes per 32-beat block:

    • the kick vote: the phase with the most kick energy;
    • the chord vote: the phase where the harmony (chroma of the other and bass stems) changes most.

    The chord vote wins when it holds at least half the blocks and changes at least 1.5× more than the other phases; otherwise the kick decides. On ESCAPE VELOCITY the kick vote was split 36% while the chord vote held 77%. Keep a manual override.

  5. If the tempo search locks onto half or double time, pin a tempo range.

Done when section starts, verse entries and impacts mostly fall on beat 1, with kicks heavy on 1 and 3, claps on 2 and 4, and harmony changing on 1. CLODYSSEY's final take needed the override: claps on 2 and 4 fooled the kick vote.

7. Sections and hits

Job. What the track does, where: the camera answers sounds you can hear, not just the grid.

Sections. Segment the harmony and timbre, snap to downbeats, and name the sections from vocals and energy. Confirm them against the lyric's tags.

  • Pickup rule: a line starting less than 0.8 s before a section's first downbeat belongs to that section.
  • List every vocal gap of 2 s or more; instrumental shots and title cards go there.
  • Print each stem's energy per bar on a 0–9 scale, and write one line per section on what the track does ("the beat stops after 'it sucks'").

Hits:

Event How it is found
kick onsets in the drum stem, 30–150 Hz
snare onsets at 150 Hz–5 kHz, minus anything within 45 ms of a kick (a broadband kick fires in every band)
hat onsets at 6–16 kHz, minus kick-coincident ones
impact a jump of the next 0.6 s over the previous 1.5 s in the mix, drum or bass envelope; score = max(mix, 0.9 × drums, 0.7 × bass) ≥ 0.22; snapped to a downbeat within 0.65 s; merged within 2 s
riser the other and bass stems rising for at least 1 s, with the hit each one lands on
fill the beat before a real hit, and bars whose last half breaks the pattern
vocal entry where the vocal comes back after silence
envelopes drums, vocals and bass at 24 fps, scaled 0–1

Done when every impact sits in a loud bar of the energy table. Keep a manual add and drop list: the detector missed CLODYSSEY's loudest bar, which grew out of an already loud chorus.

8. Display tokens

Job. What the viewer reads, which is often not what is sung. The display is the joke; the sheet is the phonetics.

How. One line per lyric line, made of tokens: TEXT|n shows TEXT for the next n sung words, from the first word's start to the last word's end.

  • +25%|3 =|1 −17%|2 shows "+25% = −17%" over the sung "plus twenty-five is minus seventeen".
  • AGI|3 shows AGI over the three sung letters.
  • Markers: *x* accent, ~x~ strike-through, _x_ mono. |0 rides on the next token.

Done when every line consumes exactly its aligned words. Fail loudly if one doesn't.

The board

9. The board file and its audit

Job. The film as data, written by commands that check every value, never by hand-editing JSON.

In → out. The beat map, sections, words and budget → shots.json: a draft, then the written board.

How.

  • The draft cuts each section at a pace: seconds per shot, per section. ESCAPE VELOCITY's measured pace was about 1.35 s in the intro, 1.6 s in verses, 2.7 s in sung choruses and 4.4 s in the outro. A calmer song might run 4–7 s a shot.
  • Every creative field starts as a TODO.
  • Writer commands set fields one shot at a time, or from a table. They refuse:
    • placeholder text;
    • unknown device kinds, bad anchors and bad timing keys;
    • a prompt in the wrong generator's field.

The audit's gates are listed in The storyboard. It exits non-zero while any gate fails, and it prints the exact entry for every clip still to make.

Done when the audit exits 0. The approval is recorded against a hash of the board, so any later change makes it stale.

10. The board sheet

Job. The board as something a person can review on a phone.

Out. A PDF and JPG pages of 12 cards each. Each card holds the frame (or the prompt's first words), the id, the time, the lyric, the framing, the action, what moves and the explaining graphic. A cover page shows the chapter map. A FACE? badge marks every plate listed in a flags file. If the film has a grade, the frames are graded.

Pictures

11. The plate book, contact sheets and picks

Job. One Midjourney prompt per plate, the images back in, and a pick for every shot.

How.

  • The prompt order: framing first, then her identity text (the full form for close and medium shots, the short form for full length and wider), wardrobe, the action and her position in the frame, any empty field kept for type, the film's style family, and "no text, no letters, no logos". Then the parameters. See Midjourney and the canon.
  • Runs: a person runs the batches in Midjourney, at about a minute a job, four images each.
  • Ingest:
    • identify each job by its own prompt text, never by its position on the page;
    • Midjourney rewrites long prompts on submit, so match those by hand;
    • keep all four candidates.
  • Contact sheets are numbered and show faces large.
  • Picks: plate → file, candidate number, and the shots that use it.
  • Variety: the three candidates of each job that weren't picked are a free variety pool. Use one image per shot, and never start two clips from the same frame.

Done when every plate on the board has a picked file, and no file sits on two shots that are not one continuous moment.

12. Plate prep: depth and boxes

Job. What the 2.5D moves, the type and the camera need from each still.

How.

  • Depth: Depth Anything V2 Small, which is Apache-2.0; the Large model is non-commercial.
  • Boxes: face and person boxes from Grounding DINO. Graphics anchor to them, and the camera keeps the face in frame through pushes and shakes.
  • Mattes, where type tucks behind her.
  • Re-run all of it after any plate edit: edited plates once rendered with their old depth.

Moving shots

13. Clip files and the spend gate

Job. One file per clip, from draft to delivered, and nothing spent without a locked prompt and a yes.

How.

  • Status: draft → locked → submitted → delivered, declined or failed.
  • Each clip file holds the kind (sung or motion), the provider, the start image, the window's song time, its exact words and the request. Each take records its file, its cost and its sync result.
  • Locking checks:
    • the words are quoted exactly;
    • the window rules hold;
    • a sung prompt stays under about 90 words and two timed beats (warn).
  • Submitting:
    • submit only when the clip is locked and the approval covers this exact prompt;
    • validate the job id: "submitted" with no id, and HTTP 500, 502 or 520, were silent failures;
    • never resubmit a request that got a job id;
    • write a manifest after each batch, so an interrupt never loses which jobs exist.
  • Polling: 70–130 s a clip on OpenRouter, and 4–11 minutes on Higgsfield.
  • Real spend comes from the provider: OpenRouter's credits and key endpoints, and Higgsfield's transactions.

Provider facts, checked with paid probes:

  • HTTPS audio only. OpenRouter's video endpoint rejects data: URIs for audio. Host each window over HTTPS: a Cloudflare quick tunnel worked, and so did a provider's media URL.
  • A first frame drops the audio. On OpenRouter, frame_images overrides the audio reference: the take makes its own soundtrack and nothing can sync. Send sung clips and covers in reference mode: the plate as an image reference (@Image1) plus the window as an audio reference (@Audio1). First-frame mode is for motion clips.
  • Face filters. They often decline photoreal faces, and the decline is free. Send the same unaltered image to the other provider.
  • Audio refusals. "Output audio may contain sensitive information" and copyright refusals on vocal references are unbilled. Retry once with a re-cut window or a reworded prompt, then try the other provider.
  • Higgsfield.
    • Ask for the price before each new configuration.
    • Upload by URL: presigned uploads flood an agent's context.
    • A 480p draft can be finalised to 1080p within seven days.

Done when nothing can spend without a locked clip and a recorded yes, and the ledger reconciles with the bill.

14. The clean vocal window

Job. Give the video model one voice to copy: only the words she sings on camera in the window, and nothing else.

How, with the constants used:

  1. Cut the vocal stem at [song_t0, song_t0 + duration], mono, 44.1 kHz.
  2. Keep the aligned words that start inside [song_t0 − 0.02, song_t0 + duration − 0.08), and only hers. Mute the tail of a word that began before the window.
  3. Merge words into phrases across gaps of 0.35 s or less.
  4. Keep each phrase from 0.05 s before its first word to its last word's end. Add a release tail that follows the stem's 10 ms energy until it falls 24 dB below the phrase peak: at most 0.40 s, and never past the next phrase's start minus 0.05 s.
  5. Put 15 ms raised-cosine ramps on every gate edge, and silence everywhere else.
  6. Normalise the kept part to −16 dBFS RMS, then limit peaks to −1 dBFS.
  7. Write a WAV (OpenRouter, by HTTPS URL) and an MP3 at 192 kbps (Higgsfield).

A window holds one phrase, starts in a word gap, and never ends inside a word.

15. The sync check: the spectrogram comparison

Job. Measure how much of a sung take is really in sync, and where it stops.

The director noticed it on ESCAPE VELOCITY's first tests: "its perfectly synced while in seedance, but the audio changes when it desyncs. if you check the spectrogram for audio against reference audio you can see the clear divergence point. you could just create a cut and make another clip to cover desyncs."

Seedance copies the reference audio into its own soundtrack, almost sample for sample while it copies, and the mouth follows that soundtrack. So compare the take's soundtrack with the window you sent, find where they part, and cut there.

In → out. The take and the exact window file → the lag, a copy flag, the in-sync spans in song time, the divergence, the cut, the words at the divergence, a diagnostic image and a review video.

How, with the constants used:

  1. Decode both to 16 kHz mono. Use 10 ms frames (hop 160), 64 ms FFT windows (1024) and 80 mel bands from 50 to 7,600 Hz.
  2. Global lag: FFT cross-correlation within ±1.0 s. The peak's normalised value is the take's NCC. The take is a copy if NCC > 0.2; otherwise the model re-sang it.
  3. Place the reference on the take's timeline at that lag. Per 10 ms frame:
    • wave: normalised cross-correlation of the waveforms over 100 ms. A verbatim copy scores 0.6–1.0.
    • mel: cosine similarity of mean-removed log-mel frames (70 dB floor), smoothed over 11 frames. It holds on fricatives, whose noise never correlates sample by sample.
    • active: either track's RMS is above 10% of its 95th percentile.
  4. A threshold per take, against a null. Compute the same mel similarity with the reference shifted by ±37, ±61, ±89 and ±131 frames; these offsets avoid eighth-note multiples. Then:
    • null90 = the 90th percentile of the null;
    • matched_med = the median mel where wave > 0.6, or the 75th percentile of active mel if there are 10 or fewer such frames;
    • thr = null90 + 0.4 × (matched_med − null90).
  5. A frame matches when it is active and wave ≥ 0.4 or mel ≥ thr.
  6. Divergence is the first 300 ms stretch (at least 60% active, under 30% matched, starting on an unmatched frame) that never recovers with a 400 ms matched run. Short dips that recover are dips, not cuts. With no such stretch, voice at least 80 ms after the last match counts as divergence. synced_until = the last matched frame minus 42 ms (one frame at 24 fps).
  7. In-sync spans are matched voice, bridged across silences and across up to 0.25 s of unmatched voice. Drop spans shorter than 0.5 s.
  8. Map everything to song time: song time = song_t0 + clip time − lag.

The diagnostic image shows our vocal's spectrogram with word labels, the take's own soundtrack, the match curves with the threshold, and the cross-similarity matrix, whose diagonal means "in sync". The review video lays our audio over the take, marked IN SYNC up to the cut and DIVERGED — CUT HERE after it. Look at both before trusting a number.

The NCC is a diagnostic, not the verdict. Over a whole take it is usually 0.1–0.3, because the copy stops somewhere. One take scored 0.157 and still held 99% of its sung seconds. Measure the lag per take, never as a constant: takes ran from −0.40 s to +0.83 s. Motion clips conditioned on the mix aren't copied (NCC about 0.05–0.08): play them at lag 0.

Done when each take has its spans and its cut, and the review video agrees with them by eye.

The sync check on an early ESCAPE VELOCITY test: the vocal that was sent, the soundtrack that came back, and how closely they match over time.
The sync check on an early ESCAPE VELOCITY test: the vocal that was sent, the soundtrack that came back, and how closely they match over time.

16. Covers and coverage

Job. Keep the in-sync part of every take, cover the rest, and know how much of the singing is really synced.

The cut. Cut on the last eighth note at or before synced_until. Eighths are the beat map's beats plus the midpoints between them. A cut on an eighth feels like part of the music; a cut at the divergence frame looks like a glitch.

A cover is a new clip:

  • Window: it starts in the word gap before the cut, and lasts clamp(ceil(end − start + 0.3), 4, 8) whole seconds.
  • Start image: the parent take's frame at the cut (clip time = cut − song_t0 + lag), unaltered, sent as the image reference, never as a first frame.
  • Words and audio: exactly the words that start in its window, with their clean vocal.
  • Prompt: the parent's prompt, re-quoted to the new words.
  • Records: the parent lists its covers; the cover records its parent and why.
  • Limits: at most two covers per root, and one on a provider where copies are rare.

Coverage, per chain (a root and its covers):

  • sung stretches = each aligned word ± 30 ms, merged when closer than 0.12 s;
  • in-sync spans = every element's spans, each start extended by 60 ms;
  • gaps = sung time not covered by any span, of 0.6 s or more, merged when closer than 0.3 s.

The bar is "most lips", like a pop video ("like grimes vids … shedoesnt sync lips on everything, just most"). Target at least 80% of sung seconds, never an out-of-sync mouth held on screen, and budget covers at about 60% of the sung seconds.

Measured:

Run Result
ESCAPE VELOCITY 41% of sung takes copied; with covers, about 86% of sung seconds verified
CLODYSSEY, Higgsfield about 21% of sung takes copied
BLISS 62% of six windows' sung seconds
A later test run, seven windows long prompts held 44% of the chorus's sung seconds; short ones 62–71%

When a window won't copy, try these in order:

  1. The short sung form. The biggest lever measured is prompt length. Long timestamped screenplays make the model perform the text and re-sing it; short prompts copy. The short form is:

    • the head, quoting every word;
    • one line of framing and light;
    • at most two timed beats;
    • "One continuous shot, no cuts, no head turns. Only she sings. No text, letters or logos.";
    • about 90 words or fewer.

    On seven windows the mean in-sync share per take went from 26% to 58%; one window went from 0% to 78%. Screenplays stay the rule for acting and motion clips.

  2. Clear the mouth. Nothing in front of it on the plate: no held microphone, hand or hair.

  3. The other provider, after two non-copies in the short form.

  4. Cut away: to the shot's motion clip, stills in rhythm, depth moves, crops to the eyes, hair or hands, inserts, or a wider or profile framing. The words stay on screen either way.

No lip-sync models. Sync, LatentSync and Wav2Lip were tried as fallbacks and dropped: "still not exact", and the mouths came out forced and glitchy. The failure was the input, not the model.

The edit and the renderer

17. The edit list

Job. Resolve every moment of every shot to a source.

How. Resolve a sung shot moment by moment in song time. At each time T, use the chain element whose in-sync span contains T, at clip time T − song_t0 + lag. A take and its covers then interleave, each switch sitting on the eighth before a divergence. Fill holes from the shot's motion clip, then from the plate. Motion clips play straight through, holding their last frame if the shot is longer.

Done when the film renders end to end at every stage, with plates standing in for clips that don't exist yet.

18. The renderer

Job. Draw every frame: footage, 2.5D, lyrics, explaining graphics, chrome and transitions, on the song's clock.

The frame contract. It renders offline in headless Chromium, one frame per 1/24 s, in parallel and out of order.

  • Every frame is a pure function of song time t. No state between frames, no Math.random (use a hash of the index), no Date.
  • Two canvases. The base holds footage and anything graded with it. The overlay holds type and chrome, composited crisp on top.
  • Memos only for data derived from static media, keyed by media name, never by frame.
  • Every 2D canvas is software-backed (willReadFrequently), and every scene starts from a reset context. A GPU-backed canvas under software GL kept the previous frame's overlay once frames got dense; that took about two hours of bisecting to find.
  • Quantise to frames. A cut shows on the first frame at or after its time. Video may run up to one frame ahead of the sound, never behind it.
  • Speed: at most 1,500 ms a frame. The 64-step parallax shader costs 0.5–1.1 s, and a full-resolution canvas blur about 0.5 s. Use depth slices (about 30 ms) and blur by resampling instead.

The clock. All musical time goes through the beat map: beat position, bar, local beat length, kick pulses and eighths.

  • A lyric token appears at its aligned start exactly, never snapped to a beat: "the motion syncs to the beat, the word to the voice".
  • An entrance takes at most an eighth of a beat.
  • Every timed thing carries a timing key:
    • W a word's time;
    • B the beat grid;
    • H a track event (kick, snare, impact, riser, fill, envelope);
    • F a fixed song time;
    • T a keyframe measured on a take's own frames.

Lyric modes. Every sung word appears in at least one, and every drawn token is marked for the lyric gate.

Mode For
subtitle sung lines; the sung word lit in the accent
coverline spoken lines set beside her in the plate's empty field, with tracking collapsing word by word
masthead payoff words, numerals, the answer in a call and response
stack short emphatic lines, one word per row
mass a declared wall of micro text in a drop, with the sung line still drawn
title the title knocked out of black

Type rules:

  • One lyric block per frame; a line keeps one mode and one position across cuts.
  • At least 40 px, and never over her face.
  • Margins are left 72, top 56, right 96 and bottom 84 at 1920 × 1080.
  • Hang type from the left axis or from a feature in the frame.
  • Nothing may cover the lyric: the renderer measures it every frame and moves the other layers off it.

The camera plays the track:

  • punch on kicks: amp 0.01–0.03, gated by the drum envelope;
  • shake with the drums: none in screen or UI shots;
  • ride a riser into the hit it lands on;
  • impact on real hits: strength at least 0.6, zoom kick 0.06;
  • whip on fills.

Hold still where the song breathes. A move never cuts the face.

Density is chosen per film:

  • dense: three or more layers, five or six as the norm;
  • budgeted: footage, one lyric block, one hero device, at most two detection marks, the chrome, and the accent on at most two elements;
  • sparse: about 70% of the film with no graphic at all.

Chapters. In the Full scale each chapter is its own code file, written by an animator agent against this contract. Chapters draw only what is unique; shared helpers live in one library. Workers report shared-file bugs; the operator fixes them once.

Rendering. The render is resumable: skip frames that exist, write each frame to a temp file and rename it, open a fresh page every 200 frames, and retry a failed frame once. ESCAPE VELOCITY's 7,354 frames took about 63 minutes with three workers on 4 CPUs; at 1080p one cut's frames take about 9.6 GB.

Checks

19. The lyric gate

For every sung token, render the frame at t0 + clamp((t1 − t0)/2, 0.03, 0.12) and check that the token was marked as drawn on it. It prints N/N; anything less fails. ESCAPE VELOCITY: 417/417. CLODYSSEY: 315/315.

  • Run it per chapter while animating, and for the whole film before any render.
  • It proves presence and timing, not legibility; contact sheets check legibility.
  • Fix a mistimed word in the alignment, never in chapter code.

20. The history check

Render frame B fresh, then render it again right after frame A in the same page. It fails if any pixel differs by more than 40 levels.

  • Run three pairs per chapter, one across each seam, with one pair whose A is a heavy 2.5D frame.
  • It must print HISTORY-INDEPENDENT.

21. The flash check

This follows WCAG 2.3.1's general flash threshold.

  • Read the video at 96 × 54, 24 fps, in linear relative luminance.
  • A flash is a pair of opposing changes of at least 10% of the range, where the darker state is under 0.8, over at least a quarter of the frame.
  • Fail any one-second window with more than three flashes. Count pairs, not transitions: an early version counted each transition, so 2.2 flashes a second read as 4.4.
  • Flashes belong to a theme, on hits: "dont need to force flashing".

22. Pace against the music

Picture motion is the mean absolute frame difference of the review copy at 160 × 90 grey, per section. Compare it with each section's stem energy and its kick and hat onsets per bar. Busy picture over quiet music reads as "off pace". The fix on one film halved the cuts (0.92 → 0.47 a bar) and slowed in-shot motion until the picture followed the music.

23. Contact sheets and the reviewer

Contact sheets are how agents look.

  • Per chapter: every shot's first, middle and last frame; every lyric line's first and last word; every hit; each graphic's entrance; each transition's middle.
  • The reviewer sees at least 24 frames, plus full-size stills of the five most important moments.
  • The whole film: mid-shot frames of every shot, 6 × 24 to a sheet, plus the seams and 100% crops of faces.
  • Open every sheet. A sheet nobody looked at is not a review.

The reviewer renders and looks for itself, and never edits files. It scores four lenses, 0–10:

  • taste: the director's taste, which counts double;
  • craft;
  • sync;
  • story.

It returns at most 12 concrete fixes, each marked blocking or polish, and copies the lyric gate's printed line.

The verdict is computed in code, not taken from the reviewer:

  • A chapter passes when no blocking fix is left, taste is at least 7, the other three lenses are at least 6, and the gate reads N/N.
  • A review that returned nothing, or never looked at a frame, is a failed review, never a pass. The original scripts read an empty review as "no fixes", and runs that hit a usage limit reported success.

Run at most two review-and-fix rounds per chapter. Calibrate the reviewer on the director's verdicts: on CLODYSSEY the reviewers scored 6–7.5 a cut the director called "perfect. change nothing."

24. The safety and canon scan

Job. Look at every generated take before release.

How.

  • Make sheets of frames per take.
  • List every take checked, and every problem with its severity: block, fix or note.
  • Whitelist intended props by name.
  • Re-scan every re-make.

On CLODYSSEY the scan found 12 blocks in 96 "final" takes, among them nudity through a corset, modern objects and earmuffs.

Delivery

25. Encode and deliver

The master. Frames are sRGB, so convert through RGB with accurate rounding to BT.709 and tag the stream:

ffmpeg -framerate 24 -i frames/f%05d.jpg -i song.wav -map 0:v -map 1:a -c:a aac -b:a 256k -shortest \
  -vf "format=rgb24,scale=flags=accurate_rnd+full_chroma_int+full_chroma_inp:out_color_matrix=bt709:out_range=tv,format=yuv420p,setparams=color_primaries=bt709:color_trc=iec61966-2-1:colorspace=bt709:range=tv" \
  -colorspace bt709 -color_primaries bt709 -color_trc iec61966-2-1 -color_range tv \
  -c:v libx264 -preset slow -crf 16 -pix_fmt yuv420p -movflags +faststart master.mp4

An untagged or BT.601 stream looks lighter and washed out in QuickTime ("stuff is v blown out").

  • The review copy: the same, with -crf 20 -maxrate 8M -bufsize 16M.
  • Verify with ffprobe:
    • codec, 24/1, and a frame count equal to duration × 24;
    • bt709 / iec61966-2-1 / tv;
    • AAC stereo;
    • no container tags beyond the encoder's own.
  • Then run the flash check on the master.

The 9:16 cut follows a subject box per plate, so an off-centre lead stays in frame. It gets its own lyric gate.

Housekeeping:

  • Run encodes as tracked background jobs, never a detached process nobody watches.
  • Delete the frames only after ffprobe has confirmed the outputs.
  • Keep big scratch on the big disk.

Running agents

26. Workflows, run records and the stop switch

The films ran their heavy stages as multi-agent workflows (Claude Code):

  • storyboard: three director agents, three judges, and a merge;
  • chapters: an animator, an art director who renders and reviews contact sheets, and a fixer;
  • research and lyrics: topic researchers, fact verifiers, writers and judges;
  • the safety scan.

Rules that cost time to learn:

  • Save every run's id and arguments the moment it starts, and resume by run id in the same session. Finished agents replay from cache. An interrupt kills every running workflow, so keep phases short; ESCAPE VELOCITY's merge restarted three times.
  • Check that every agent actually finished. A run can report success while agents that hit a usage limit returned nothing. Carry open items forward explicitly.
  • Parallelism has a ceiling. On 4 CPUs, run two agents a workflow. Allow one rendering workflow at a time: six rendering agents filled the swap and the root disk.
  • Put long lists in a file, not in the launch arguments: a placeholder once launched in place of 96 paths.
  • The stop switch. On any quality complaint from the director:
    1. stop every job and workflow that can spend, after searching each running script for its generate calls;
    2. confirm at each provider that nothing is queued;
    3. log the director's words as a rule;
    4. fix the prompts, then re-make. Don't review the bad takes.
  • Write the acceptance test before the brief, and make the check something the worker can't edit. Every human observation becomes a metric the same day: "if a reviewer can see it and the judge cannot, the judge is wrong".