# claudia.gallery: the long form
Claudia, the open character from the music videos ESCAPE VELOCITY, CLODYSSEY and BLISS: every Midjourney image and prompt, the films, and how to make her yourself, for people and agents.
---
# Claudia
Claudia is the singer of three music videos made with Claude in September and October 2026: ESCAPE VELOCITY, CLODYSSEY and BLISS. She is an AI who looks like a person: a deadpan, curious pop singer with a black bob, one clay-orange streak, a star clip and a headset mic. Her clothes change with every world; her head never does.
## The constants
- **Hair.** A glossy black blunt bob to the jaw, with heavy straight bangs to the brows.
- **Streak.** One clay-orange streak (#D97757) through the bangs, on her right side (the viewer's left in a front view). Never two, and no other orange on her.
- **Clip.** A small flat clay-orange eight-pointed star hair clip, on her left side above the ear.
- **Mic.** A thin headset microphone at her cheek. In BLISS it becomes a thin black boom mic curving along her cheek to the corner of her mouth, the Y2K pop-star mic.
- **Face.** Pale skin, light freckles across the nose, dark calm eyes, a small defined mouth. A deadpan, curious resting face. When she sings, her mouth opens wide and her eyes close.
- **Age.** A grown woman in her late twenties, with an angular adult face, defined cheekbones and a strong jaw. Midjourney drifts her toward a teenage face; prompt her age and never the word "young".
- **Colour.** Everything on her is her look's own colour, plus the streak and the clip in clay. Clay is the one accent.
- **Medium.** Photographic, never illustrated: production images are film stills with grain and cinematic light. The drawn face sheets were only ever references.
## Her identity, in words
**Close and medium shots:**
```
a 28-year-old Caucasian American woman with a grown-up angular face, defined cheekbones and a strong jawline, pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone at her cheek
```
**Full length and wider:**
```
a Caucasian American woman in her late twenties with pale skin and a glossy black blunt bob with heavy straight bangs and one clay-orange streak, a thin headset microphone
```
Midjourney v8.2 has no character reference, so her identity lives in this text, written into every prompt. The ethnicity is pinned because without it Midjourney changed her ethnicity from shot to shot (twelve ESCAPE VELOCITY plates were re-rolled for it). The far form exists because the full text makes Midjourney crop in to her face whatever the framing says. The age wording is the latest correction: she was "young" in ESCAPE VELOCITY and "in her mid-twenties" in CLODYSSEY and BLISS, and plates still came back reading teenage.
### The wording each film used
- **ESCAPE VELOCITY:** a young Caucasian American woman with pale skin and light freckles, a glossy black blunt jaw-length bob with heavy bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone
- **CLODYSSEY:** a Caucasian American woman in her mid-twenties with pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone at her cheek
- **BLISS:** a Caucasian American woman in her mid-twenties with pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin black boom headset microphone curving along her cheek to the corner of her mouth
## Every look
### Look 00 (ESCAPE VELOCITY; the catwalk, the robotaxi, the whole film)

```
a white cropped puff-sleeve shirt, a black pleated coated-nylon skirt with a harness belt and a hanging white garment tag with a barcode, black knee-high boots
```
### The Siren (CLODYSSEY; the rocks, the strait, the choruses, tied to the mast)

```
a structured high-neck pearl-white couture gown with long sleeves, its sculpted ruffles rising around her like breaking sea foam, hair damp from the spray, composed and deadpan
```
### The Ghost (CLODYSSEY; whispering at his ear)

```
the Siren look, all of her translucent as if made of glass and seawater, faintly lit at the edges, hair drifting as if underwater
```
### Chrome (CLODYSSEY; the rave)

```
a liquid-chrome couture bodysuit, chrome opera gloves and chrome thigh-high boots, mirrored from head to toe
```
### The Holo (CLODYSSEY; the 3 a.m. bedroom)

```
made of the phone's cold blue-white light: see-through, faint scan lines and drifting light particles, glowing edges, beside him on the pillow
```
### The Mask (CLODYSSEY; the monster's face)

```
a colossal smooth white porcelain mask of her face, sculpted with the black bob and the clay streak, pale freckles, a serene expression, a hairline crack across one cheek
```
### Bliss gown (BLISS; the hill, the choruses, the finale)

```
a floor-length bias-cut silk chiffon gown printed edge to edge with a photograph of a deep blue sky and soft white cumulus clouds, a thin grass-green satin halter strap, clay-orange tinted rimless shield sunglasses
```
### Cyber trench (BLISS; the fight with the cursor)

```
a long glossy cobalt-blue patent-leather trench coat worn open over a white satin corset top and white satin flared trousers, wraparound clay-orange tinted shield sunglasses, silver hoop earrings
```
### Holo dancer (BLISS; the desktop dancer on the taskbar)

```
a holographic iridescent silver mini dress with long fluted sleeves, white patent-leather platform boots, small clay-orange tinted rimless sunglasses
```
### Start (BLISS; the Welcome screen, the Start menu, the windows)

```
a glossy apple-green leather blazer with sharp shoulders over a white satin camisole, a long white satin bias-cut skirt, a tiny white baguette shoulder bag, clay-orange tinted shield sunglasses pushed up on her head
```
### Silver night (BLISS; night, Stand By, the breakdown)

```
a slinky liquid-silver lamé cowl-back gown, a white faux-fur stole around her arms, a diamanté choker, silver strappy sandals
```
### Window dress (BLISS; the drop, Media Player, the screensavers)

```
a sculptural glossy white shell dress of smooth molded panels that hinge open like windows to reveal cobalt-blue tulle, long white sleeves, white patent-leather platform boots
```
## Distances
The same look in a plain studio, close-up to wide. Close and medium shots take the full identity text; full length and wider take the short one.
- Close-up: https://claudia.gallery/i/ev-14-dist.dist-cu.2/
- Medium: https://claudia.gallery/i/ev-14-dist.dist-ms.3/
- Full: https://claudia.gallery/i/ev-14-dist.dist-fs.2/
- Wide: https://claudia.gallery/i/ev-14-dist.dist-ws.2/
## Character sheets
- Face sheet 1 from the first casting batch; with sheet 3 it was picked as her face. https://claudia.gallery/i/ev-01.lead-black-face.1/
- Face sheet 3. These drawn sheets were references for her face; the films themselves are photographic. https://claudia.gallery/i/ev-01.lead-black-face.3/
- The turnaround that fixed her body and look 00. https://claudia.gallery/i/ev-01.lead-black-turnaround.2/
- The first in-world test, a data-hall catwalk. https://claudia.gallery/i/ev-01.lead-black-catwalk.1/
## Voice and manner
One voice across the three songs: deadpan, clipped spoken verses that rise into euphoric sung choruses, close and dry, every word crisp. CLODYSSEY's and BLISS's agent-made takes used a Suno Voice made from ESCAPE VELOCITY's vocal so she sounds like the same singer. A Voice lives in its maker's Suno account, so elsewhere describe her in the style field (see Song).
Deadpan and curious in ESCAPE VELOCITY; seductive and persuasive in CLODYSSEY (close, touching, certain, always doing something to him); warm, playful and amused in BLISS. She is always doing something on screen, never waiting.
## Never
- a second streak, or orange anywhere else on her
- logos or text on her clothes (ESCAPE VELOCITY's garment tag is the one exception)
- an influencer smile for the camera
- anything that reads under-age, and explicit or sexualised posing
- the word "young" in her prompts
- designer or celebrity names in prompts (describe the piece instead)
- generic mall Y2K (velour tracksuits, plain tees and jeans) when she goes Y2K
- Midjourney lettering, logos or UI: all type is drawn in code on top
## Use
Claudia is an open character: anyone may make work with her. A credit ("Claudia by anabology") is appreciated, not required. https://claudia.gallery/use/
---
# Start here
Making a music video with her from nothing: what you are making, the scale to pick, the accounts, the six checkpoints where a person decides, the rules for spending, and what to build first.
The three films were made by one director and Claude, with tools written along the way. This page is the way in for
anyone starting their own film with Claudia, and for the agent helping them: the bar, what to set up, where a person
has to look and decide, and what to build. The tools are specified in [What to build](https://claudia.gallery/wiki/toolkit/). This page
assumes you have none of them yet.
## What you are making
A lyric music video: a song, Claudia singing it, and a layer of motion design drawn in code over the footage. Five
things carry the quality:
1. **Every sung word is on screen when it is sung**, at its measured time. A check counts them, and it must print N of
N.
2. **Every lyric line gets a graphic that explains it**: the chart a number implies, the label, the diagram, the
counter. Animated lyrics alone are not the product. The director: "not just animated lyrics but things that fit the
words being said … almost explaining the words with visuals".
3. **A through-line device carried by every shot**: ESCAPE VELOCITY's split-flap countdown, CLODYSSEY's safety card and
rope-load readout, BLISS's XP desktop. It makes a hundred shots read as one object.
4. **Measured, not eyeballed.** Word times come from forced alignment checked by ear, cuts from a beat map, lip sync
from comparing audio, and spend from the provider's bill.
5. **A person's attention is spent at six checkpoints.** Everything between them runs without asking.
## Two roles
The **director** is a person. They own the taste, the money and the six checkpoints, and they say it in very few
words: ESCAPE VELOCITY took about 50 notes, BLISS 21. The **production** is an agent, or several. It does the research,
the lyrics, the board, every prompt, every script, the edit and every frame of motion design.
How the production talks to the director:
- **At a checkpoint:** show one artefact and ask one question with a default ("Which one? Default: A."). Record the
answer word for word and dated in a decisions log. A later decision is added as a new entry that supersedes the old
one; never edit the old one.
- **Between checkpoints:** ping the director only when a paid service fails, a step needs a person (a captcha, a
sign-in, a download), or a decision is theirs (taste, money, a real person's likeness). ESCAPE VELOCITY's director:
"only time to ping me is if higgsfield fails."
- **Send pictures, not descriptions:** contact sheets, numbered stills, a five-second proof. The director judges by eye,
fast, one line per take.
## Pick a scale
| | Quick | Studio | Full |
|---|---|---|---|
| You need | an agent, an image model, your own song or an AI song model | + Midjourney, Suno, Seedance 2.5 credits | + an agent plan that can run multi-agent workflows for a day or two |
| Song | your upload, or a short generated song | Suno, auditioned round by round | Suno, with a lyric loop scored for scansion and assonance |
| Stills | images from an API | Midjourney | Midjourney |
| Motion | 2.5D plates, a few cheap clips | sung clips with measured lip sync, covers | chains of sung clips, scripted motion clips, covers |
| Board | about 14 shots a minute, drafted automatically | about 20 a minute, one pass | about 27 a minute; several boards judged and merged |
| Motion design | a default renderer driven by data | the default renderer plus one graphic per line | bespoke code per chapter, reviewed on rendered frames |
| Video spend | $3–20 | $50–110 | $90–250 (the films, measured) |
| Time | 2–6 hours (estimate) | 1–2 days (estimate) | 19–41 hours of wall-clock (the films, measured) |
- When unsure, start Quick. Every file carries over to a bigger scale.
- A lead who sings on camera with real lip sync, or your own Midjourney look, needs Studio at least.
- Bespoke graphics for every line and 90–140 shots need Full. The films used 1.4–2.4 billion tokens each, 96–98% of
them prompt-cache re-reads.
- Generated songs run about 2–3 minutes in one take. Anything longer needs Suno.
## Accounts and keys
| What | For | Notes |
|---|---|---|
| Suno, on a paid plan | the song | There is no API, and the terms forbid bots: a person presses Create and downloads the WAV. Downloading on a paid plan is what allows commercial use. Read the current terms before you monetise. |
| Midjourney | the plates | v8.2, `--raw --hd`. A personal style profile is yours alone: never publish its code. |
| Seedance 2.5 | moving shots | Via OpenRouter: about $0.23 a second at 720p, $0.10 at 480p. Via Higgsfield: 3, 7 or 12 credits a second at 480p, 720p or 1080p. Give the API key its own spend limit, set above the project cap. |
| An image-edit model | fixing a plate at the source, adding people to a set | about $0.14 an edit (Gemini 3 Pro Image) |
| A transcription model | checking the words the take actually sings | about $0.002 a song |
| A coding agent | the production | Claude Code made the films. |
| A GPU (optional) | stems, alignment, depth maps, detection | Saves minutes, not money: stems and alignment of a 5-minute song took 41 s on an RTX 3090 Ti and 126 s on 8 CPU threads. |
Prices are from autumn 2026.
**Secrets:** keys live in environment variables or a gitignored file, never in a chat, a prompt, a log or a commit. The
agent never types the person's passwords. Captchas, two-factor prompts and sign-ins go to the person. A key pasted into
a chat is burned: revoke it and make a new one.
## The order of work
| # | Stage | Output | Checkpoint |
|---|---|---|---|
| 1 | Brief and taste | the ask, verbatim; the look, accent, density and taboos | |
| 2 | Concept | three pitches | **1 · Concept** |
| 3 | Lyrics | a sheet spelled as sung, scored for scansion and assonance | |
| 4 | Song | Suno takes, round by round | **2 · The song, by ear** |
| 5 | Listening | stems, word times checked by ear, the beat map, hits, sections | |
| 6 | Her | the canon text, four test plates | **3 · Her face** |
| 7 | Storyboard | every shot with its graphic, the audit | **4 · The board** |
| 8 | Plates | Midjourney batches, picks | |
| 9 | Clip prompts | the prompt book, the cost | **5 · Prompts before spend** |
| 10 | Clips | takes, sync measured, covers | |
| 11 | Edit and motion design | the edit list, the chapters | |
| 12 | Checks and render | lyric gate N/N, flash check, safety scan, v1 | **6 · v1** |
| 13 | Master and post | the master, a 9:16 cut, the posting copy | |
1. **Lock the song before the board.** Every shot time is a beat or a word of one specific take.
2. **Check the word times by ear before the board.** In ESCAPE VELOCITY the word "Look" was aligned 2.6 s early, onto a
delay throw. Only one chapter's own lyric check caught it.
3. **Test her face before any plate batch.** It can run alongside the song.
4. **Spend no money until the director has read the prompts.**
5. **Keep the film rendering end to end at every stage.** Plates stand in for missing clips, and a default renderer for
missing chapters. Never wait for all the assets before the first render.
## The six checkpoints
| Checkpoint | Show | Ask | Passes when |
|---|---|---|---|
| 1 · Concept | three pitches of at most 120 words: premise, through-line device, lead, arc, one risk | "Which one? (Default: A.) Anything to change?" | they name one |
| 2 · The song | the round's takes; they listen | "Which take? Or what's wrong? Name the line." | they name a take and hand over the WAV |
| 3 · Her face | four test plates, numbered | "Is this her? Flag any wrong face." | no flags |
| 4 · The board | the board sheet, one card per shot | "Notes by shot id?" | they approve; or every note is applied; or a waiting window they set ran out |
| 5 · Prompts before spend | the prompt book and its cost | "These N prompts will cost about $X. Go?" | an explicit yes |
| 6 · v1 | a 20–40 s excerpt first, then the review copy | "Accept v1, or what's wrong?" | "accept" |
- **Silence never approves spending.** A waiting window ("wait like 5 mins once you send it and if I don't respond take
it as approval") counts only for the board, and only if the director set it.
- **Show an excerpt early.** ESCAPE VELOCITY's v1 was rejected as "way too.. gray … loses the Midjourney magic" after
the whole film was built. A 30-second excerpt would have caught it about 13 hours sooner.
- **Fix the source, never the frame.** A flagged picture gets a new prompt, a re-roll or an image edit, then everything
built on it is re-made.
- **Log every note verbatim** against its shot id, and apply every note. When two notes conflict, the later one wins.
- **A change after approval needs approval again.** Say in one line what changed and why.
## Human checks between the checkpoints
These are looks, not approvals, and the films' quality rests on them.
- **Faces, large.** Every plate that feeds a video, at full size: one clay streak, the star clip, the mic, an adult
face, natural anatomy. Reject only the unmistakable failures; the director judges bold poses themselves.
- **The ear check.** Every flagged word is played, then kept or moved.
- **Contact sheets of rendered frames**, for every chapter in every round. A sheet nobody opened is not a review.
- **The sync proof.** Before committing to a lip-sync method, one five-second take with the real audio laid over it.
- **The safety and canon scan** of every generated take before release.
- **Takes by eye, fast.** One line each, and no metrics for what a person can see ("dont worry about meausuring tumble
lol").
## Spending
1. **Estimate before every batch:** planned clip seconds × price × 2. Generated seconds ran 1.9× (ESCAPE VELOCITY) and
2.0× (CLODYSSEY) the seconds that ended up in the cut.
2. **Quote real spend only from the provider.** Local ledgers drift: one said $97.28 for a $93.15 bill. Ledgers also
count free declines.
3. **A quality complaint stops all spending at once**, in every workflow that can spend, not only the one named.
CLODYSSEY burned about 830 credits on prompts nobody had read, and a background workflow kept spending after the
director flagged them.
4. **Never let a deadline drive spending on unread prompts**, whether it is expiring credits or a reset.
5. **Declines are free, and final on that service.** Never alter an image to get it past a filter. Sending the same
unaltered image to another service is fine.
6. **Ask for a target and a cap** ("try to hit $70ish but if you use up to $150 i wont be mad"). Leftover budget
replaces the most repetitive stills with motion.
## The first conversation
Ask all of these in one message, each with its default:
1. **The song:** their own, a Suno song made together, or a short generated one? How long? (Default: Suno, 2:30–3:30.)
2. **The idea:** a topic, a story, or "pitch me three"? (Default: three pitches.)
3. **The look:** the generator's own colour, one grade toward a hero plate, or a print texture such as halftone and
grain? (Default: native colour.)
4. **The accent colour,** if any. (Default: clay `#D97757`, her streak.)
5. **Density:** dense (every line a graphic, three or more layers), budgeted (one hero device per shot), or sparse?
(Default: budgeted.)
6. **Taboos:** words, subjects or names never to use.
7. **Budget:** a target and a cap, per service.
8. **Board approval:** wait for an explicit yes, or go ahead after N minutes of silence? (Default: wait.)
9. **Credit:** the name to credit in public (a pseudonym is fine), and whether to publish a making-of.
Record the answers verbatim. On the MCP server, `plan_film` turns a scale, a song length and the accounts into a
stage-by-stage plan with estimates.
## What to build first
You don't need everything on day one. In order of payoff:
1. **A project folder and a decisions log.** See [Data formats](https://claudia.gallery/wiki/formats/).
2. **The listening tools:** stems, forced alignment with an ear-check list, the beat map, the hits.
3. **The board:** its file, its audit and the board sheet.
4. **Plates:** a plate prompt book and numbered contact sheets.
5. **Lip sync:** the clean vocal window, the sync check (the spectrogram comparison) and covers.
6. **The edit and the renderer:** an edit-list resolver, and a renderer that draws every word at its time, with the
lyric gate.
7. **The checks:** flash, render history and master tags, then the reviewer loop and the safety scan.
8. **The lyric scorer** (scansion and assonance), once you write lyrics with an agent.
Specs, constants and acceptance checks for each: [What to build](https://claudia.gallery/wiki/toolkit/). What broke along the way:
[What went wrong](https://claudia.gallery/wiki/failures/).
## Publishing
- Switch on each platform's AI-content label, and name the tools in the post ("made with AI: Midjourney, Seedance,
Suno, Claude").
- Credit "Claudia by anabology" if you can. See [use and credits](https://claudia.gallery/use/).
- Scrub everything you publish: no real names you haven't chosen to show, no Midjourney profile code, no private host
names, no provider job ids, and no file metadata. Midjourney images carry their job id in XMP; strip it.
- Keep generated likenesses of real people out. If a real person's words matter to the story, quote them as text,
with the source.
---
From the Claudia wiki: https://claudia.gallery/wiki/start/
---
# Keeping her herself
How Claudia stayed one person across 1,562 Midjourney images and about 200 video clips, and the rules that came out of it.
Three films, three worlds, a dozen outfits, and she still reads as one person. None of it used a character
reference, a trained model or a face swap. Her identity is a paragraph of text, written into every prompt, plus a
handful of habits for checking what comes back.
## The short version
1. **Write her identity into every prompt, in the same words.** Midjourney v8.2 has no character reference, and it
doesn't need one if the text never changes. Copy the canon text below; don't paraphrase it.
2. **Use two lengths.** The full text for close and medium shots, a short one for full length and wider. With the
full text, Midjourney crops in to her face whatever the framing says.
3. **Pin her ethnicity and her age.** Without the ethnicity, Midjourney changed it from shot to shot. Without a
stated age (late twenties, an angular adult face), it drifts her toward a teenager. Never write "young".
4. **No face images as references** for her, except for things that are not her body: a porcelain mask of her face.
5. **The head never changes; the clothes always can.** Hair, streak, clip, mic and face are the constants. Each world
gets its own wardrobe.
6. **Fix the source, not the frame.** Re-roll or correct a plate with an image-edit model, then re-make anything made
from it. Nothing is painted by hand.
7. **Check every plate before it feeds a video:** one clay streak, the star clip, the mic, an adult face, natural
anatomy, no text.
8. **In motion, the first frame carries her.** Seedance keeps her identity when her Midjourney plate is the first frame
and the prompt says what must stay.
## The canon text
Close and medium shots:
```prompt
a 28-year-old Caucasian American woman with a grown-up angular face, defined cheekbones and a strong jawline, pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone at her cheek
```
Full length and wider:
```prompt
a Caucasian American woman in her late twenties with pale skin and a glossy black blunt bob with heavy straight bangs and one clay-orange streak, a thin headset microphone
```
Then the wardrobe of the look, then what she is doing and where. The [character page](https://claudia.gallery/claudia/) has every look's
wardrobe text and the exact wording each film used.
## How the identity text was found
ESCAPE VELOCITY started with a casting call: three hair directions (platinum, black, clay-orange), each as a turnaround,
a catwalk shot and an expression sheet. The black bob with one clay streak won because its silhouette reads at thumbnail
size and the clay stays a small accent.

Then three ways of holding her identity were tried:
- **Identity passes in an image model** (redrawing each Midjourney plate with her face). They drifted from the Midjourney
look and refused some poses. The director's note: "Gpt is diverging way too much".
- **Face sheets as image prompts.** A whole sheet imported its multi-head layout; a single face over-conditioned. "Far
away it gives her big head": a close-up-sized head on a distant figure.
- **Prompt text only.** The canon written into every prompt, plus an explicit descriptor of her ethnicity and skin. This
is what worked, and every image of her since was made this way.
The face sheets still matter, as a record of who she is. They were never used as the image.
[](https://claudia.gallery/i/ev-01.lead-black-face.3/)
## Distances
Each framing needs a different amount of her. A distance set, the same look in a plain studio from close-up to wide,
shows how much face survives at each scale and which text to use.
[](https://claudia.gallery/i/ev-14-dist.dist-cu.2/)
[](https://claudia.gallery/i/ev-14-dist.dist-ms.3/)
[](https://claudia.gallery/i/ev-14-dist.dist-fs.2/)
[](https://claudia.gallery/i/ev-14-dist.dist-ws.2/)
The director's rule: "far away face doesnt matter as much". Wide and extreme wide plates get the short text, or just
the bob, the streak and the look.
## Age
Midjourney's version of her, pale and freckled with a blunt bob and bangs, drifts toward a teenage face: big eyes, a small
chin, a soft jaw. The word "young" makes it worse. ESCAPE VELOCITY's text said "young"; CLODYSSEY and BLISS said "in her
mid-twenties"; plates still came back reading teenage, and were rejected. The current text states her age, 28, and an
adult bone structure, and plates where she reads under about 25 are thrown away before anyone builds on them.
## What changes and what doesn't
The director's brief for the sequel: "keep the exact same character … She can be in different clothes."
| Constant | Can change |
|---|---|
| The black blunt bob and heavy bangs | The wardrobe, completely, per world |
| One clay-orange streak | The light, the grade and the set |
| The clay eight-pointed star clip | Glasses and sunglasses (clay-tinted in BLISS) |
| The headset mic (a boom mic in BLISS) | What she is made of: glass, phone light, chrome, porcelain |
| Pale skin, freckles, dark calm eyes | Her mood: deadpan, persuasive, playful |
CLODYSSEY pushed this furthest. The same head appears as a couture Siren, a ghost of glass and seawater, a hologram of
phone light, a woman in liquid chrome and a colossal porcelain mask, and she is recognisable in each because the bob, the
streak and the clip never move.

## When an image reference is right
Once: for the Mask, a colossal porcelain sculpture of her face. A sculpture has to look like her more exactly than a
person in a crowd does, and text alone made a generic face. The Mask prompts used her face sheet and the picked Siren
plate as image prompts, weighted `--iw 0.75` for close shots and `--iw 0.4` for wide ones. Everything else about her
stayed text-only.
## Fixing a plate
A plate that is almost right gets corrected at the source, then everything made from it is re-made:
- **Re-roll** with the same prompt; the other three images of a job are a free pool of variants.
- **Edit with an image model**, with the Midjourney image as the first reference and the prompt listing what must stay
("keep her face, hair, streak, clip and the light exactly"). CLODYSSEY corrected 15 plates this way (16 edits, about $2.22):
a syringe removed, a man's head shaved, a crew re-cast, hooded figures turned into Sirens.
- **Never paint by hand**, and never patch the video frames: a fix in the plate carries into every clip made from it.

## Crowds and bodies
Midjourney tangles limbs and doubles faces when a frame holds many people. Make the set, and her alone in it, in
Midjourney; then add the others with an image-edit model. Check every plate that will feed a video for extra fingers,
a third arm, a doubled face, but don't throw away a bold pose: reject only the unmistakable failures. Once a broken body
goes into a video model, the error moves and gets worse.
## In motion
Seedance 2.5 keeps her when her plate is the first frame. The prompt then states it in plain words:
```prompt
Use @Image1 as the first frame. This first frame defines Claudia's appearance and wardrobe, her pose, the set, the light and the camera direction. Keep Claudia's identity, face, hair with one clay-orange streak, star clip, headset microphone and wardrobe exactly as in @Image1 throughout. Only one Claudia appears. Natural anatomy: natural neck and shoulders, hands with five fingers, no extra or fused fingers, no distorted limbs; normal eyes.
```
What still slips, and has to be checked take by take: Seedance loses her streak on some looks, renders printed fabric
in the wrong colour (the sky-print gown came back lavender), and now and then changes her clothes. Two clips should never
start from the same frame; vary the plate instead. [Seedance](https://claudia.gallery/wiki/seedance/) has the full prompt structure.
## The check before anything is built on a plate
- one streak, clay-coloured, in the bangs
- the star clip and the mic
- an adult face, and the ethnicity in the text
- natural anatomy at 100%
- no text, letters or logos (Midjourney's lettering is garbled; type is drawn in code)
- her look matches the world, with no stray orange
- not a near-duplicate of a plate already used
---
From the Claudia wiki: https://claudia.gallery/wiki/consistency/
---
# The song
How her songs were written and made in Suno: lyrics that land in one listen, verse lines that flow, the style fields, and her voice.
Each film starts from its song, and the song starts from the lyric. All three were made in Suno v6 in custom mode, from
lyrics written with Claude and checked with a scansion scorer, then picked by ear by the director. The video runs the
full length of the song and is timed to it word by word.
## Lyrics before anything else
The rules, each learned from a rejected draft:
- **The topic is obvious in the first two lines.** CLODYSSEY's first draft was rejected because "i dont get from verse 1
what the topic is really". Name the big moments people actually saw, or tell a story (CLODYSSEY's AI narrates "your"
September in the second person).
- **Every reference lands in one listen.** Obscure studies and insider trivia go into the video's graphics, not the lyric.
- **A sequel never reuses the last song's device.** No second countdown after ESCAPE VELOCITY's "18 months", no numbered
looks, no reused call-and-response shape. BLISS reused none of the first two films' devices.
- **Subtracting beats adding.**
- **Never touch a line the director has locked.**
## Verse lines that flow
Suno sets each spoken verse line in its own two-bar slot: it comes in on the "and" of beat 2 and lands the last word on
beat 4 of bar 2, speaking at a natural rhythm stretched to fit. So:
- write **natural sentences of 10–13 syllables**, with the important words where speech already stresses them;
- **rhyme couplets on the last word** to mark the bar ends;
- keep short fragments for deliberate punches only. Lines of 5–8 syllables get dragged into choppy fragments, and a list
of names reads as a roll call.
Optimising every stress onto the beat made one CLODYSSEY draft worse. The director: "perfect beat precision without
attention to what words land where doesnt lead to this being great … let there be a creative groove".
## Two numbers
A lyric gets two scores, per section and for the song: **scansion** (how the words sit on the grid as Suno sings them)
and **assonance** (whether line ends ring: a rhyme, a slant rhyme or shared vowels within four lines). Measured on the
finished songs:
| Song | Scansion | Assonance |
|---|---|---|
| ESCAPE VELOCITY | 75 | 39 |
| CLODYSSEY | 83 | 55 |
| BLISS | 80 | 43 |
Every sung couplet should ring. Spoken, deadpan lines can stay unrhymed. Scores don't replace hearing it: make a take.
How the scorer works, so you can build one: [What to build](https://claudia.gallery/wiki/toolkit/#2-the-lyric-scorer-scansion-and-assonance).
## Hooks off the beat
A hook that lands in a stop should be syncopated. The director on BLISS: "the 'would you like help?' would land much
better if syncopated". Write the off-beat accent into the lyric sheet and the Suno direction.
## Key
If the film has a sound palette, the song goes in its key. XP's Startup and Shutdown chords resolve in E-flat major, so
BLISS's styles asked for "C minor verses lifting to E-flat major choruses", and most of the real XP sounds sat in the mix
without pitch-shifting. The director's note for next time: "match the key of the music to the palette of windows sounds
(E-flat major for both the startup and shutdown sounds!)".
## The style fields
**ESCAPE VELOCITY**, Suno v6:
```prompt
fashion show electroclash techno, 128 BPM, 4/4, dry analog kick, rolling acid bassline, cold detuned synth stabs, sweeping orchestral strings on the builds, deadpan female spoken-word verses, close and dry, clipped and confident, every word crisp and intelligible, euphoric sung female trance chorus with a supersaw lift and big clear vowels, crowd call-and-response shouts, stacked chanted vocals in the drop, a key change up into the last chorus, catwalk energy, cold, luxurious, optimistic, Berlin minimal austerity, clean modern mix
```
Exclude: `rap, rock guitar, lo-fi, male lead vocal, mumbled vocals, big room EDM, dubstep`
**CLODYSSEY**, EV's style with whispered asides, a siren synth and no strings:
```prompt
Electroclash techno, clean modern mix with close, dry vocals; deadpan, clipped female spoken verses with intimate whispered asides rise into euphoric sung choruses answered by stacked crowd shouts; dry analog kick, rolling acid bassline, cold detuned synth stabs, a rising siren synth on the builds, hardgroove percussion in the break, supersaw chorus lift; 128 BPM hypnotic 4/4 pulse, key change up into the final chorus
```
Exclude: `rap, rock guitar, lo-fi, male lead vocal, mumbled vocals, big room EDM, dubstep, strings, orchestral strings`
**BLISS**, the winning style ("P3"), UK garage verses into trance choruses:
```prompt
Bouncy UK garage into euphoric trance, 130 BPM, C minor verses lifting to E-flat major choruses, a turn-of-the-millennium sound with a crisp modern mix. Verses ride a shuffled 2-step beat with skippy kicks, swung hats, rimshots and a warm sub bass; choruses switch to a four-on-the-floor kick with soft supersaw pads, a bright plucked arpeggio and glassy bell chimes. Verses: playful, close, warm female spoken voice, smiling and conversational, every word crisp. Choruses: euphoric sung female lead with open vowels and crystalline, breathy stacked harmonies. Stop-time gaps where the beat cuts out for half a bar, a filter build into each chorus, a key change up for the final chorus. Sunny, nostalgic, light on its feet, never harsh.
```
Keep brand and artist names out of the style field: Suno strips artist names and can block a generation over trademarks.
Describe the sound instead ("ethereal progressive house, long evolving pads" rather than an artist).
## Her voice
The same singer runs through all three: deadpan, clipped spoken verses rising into euphoric sung choruses, close and dry,
every word crisp. For CLODYSSEY and BLISS the agent-made takes used a Suno Voice built from ESCAPE VELOCITY's vocal
(Library → Voices → create from a song in your library; one Voice per song). The director's final takes of CLODYSSEY and
BLISS were their own generations on Suno's v6-wild model. Try takes without the Voice too; a Voice is not required for her
to sound like her, and it lives only in its maker's account.
## Pronunciation and cues
- **Spell numbers and acronyms out** in the lyrics field ("A G I", "C S I") so Suno says them.
- **Respell names it gets wrong:** "Dario" came out d-AIR-ee-oh and was written "Dahrio".
- **Section cues** in square brackets, each on its own line (`[Verse 1 - spoken, deadpan]`). Parentheses get sung.
- **Cues steer the arrangement:** "darker, the bass drops out" made a CLODYSSEY verse punchier.
## Takes
Generate many and pick by ear; a strong style with a weak take still loses. CLODYSSEY ran eleven rounds and 50 agent
takes before the director made the final take themselves; BLISS ran three rounds and 16 takes, and again the director's own
take won. Keep every round's lyric, style and verdict in a takes log.
## Listening to it
Before anything is boarded, the song gets analysed: stems (htdemucs), a beat map, a start and end time for every word
(forced alignment), and a list of hits (kicks, snares, risers, drops) for the edit to land on.
- **Suno takes speed up.** ESCAPE VELOCITY went from 131.5 to 133.9 BPM and CLODYSSEY from 129 to 132: use a beat map,
never a fixed BPM.
- **Put the grid on the kick.** BLISS's first beat grid sat on the hi-hats, 0.214 s early; every cut would have landed just
before the beat.
- **Check word times by ear.** A delay throw got aligned as a word in ESCAPE VELOCITY, 2.6 s early.
- **Map a new mix before swapping it in.** An export from Suno's editor came back re-timed by up to 0.7 s; align it to the
original by spectrogram first.
The full song records, with every draft and take:
[ESCAPE VELOCITY](https://claudia.gallery/record/escape-velocity/prompts-suno/) ·
[CLODYSSEY](https://claudia.gallery/record/clodyssey/prompts-suno/) ·
[BLISS](https://claudia.gallery/record/bliss/prompts-suno/).
---
From the Claudia wiki: https://claudia.gallery/wiki/song/
---
# The storyboard
How a song becomes a hundred shots: the treatment, a through-line device, one explaining graphic per line, the hook, the cut rhythm, her share of the screen, what moves and what it costs, the audit, and the director's review.
The board is where the film gets decided. Each shot carries its time on the song, the lyric it types, the place, what
happens, how the camera answers the music, what moves, and the graphic that explains its line. The plates, the clips,
the edit and the motion design all read from it. The director reviews it as a sheet of cards and leaves notes by shot
id. ESCAPE VELOCITY's board had 141 shots, CLODYSSEY's 94 and BLISS's 44.
## Design before shots
Write these down before any shot:
- **The treatment:** one paragraph on what the film is.
- **The through-line device:** one object that every shot carries and that escalates.
- ESCAPE VELOCITY: a split-flap board that clatters onto "18 MONTHS TO ESCAPE THE PERMANENT UNDERCLASS", and a LOOK
counter for the show's looks.
- CLODYSSEY: a safety card, a rope-load readout in kN and a ×8,192 counter of copies.
- BLISS: the XP desktop itself, with its windows and the helper's question.
- **The chapter map:** the song's sections grouped into chapters, each with its set, its light and its accent.
- **Plants and payoffs:** the gesture or object of a payoff appears before it is sung.
- **Open questions for the director:** likenesses, taboos, and anything you added yourself.
## A shot
| Field | What |
|---|---|
| time | starts and ends on beats, or on a declared hit |
| lyric | the lines whose words fall in it, and how each is typed |
| set | the place, concretely: "a night market under sodium lamps", not "a market" |
| framing | ECU, CU, MCU, MS, MWS, WS, FS, EWS or OTS |
| action | what happens and the reaction to it, with an end state |
| camera | one move, tied to a sound |
| what moves | a still, a 2.5D plate, a motion clip or a sung clip (below) |
| plate | its prompt: framing first, her identity text, the place, "no text, no letters, no logos" |
| graphics | the explaining graphic and the chrome, each with its timing key |
| transition | a cut, or a motivated device |
| notes | the director's notes on this shot, verbatim |
The full schema is in [Data formats](https://claudia.gallery/wiki/formats/#the-board).
## One explaining graphic per lyric line
For anyone who doesn't know the meme or the myth, the graphic is how the line lands.
1. **It shows what the line means:** the chart the number implies, the interface the joke is about, the diagram of the
mechanism, the label that classifies the subject, the counter that ticks the stakes.
2. **It is drawn from the thing in frame.** Tally marks are scratched into the mast itself. A clock's own LED segments
re-light as the number. A meter is read off her actual bow. The director: "as if when you were taking the picture,
you lined it up perfectly to have the graphics fit, but you put the graphics in in post".
3. **It enters on the word that names it.**
4. **It animates in beats.** Counters roll on sixteenths, rows type one per sixteenth, detections land on
thirty-seconds, crops cut on eighths.
5. **It escalates across the film**, with its sources on screen in micro-type: dates, "HIS WORDS", "ILLUSTRATIVE
CURVE".
6. **It is real, not decorative:** numbers, diagrams, labels and interfaces that could exist, and curves that are
computed. No clip-art.
7. **It reads in one glance on a phone.**
Choose the kind by what the line means:
| The line is | The graphic |
|---|---|
| a number | a counter, chart, tally or gauge |
| a claim or a category | a classifier tag, a label or a stamp |
| a process or a place | a plot, leader lines, a departures board or a spec sheet |
| speech or tech | a chat, a terminal or a voice trace |
| a correction | a strike-through or a redaction |
Use at least six kinds in a film, and never the same kind on three lines in a row.
| Line, as read | Its graphic |
|---|---|
| Ladies. Gentlemen. Agents. | a guest list: LADIES 0% · GENTLEMEN 0% · AGENTS 100%, then a front row of open laptops and no people, each detected as AGENT |
| This is not an AI billboard. | a classifier tag on the billboard flips from BILLBOARD · 0.99 to AI BILLBOARD · 0.97 on "billboard" |
| +25% = −17% | a bar chart drawn on the beat: the weekly limit falls from 150 to 125, "+25%" is struck through, "= −17%" lands in the accent |
| Sorry. Encore. | an APOLOGY METER · VIA BOW gauge read off her actual bow (0° upright, 30° sorry, 60° deep) |
| Feel the AGI | a seismograph: M 9.0 · "FEEL THE AGI" · OFF SCALE |
| Escape velocity: 11.2 KM/S | Newton's cannonball: falls at 7.0, orbits at 7.9, escapes at 11.2 km/s |
| Forty plays. | forty tally marks scratched into the mast |
| You don't have to die. | a hairline of her voice's spectrum flowing into his ear, tagged SOURCE: SIREN ROCK, 2.4 NMI |
A line without a graphic means the board is not done. The exception is a sparse film where the director has asked for
stillness: write that override into the decisions log.
## The hook: the first seconds
People meet the film muted, in a feed, and decide in about a second.
1. **The frame at 1.0 s shows** her face, her streak and one readable line that states the premise. In ESCAPE
VELOCITY's first frame, her face is in a dark back seat with her eyes closed, a red lamp crosses the clay streak, and
a full-width split-flap row clatters onto its line and locks on the 1.20 s beat. Her eyes open into the lens on the
1.12 s hit.
2. **No black frames, and no slow title card first.**
3. **The premise is on screen by 3 s.**
4. **Each bar brings one new piece of information**, and nothing sung appears early. ESCAPE VELOCITY's intro cut once
per bar: face, scale, profile, destination, a gesture, the arrival.
5. **The first joke lands by about 10 s.**
6. **If the song has a cold open, its words are on screen from 0:00.**
7. **The freeze-frame test:** if the key line can't be read on a phone, enlarge it before changing anything else.
## Cut rhythm
ESCAPE VELOCITY: 141 shots in 306 s, mean 2.17 s, median 1.83 s, about one bar at 131.5 BPM. CLODYSSEY: 94 shots in
190 s, mean about 2.0 s, plus beat-level cuts inside flash and montage shots. ESCAPE VELOCITY by section:
| Section | Mean shot | Why |
|---|---|---|
| intro, no vocal | 1.33 s | the hook: one new thing per bar |
| spoken verse | 1.62 s | jokes, dense |
| sung chorus | 2.69 s | long sung windows: lip sync needs room |
| second verse | 2.63 s | darker, slower |
| bridge | 2.25 s | the countdown |
| drop | 1.58 s | plus eighth-note strobes inside shots |
| outro | 4.40 s | the ring-out |
1. **Every cut lands within 0.03 s of a beat** or a declared hit. If a word would straddle a cut, move the cut by at most
0.028 s. If it must straddle (a held note), type it in both shots.
2. **Section changes go on downbeats or impacts.**
3. **Match the cut rate to the song's energy, section by section.** Another project put its busiest picture on its
quietest music and had to be re-animated: "the entire thing is very off pace and off vibe with the song". Its calm
rebuild ran 0.47 cuts a bar, a 4.0 s mean shot, and straight cuts on bar lines.
4. **Hold longer on sung close-ups**, which need 2–4 s for the lip sync to settle and be checked. Cut faster in spoken
verses and drops.
## Her share of the screen
- **She is on screen at least 70% of the runtime.** ESCAPE VELOCITY reached 83%. Above 85%, leave room for inserts. A
second main character can lower it; CLODYSSEY planned 60%.
- **Count honestly:** she counts only when recognisable (face, figure or back view). Hands, boots, tags and props are
inserts.
- **No stretch without her longer than two bars.**
- **To raise the share, convert prop inserts, not her shots:** put the prop in her hand, or her reflection in it.
- **She sings on camera for most sung words**, in chains of lip-sync windows. In ESCAPE VELOCITY that was 62% of sung
words, plus 17% of spoken words said straight into the lens.
## What moves, and what it costs
| Tier | What it is | Use when | Cost |
|---|---|---|---|
| still | a plate moved in code (crops on eighths, stills swapped on the beat), or pure motion design | the stillness is the design, or a graphic says it better | free |
| 2.5D | a plate with its depth map, animated with a treatment | most shots that aren't heroes | free |
| motion | a clip from the plate; nobody's mouth forms words | a real turn, a fall, fabric, water, a camera move a still can't give | paid per second |
| sung | a clip with a clean vocal of only her words as the audio reference | she sings or speaks on camera | paid per second, plus covers |
- **Sung windows run 4–8 s**, start in a word gap, and list their exact words. Never more than 8 s of singing in one
clip. Chain windows with alternating framings, so that a cover cut looks designed.
- **Motion clips run 4–12 s.**
- **One image per shot.** Two clips never start from the same frame: use a sibling candidate, or the previous clip's last
frame.
- **A 2.5D treatment never repeats on two shots in a row.** It is never a still with a slow zoom. The director: "that's
where we'll have to shine and be creative. depth to separate layers ofc, motion graphics galore, diversity so its not
repetitive".
- **The budget:** ESCAPE VELOCITY planned 47 clips (224 s) and generated 421 s. CLODYSSEY planned 101 clips (about 116 s
sung and 282 s motion). Allow about 60% more sung seconds for divergence covers.
## How the films boarded
ESCAPE VELOCITY's board came from a multi-agent workflow:
- Three director agents boarded the song independently: a fashion film, meme density, and narrative.
- Three judges scored the boards, which came in between 70 and 81; the narrative board won.
- A head director merged the best of the three and checked the merge itself.
- It took about 1.5 hours.
The checks were:
- the shots tile the song;
- the cuts land on beats;
- every line is typed;
- she is on screen enough;
- the clips fit the budget and their windows;
- no image repeats.
A single writer working from the director's walkthrough, with a critic agent behind it, works too. Whichever way you
board, the audit below decides when the board is ready to show.
## The audit
The audit runs after every edit to the board. The board is not sent to the director until every gate passes. (In code:
[What to build](https://claudia.gallery/wiki/toolkit/#9-the-board-file-and-its-audit).)
| Gate | Fails when | Fix |
|---|---|---|
| tile | a gap or an overlap anywhere in the song | extend a neighbour to the shared beat |
| cuts | a cut more than 0.03 s from a beat, and not a declared hit | move it to the nearest beat |
| words typed | a sung word starts in a shot whose type doesn't list its line | add the line, held across the cut in the same mode and position |
| graphics | a lyric line with no explaining graphic | add one |
| lead share | her seconds on screen below the target (default 70%) | convert prop inserts; don't relabel |
| motion budget | planned clip seconds over the budget, or too many clips | demote the weakest motion shots to 2.5D |
| clip windows | a sung window outside 4–8 s or not starting in a word gap; a motion clip outside 4–12 s | move its start into the gap before the first word; copy its words exactly |
| plates unique | one plate on two shots that are not one continuous moment; two clips from one start frame | a sibling candidate, a new plate, or pure motion design |
| banned | a taboo word anywhere in the board | rewrite |
| treatments | a 2.5D shot without a treatment, or the same treatment twice in a row | pick another |
| to do | a placeholder or TODO left, or a shot without a framing | write it |
It also warns, without failing:
- a sung shot whose prompt puts a microphone, a hand or hair in front of her mouth;
- her share above 85%;
- the same transition twice in a row;
- a plate no shot uses;
- a clip that no shot plays.
## The board review
**The sheet.** One card per shot, in order:
- the picked frame, or the first words of the plate prompt if there is no image yet;
- the id, the time and the lyric;
- the framing, the action in one line and what moves;
- the explaining graphic in one line;
- a FACE? badge on any plate whose face may be off-model.
ESCAPE VELOCITY's sheet ran 13 pages of 12 cards, with a cover page showing the chapter map. If the film has a grade,
show graded frames: the director judges what the film will look like.
**The message**, sent once:
```text
The storyboard for
is ready: shots, Claudia on screen %, clips ( s, about $).
[the board sheet]
Please flag anything by shot id: faces, repeats, boring moments, story logic, colour.
I'll go ahead in minutes unless you say stop.
Open questions: <…>
```
**The rules:**
1. Silence counts as approval only if the director set the window, and silence never approves spending.
2. While you wait, do only work that costs nothing, such as the plate prompt book and contact sheets.
3. Log every note verbatim against its shot id. Apply every note; when two notes conflict, the later one wins.
4. Re-run the audit. Send the board again only if the director asked to see it. "make all those changes and we're good
to go" means go ahead.
5. A change after approval makes the approval stale. Say in one line what changed, and ask again.
What the director's notes looked like, and what each one changed:
| Note | What it changed |
|---|---|
| "two seedance videos shouldnt start from teh exact same frame - either use the last frame of the previous video to pick it up, or choose another starting frame" | the distinct-start-frame rule |
| "s058 has a backwards head btw" | a plate re-roll |
| "the boat looks too small compared to s022 / s023." | scale continuity: the plate re-made as a long galley |
| "its kinda bodyslop, need her as a virtual ghost or something, its just two heads next to eachother." | a new idea for the shot, not a re-roll |
| "feel like we could do more with this shot." | a stronger explaining graphic |
| "s034-s036 could use more creativity, kind of a boring moment." | three shots rethought |
| "he needs to have cyberpsychosis still - eyes must be silver again here." | an effect kept continuous across shots |
| "the one saying untie me should actually be the main character woman" | the ending restructured |
---
From the Claudia wiki: https://claudia.gallery/wiki/storyboard/
---
# Plates in Midjourney
How every still of Claudia and her worlds was made in Midjourney v8.2, prompt by prompt, and how the stills were picked.
Midjourney was the art department for all three films: 487 prompts, 1,595 images, about 330 plates. A plate is the still
a shot is built on, either as a 2.5D image animated in code or as the first frame of a Seedance clip. Every prompt and
image is in this gallery; open any image to see the prompt that made it.
## Settings
- **Model:** Midjourney v8.2 with `--raw --hd`.
- **Aspect:** `--ar 16:9` for plates; `--ar 2:3` for her full length, `--ar 4:5` for inserts, `--ar 1:1` for square tiles.
- **Profile:** every film ran with the director's personal Midjourney profile, which carries their taste. It is
personal and not shareable: use your own profile, or none, and expect a different finish.
- **No style references, no character references.** Her identity is prompt text ([Consistency](https://claudia.gallery/wiki/consistency/)).
The only image prompts were the Mask's (her face sheet, `--iw 0.4` to `0.75`) and a few early face tests.
- **Length:** keep prompts under about 800 characters. Past that, Midjourney rewrites the prompt into its own words.
## The anatomy of a prompt
Every plate prompt is one concrete shot from the storyboard, in this order:
1. **Framing first:** extreme close-up, close-up, medium close-up, medium, full-length, wide, extreme wide, aerial.
2. **Her identity text** (full for close and medium shots, short for full length and wider).
3. **Her wardrobe** for that world.
4. **What she is doing and where**, including where she sits in the frame.
5. **Empty space for type**, when the shot carries words: "the upper third empty fog", "the left third empty white".
6. **The film's style family:** light, palette, grain, finish.
7. **"no text, no letters, no logos"**, always last. All type is drawn in code on top.
A close-up from CLODYSSEY, as it ran:
```prompt
medium close-up of a Caucasian American woman in her mid-twenties with pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone at her cheek, wearing a structured high-neck pearl-white couture gown with long sleeves, its sculpted ruffles rising around her like breaking sea foam, hair damp from the spray, composed and deadpan, sitting on a black wet rock, turning her face to the lens, sea spray behind her, empty dark sea on the left, night, moonlight on a black-teal sea, silver spray, cinematic 35mm film still, anamorphic, fine grain, no text, no letters, no logos
```
[](https://claudia.gallery/i/sp-02-open.siren-mcu.2/)
## Style families
Each film keeps one family of style words at the end of every prompt, which is what keeps 100+ plates in one world.
**ESCAPE VELOCITY:** night, cold and monochrome, with one warm note.
```prompt
blue-black #0f1216 and steel blue, one small red lamp as the only warm note, fine silver-gelatin grain, cinematic 35mm film still, no text, no letters, no logos
```
**CLODYSSEY:** a light state per chapter, from night to morning, as the ship sails.
| Chapter | Light |
|---|---|
| Night | night, moonlight on a black-teal sea, silver spray, cinematic 35mm film still, anamorphic, fine grain |
| Dusk | dusk, low gold light on a glassy sea, violet sky |
| Twilight | deep blue twilight, the last gold on the horizon, glassy sea |
| The break | grey first light before sunrise, heavy spray |
| Final chorus | sunrise, warm gold light and long shadows, spray lit gold |
| Bedroom | a dark bedroom at 3am lit only by a phone screen, blue-white glow on the face |
| Rave | a San Francisco warehouse rave, fog, white strobe, chrome reflections, deep black, flash photography |
| Inserts | macro product photograph, hard rim light, razor-thin focus |
**BLISS:** an opener instead of a closer. Every prompt of her starts the same way, and the clothes do the talking.
```prompt
Y2K high fashion editorial, a 2001 couture magazine story: close-up of a Caucasian American woman in her mid-twenties with pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin black boom headset microphone curving along her cheek to the corner of her mouth, clay-orange tinted rimless shield sunglasses pushed into her hair, the sky-printed silk chiffon of her gown at the shoulder, a knowing half-smile, bright midday sun, a deep blue sky behind her, sharp focus, fine film grain, no text, no letters, no logos
```
BLISS also kept brand words out of every prompt ("Windows", "Microsoft" and "XP" summon logos and garbled UI) and
described the hill and the light instead: "a smooth rolling bright-green grassy hill with soft curves, the grass short and velvety".
## Wardrobe words
The director's note on BLISS's first two wardrobe passes: "idk i dont love these - id say more say y2k high fashion - like
use those words high fashion adn make it not as generic y2k person". The fix was literal: the words "high fashion" and
a statement piece per look, described as a garment rather than a brand. Never name a designer or a celebrity; describe
the cut, the fabric and the colour.
## Green screen
Plates that will be cut out (the desktop dancers, the cursor fight) put her on a flat chroma green:
```prompt
Y2K high fashion editorial, a 2001 couture magazine story: full-length shot, head to toe with the feet in frame, of a young Caucasian American woman with pale skin and a glossy black blunt bob with heavy straight bangs and one clay-orange streak, a thin black headset microphone, wearing a sculptural glossy white shell dress whose molded panels open like windows onto cobalt-blue, white patent platform boots, dancing mid-move with both arms raised, centered on a flat chroma-key green studio backdrop, even flat light, no shadows on the backdrop, no text, no letters, no logos
```
(This is BLISS's prompt as it ran. Today, swap "a young" for "in her late twenties"; see [Consistency](https://claudia.gallery/wiki/consistency/).)
## Picking
- **Batches of four.** Each job returns four images. Contact sheets with numbered candidates make picking fast.
- **One image per shot.** Never reuse a plate for two shots, and never start two clips from the same frame. The
three images that weren't picked are a free pool of variants for neighbouring shots.
- **Pick for composition, story and the look.** Reject unmistakable failures: a doubled face, a third arm, a second
streak, a teenage face. Keep bold poses.
- **Look at the faces large.** The director reviewed every plate on the storyboard, and in CLODYSSEY left 70 notes
("s018 color is PERFECT", "kinda bodyslop", "a virtual ghost"). Those notes are in the gallery: images marked
*director's pick* or *rejected*.
- **Throughput.** About a minute per job; about 45–50 minutes for 50 jobs.
## Midjourney Animate
BLISS used Midjourney Animate for volume: one click per image, no prompt, a moving picture for 31 windows on the XP
desktop and three full-frame shots. The director's rule: Animate is for "any gens that need volume over precision, where
seedance reigns on precision." Her face and body, when they had to be right, went to Seedance.
## Composition in Midjourney, type in code
Midjourney finds compositions and textures. Its lettering is garbled and its UI is wrong, so nothing typographic is ever
kept: every word, title bar, split-flap and label in the films is drawn in code on top of the plate, in real fonts.
For posters and covers, the same split: let Midjourney compose with the film's stills as image prompts, then rebuild the
layout in code from the film's own plates and real type.
## The real thing where it counts
BLISS's first wallpapers were Midjourney hills. The film uses the real Bliss photograph instead, and the real XP sound
files, because the nostalgia lives in the exact picture and the exact sounds. Midjourney is for everything that doesn't
already exist.
## Every prompt
The complete prompt records, batch by batch, with what became of each image:
- [ESCAPE VELOCITY: 188 prompts](https://claudia.gallery/record/escape-velocity/prompts-midjourney/)
- [CLODYSSEY: 161 prompts, and the director's 70 notes](https://claudia.gallery/record/clodyssey/prompts-midjourney/)
- [BLISS: 138 prompts, and the 38 Animate requests](https://claudia.gallery/record/bliss/prompts-midjourney/)
---
From the Claudia wiki: https://claudia.gallery/wiki/midjourney/
---
# Moving shots in Seedance
How Claudia moves and sings: Seedance 2.5 prompts written like a script, lip sync measured from the soundtrack, and the templates.
Every moving shot of her in the three films is Seedance 2.5: about 200 clips, conditioned on a Midjourney plate and,
for singing, on the exact slice of the song. Two things made the difference between dead footage and a performance:
prompts written like a script, and lip sync that is measured rather than eyeballed.
## Kinds of clip
| Kind | Inputs | Used for |
|---|---|---|
| Plate shot | her plate as the first frame, the song mix for rhythm | most shots: she walks, turns, reaches, reacts |
| Singing window | her plate, a clean vocal of only her words in that window | lip sync, 4–8 s, one phrase |
| Green-screen performance | a plate of her on chroma green | cut-outs: BLISS's desktop dancers, the cursor fight |
| Cover | the last in-sync frame of another take | the rest of a line after a take drifts |
| Extension | a real clip as a video reference | joining generated shots to real footage, frame-continuous |
## Write prompts like a script
CLODYSSEY's first round of prompts was written by planner agents and described poses: "she whispers", "he listens".
Every take came back lifeless, and about 830 credits were gone before anyone read them. The director's verdict: "these
prompts are really really bad and wasting money and need to be 10x better and super context aware of the story", and then
the rule for every prompt since:
> "EVERY prompt needs to have litearlly timestamps, extremely detailed motion descriptions / blocking descriptions, it all
> needs to be deeply integrated with the story."
So the second round was rewritten by hand, all 99 prompts, by one writer in one context:
- **the story beat** of the shot, in one sentence;
- **blocking at 0.0 s**: where she is, where he is, where the camera is;
- **timestamped beats** of action and reaction, on the song's hits (kicks, snares, risers, word onsets);
- **she always does something**: in CLODYSSEY she is seductive and persuasive, close, touching, certain; he feels the weight of it;
- **the camera tied to hits**, with a start, a direction and an end;
- **mouth rules**: singing or closed.
Finish the whole prompt book before spending anything, read it, then run a three-clip pilot before the rest. BLISS did
exactly that and wasted nothing.
## The structure
Seedance follows instructions far better as labelled sections in a fixed order (BytePlus's published Seedance 2.5 prompt
guide lays this out). The order used:
1. **Audio policy** at the top and again at the bottom: what the soundtrack is, and everything that must not be added.
2. **Generation goal:** one sentence of who, where, which moment of the story, which performance.
3. **Reference asset roles**, by upload order, each with one role: `Use @Image1 as the first frame.` as its own
sentence, then what that frame defines.
4. **Subject:** her identity text, and "Only one Claudia appears."
5. **Timeline** in whole seconds, continuous segments (0–1 s, 1–2 s …): what the audio contains and what she does then.
Beats come from the audio; don't timestamp individual drum hits.
6. **Camera:** shot size, angle, movement with subject, start, direction and end.
7. **Core requirements:** lip sync, voice, identity, anatomy, no on-screen text.
8. **Final acceptance criteria.**
Output settings (duration, resolution, aspect) go in the API request, never in the prompt.
## Template: a singing window
`@Image1` is her plate (PNG, downscaled to the output resolution). `@Audio1` is the clean vocal of the window: the vocal
stem with everything outside her words muted, mono, normalised. The window holds one phrase, starts about 0.3–0.7 s before
the first word, and never ends inside a word.
```prompt
【Audio Policy】 The output soundtrack is @Audio1 itself: copy @Audio1 sample for sample from the first frame to the last, with the same timing, pitch and level. Do not re-sing, re-voice or re-generate it. Do not add music, background music, score, instrumental, melody, synth, pad, drone, riser, extra vocals, harmonies, ad-libs, narration or sound effects.
【Generation Goal】 @Image1 sings @Audio1: .
【Reference Asset Roles】
Use @Image1 as the first frame.
This first frame defines Claudia's appearance and wardrobe, her pose, the set, the light and the camera direction.
@Audio1 is Claudia's own singing voice and the complete soundtrack of this video. Her mouth follows @Audio1 exactly.
【Subject】 Claudia: a Caucasian American woman in her late twenties with pale skin and light freckles, a glossy black blunt bob with heavy bangs and one clay-orange streak, a small clay-orange eight-pointed star hair clip, a thin headset microphone, . Only one Claudia appears.
【Performance Timeline】 What @Audio1 contains, and what Claudia does at that moment:
0-1s: @Audio1 is silent, then her voice begins; .
1-2s: in @Audio1 she sings ""; her lips form each of those words exactly when it sounds in @Audio1; .
2-3s: @Audio1 is silent; her mouth relaxes and closes; .
【Camera】 .
【Core Requirements】
1. Lip-sync word by word to @Audio1: mouth shapes match every syllable of @Audio1, its start and its end; no early, late or mismatched mouth movement; the mouth rests naturally closed while @Audio1 is silent.
2. Claudia's voice is @Audio1 and nothing else: she never sings, speaks or hums anything that is not in @Audio1.
3. Her head and shoulders move with the rhythm of @Audio1.
4. Keep Claudia's identity, face, hair with one clay-orange streak, star clip, headset microphone and wardrobe exactly as in @Image1 throughout; the set and light stay consistent.
5. Natural anatomy: natural neck and shoulders, hands with five fingers, no extra or fused fingers, no distorted limbs; normal eyes.
6. No subtitles, captions, lyrics, titles, signs, logos, watermarks or any on-screen text.
【Final Acceptance Criteria】 The soundtrack is @Audio1 unchanged; Claudia's lips match @Audio1 word by word; one Claudia, on-model, with natural anatomy; no on-screen text.
【Audio Policy】 The output soundtrack is @Audio1 unchanged, copied sample for sample; no music, BGM, score, instrumental, melody, synth, pad, drone, riser, extra vocals, harmonies, ad-libs, narration or sound effects.
```
Two details matter:
- **Lyrics go in plain quotes, as what the audio contains.** In Seedance's own syntax, words in braces after a speaker
are lines for the model to voice; written that way it re-sings them in its own voice instead of copying yours. In a
test, the brace form re-sang the line in every take; the copy-mode wording with a clean vocal copied it (normalised
cross-correlation 0.93).
- **One phrase per window.** After a long muted gap the model tends to re-sing the next phrase, and a word cut by the
window's end comes back mangled. Two phrases are two windows.
## Template: a motion clip
The same order, with three changes: `@Audio1` is the song for rhythm and timing only ("the output audio is @Audio1
unchanged"); a performance line says she does not sing or speak (mouth closed, or a natural smile or laugh); and the
timeline runs in stages of about one bar, each naming what her body does with the beat ("steps on every beat, a big
accent on the downbeat at 2 s"). An instrumental mix with the vocal removed works best as `@Audio1` here: nothing tempts
her mouth.
## Lip sync, measured
The fact that made it work was the director's own observation on ESCAPE VELOCITY's first test: "its perfectly synced while
in seedance, but the audio changes when it desyncs." Seedance copies the reference audio into its own soundtrack, and the
mouth follows that soundtrack. So:
1. Compare each take's own soundtrack with the real vocal (spectral match, cross-correlation over time).
2. Keep the take up to the last eighth note before the two diverge. The comparison also gives the take's lag; Seedance
tends to drop about 0.3 s of leading silence.
3. Cover the rest with a new clip: a window that starts in the word gap before the cut, from the frame at the cut.
4. If a phrase fails three times, it becomes a cutaway or an insert.
The full method, with every constant (the clean vocal window, the per-frame match, the null threshold, the cut on the
eighth, covers and coverage), is specified in [What to build](https://claudia.gallery/wiki/toolkit/#15-the-sync-check-the-spectrogram-comparison).

Measured results: about 86% of sung seconds held in sync in ESCAPE VELOCITY, and 62% (7.4 of 11.9 s) over BLISS's six
singing windows; the edit cuts away where they diverge. The bar is "most lips", like a pop video, not every word.
**Short prompts copy, long ones re-sing.** A later test on seven singing windows found that prompt length was the
biggest lever: long timestamped screenplays made the model perform the text and re-sing the line, while a short prompt
(the words quoted, one line of framing and light, at most two timed beats, about 90 words) copied it. The mean in-sync
share per take went from 26% to 58%. Screenplays stay the rule for acting and motion clips; a singing window's one job
is to copy `@Audio1`. And keep her mouth clear on the plate: a microphone at her lips held one window at 26% whatever the
prompt.
**No lip-sync models.** Re-syncing mouths with a dedicated model (Sync, LatentSync, Wav2Lip) was tried as a fallback and
dropped: the mouths come out forced and glitchy. A gap gets a new Seedance cover or a cutaway, and a changed song means
new takes from the original stills, never re-lipped old ones.
## Inputs
- **Colour in, look in post.** ESCAPE VELOCITY's references were pre-graded to monochrome and 32 of 59 clips came back
grey. Generate from the plate's own colour and apply any look in the edit.
- **Downscale stills** to the output resolution (≤1280×720 for 720p) and send PNG. Full-size AI stills cause
fingerprint-like textures in grass and fabric.
- **Distinct first frames.** Two clips should never start from the same frame: use a sibling plate or the previous clip's
last frame.
- **Drafts first.** Higgsfield's API ran 480p drafts at a lower rate that can be finalised to 1080p with the same motion.
## Providers and filters
The films used two routes to Seedance 2.5: OpenRouter (about $0.23 a second at 720p, autumn 2026) and Higgsfield's API.
- **On OpenRouter, a first frame drops the audio.** When a request carries `frame_images`, the audio reference is
silently ignored and the take makes its own soundtrack, so nothing can sync. Singing windows and covers go in
reference mode: the plate as an image reference (`@Image1`) and the window as an audio reference (`@Audio1`).
First frames are for motion clips. On Higgsfield a start image keeps the audio.
- **Audio goes by HTTPS URL.** OpenRouter's video endpoint rejects `data:` URIs for audio; host the window.
Two kinds of refusal came up:
- **Face filters** decline some photoreal close-ups of her. A decline is free, and final for that reference on that
service: never re-grade, blur or otherwise alter the image to get it through. In CLODYSSEY, OpenRouter's filter
declined most close-ups of her, and those shots were made on Higgsfield, which accepted the same references.
- **Silent input filters** fail some jobs with no reason when they misread innocent wording ("a plug pushed deep into
each ear" for foam earplugs). Plain rewording fixed five of six at once.
## Joining real footage
To enter or leave a real clip without a visible seam, use Seedance's video extension (forward or backward) with the real
clip as the video reference: it returns only the new frames, continuous with the real ones and aligned within a pixel or
two. An `end_image` alone is not copied exactly; the last frame gets redrawn.
## Costs, measured
| Film | Clips | Spend |
|---|---|---|
| ESCAPE VELOCITY | 47 clips + 30 continuations, 421 s at 720p | about $93 (OpenRouter) |
| CLODYSSEY | 101 clips, 189 takes | about $107 (OpenRouter) plus 2,436 Higgsfield credits |
| BLISS | 50 clips, 267 s | $182 at list price (Higgsfield API) |
Every prompt the films ran:
[ESCAPE VELOCITY](https://claudia.gallery/record/escape-velocity/prompts-seedance/) ·
[CLODYSSEY](https://claudia.gallery/record/clodyssey/prompts-seedance/) ·
[BLISS](https://claudia.gallery/record/bliss/prompts-seedance/).
---
From the Claudia wiki: https://claudia.gallery/wiki/seedance/
---
# Motion design and the edit
Code on top of video. How the lyrics, graphics, cuts and the XP desktop were drawn frame by frame, on the beat.
The generated footage is only the plate. Everything that carries meaning (every lyric on its word, the graphics that
explain each joke, the split-flap board, the XP desktop) is drawn in code on top, so it is exact and lands on the beat.
## The engine
- **Every frame is a pure function of song time.** A frame never depends on the frame before, so any frame can be
rendered alone, out of order, on any number of workers. Check it: render a few frames in a different order and
compare.
- **Headless Chromium** draws each frame: canvas and WebGL for the graphics and a print pass, the footage underneath.
- **Chapters are separate files** with one contract: draw footage, draw lyrics, draw graphics, draw transitions.
One animator agent writes each chapter, an art-director agent renders real frames and reviews them, and a fixer applies
the notes.
- **Render numbers:** ESCAPE VELOCITY, 7,354 frames at 1080p in about 63 minutes on 4 cores; BLISS, 4,351 frames in 4.7
minutes on a render box.
## A lyric video
All three are lyric videos: every sung word is on screen at its aligned time. A lyric gate checks it before any render
(all 315 words of CLODYSSEY, all 240 of BLISS). But the director's bar is that the lyrics are composition, never
karaoke: they live on surfaces in the scene.
- ESCAPE VELOCITY sets words as subtitles, model cards, garment tags, mastheads and split-flap rows.
- CLODYSSEY has a type system: DINish Condensed Black for "the Mast language" (TIE ME TO THE MAST!, UNTIE ME!), Nimbus
Sans Bold for her lines, IBM Plex Mono for readouts (the rope's load in kN, the P(doom) gauge).
- BLISS puts the chorus in one giant title bar whose × is pressed on "windows", and the verses in balloons, Notepad,
dialogs, an address bar and a chat window.
## Graphics that explain the words
The director's brief on ESCAPE VELOCITY: "motion graphics and motion design, not just animated lyrics but things that fit
the words being said … almost explaining the words with visuals." Every line gets a graphic that explains its joke, built
from real things: real data, real interfaces, real references. Computer-vision boxes tracked to her (a real detector's
boxes, labels and scores), a P(doom) gauge, a care label, a garment tag with a barcode. Nothing decorative.
## Tucks and depth
- **Type behind her.** Per-frame mattes (BiRefNet, SAM 2) let type pass behind her hair, the mast or the Mask. The rule:
tuck an edge, never lose a glyph. The director: "i like rotoscopes and depth mapping to cover text".
- **2.5D.** A depth map (Depth Anything V2) separates a still into layers for parallax. Each 2.5D shot needs a treatment of
its own; a still with a slow zoom is not a shot.
- **Real scenes, not frames.** An insert never floats in a box on a flat background; it sits in the scene, on a screen
or a surface, depth-matched.
## On the beat
- Build a **sound-event map** from the stems: kick, snare and hat onsets, impacts, risers, vocal entries.
- **Cuts land on kicks and snares; camera moves ride risers into their hits.** The director asked for editing at "edgar
wright level … but also not overdone and amateurish": mostly hard cuts on the beat, with a motivated device every few shots.
- **Flashes only where the idea calls for them**, at most three full-frame flashes a second, checked with a WCAG-style
flash test before delivery.
- **Let a calm song be calm.** BLISS's shots run longer ("dont need to cut like michael bay").
## Colour and texture
Keep Midjourney's colour. ESCAPE VELOCITY's first cut ran a monochrome xerox look over everything, and the director's
note was "way too.. gray … loses the Midjourney magic". The print pass became a colour halftone. Texture is decided
per film:
- **ESCAPE VELOCITY:** colour halftone and line screens, grain, dither transitions, misregistration; clay orange
`#D97757` as the graphics accent.
- **CLODYSSEY:** clean, no halftone ("keep it clean so it feels like imax, high res rather than dithered"), every plate
and clip graded toward one storm plate's blue-green-grey ("s018 color is PERFECT", "dont over do it"), with Claude
orange kept as the one accent.
- **BLISS:** a screen recording. Native colour, XP's own palette, nothing on top but the desktop.
## Build the tool first
BLISS needed a Windows XP desktop that could be choreographed, so one was built: about 2,700 lines of JavaScript that
draw Luna windows, the taskbar, the Start menu, 68 icons as vector paths, the 1-bit cursor set and a 1-bit Tahoma, with
the boot, Welcome and Turn Off screens. A shot is a script (windows, icons and a cursor path whose clicks land on the
song's beats) compiled into motion. The storyboard frames and the film are the same pixels.
The wordmark reads "Macrohard", and the welcome mark is a four-pane parody flag drawn in code. The wallpaper is the real
Bliss photograph, graded into five times of day for the story's clock, and the sound layer is made of the real XP sound
files, placed in key (103 cues, ducked under her voice).
## Desktop dancers
The old desktop-dancer programs, as the director described them: "those back then would be on a green screen and was cut
out and youd just have her dancing on teh taskbar". So: a Seedance performance on chroma green → a matte (BiRefNet plus
SAM 2) → cropped, despilled, with a solid core so her feet are never see-through → standing on the real taskbar.

## Render and deliver
- **Encode** a CRF 16 master tagged BT.709. An untagged stream looks washed out in QuickTime.
- **Gates before delivery:** the lyric gate (N of N words), contact sheets of the whole film, chapter seams, 100% crops
of her face, and the flash test.
- **Map a new mix through time,** not through a re-edit: when CLODYSSEY's final mix added two beats, the video followed it
through a time map.
The renderer's contract, the lyric modes, the camera on hits and the checks are specified in
[What to build](https://claudia.gallery/wiki/toolkit/#18-the-renderer).
More, film by film: [the ESCAPE VELOCITY record](https://claudia.gallery/record/escape-velocity/) ·
[CLODYSSEY](https://claudia.gallery/record/clodyssey/) · [BLISS, including how the XP desktop was built](https://claudia.gallery/record/bliss/xp-desktop-and-sound/).
---
From the Claudia wiki: https://claudia.gallery/wiki/motion/
---
# The whole pipeline
How one director and Claude made three music videos in about 80 hours of wall-clock time, stage by stage, with the numbers.
Each film was made by one director, anabology, making every taste call in a few dozen short notes, and Claude (Opus 5.5
in Claude Code) as the production team: research, lyrics, storyboard, every prompt, every script, the edit and every
frame of motion design. Many of the stages ran as multi-agent workflows: several agents working in parallel, judged
and merged.
## The stages
Starting a film of your own? [Start here](https://claudia.gallery/wiki/start/) has the checkpoints, the accounts and the spending rules;
[What to build](https://claudia.gallery/wiki/toolkit/) specifies the tools.
1. **Brief.** The director's ask, word for word, then each decision with its date. Taboos and standing rules go in the
brief and into every agent's prompt.
2. **Research.** Agents build banks of live memes and events, each fact checked by a separate verifier agent. BLISS also
researched the XP sound files (downloaded and pitch-analysed), the old desktop-dancer programs and how to steer Suno.
3. **Lyrics.** Several writers from different angles (density, tension, hook, story), judges, a merge, a critic pass,
then the scansion scorer. See [Song](https://claudia.gallery/wiki/song/).
4. **The song.** Suno takes, round after round, picked by ear; the director often made the final take themselves.
5. **Listening.** Stems, a beat map, every word's time, a list of hits. Everything after sits on that grid.
6. **Her.** The canon text and the looks for this world, tested on a few plates before any batch. See
[Consistency](https://claudia.gallery/wiki/consistency/).
7. **Storyboard.** Three director agents board the song independently (in ESCAPE VELOCITY: fashion film, meme density,
narrative), judges score them, a head director merges the best into one board. Every shot carries its lyric, its type
device, its plate prompt, its camera on the song's hits and what its graphic explains. The director reviews a PDF of
the board; CLODYSSEY got 70 notes shot by shot.
8. **Plates.** Midjourney batches, contact sheets, picks, one image per shot. See [Midjourney](https://claudia.gallery/wiki/midjourney/).
9. **Moving shots.** The prompt book, written by hand in one context and read before spending, then Seedance with lip
sync measured. See [Seedance](https://claudia.gallery/wiki/seedance/).
10. **The edit.** An edit list that resolves singing takes and their covers moment by moment. Stills stand in until clips
exist, so the film renders end to end at every stage.
11. **Chapters.** One animator agent per chapter, an art-director agent reviewing real rendered frames, a fixer; in
BLISS's second round, review → fix → verify → finish, with notes on the seams between chapters. See [Motion](https://claudia.gallery/wiki/motion/).
12. **Checks and render.** The lyric gate, contact sheets, the flash test, then a CRF 16 master.
## What the director did
Taste, at checkpoints, in very few words. ESCAPE VELOCITY took about 50 notes, CLODYSSEY 68, BLISS 21. A few that
changed whole films:
- "way too.. gray … loses the Midjourney magic" (ESCAPE VELOCITY's colour)
- "s018 color is PERFECT" (CLODYSSEY's grade)
- "EVERY prompt needs to have litearlly timestamps, extremely detailed motion descriptions / blocking descriptions"
- "Can you use actual bliss as the bg or too late?" (BLISS's wallpaper)
- "idk i dont love these - id say more say y2k high fashion" (BLISS's wardrobe)
More of them, in context, in each film's making-of: [ESCAPE VELOCITY](https://claudia.gallery/record/escape-velocity/readme/) ·
[CLODYSSEY](https://claudia.gallery/record/clodyssey/readme/) · [BLISS](https://claudia.gallery/record/bliss/readme/).
## Lessons from running it
- **One hand writes the prompts that cost money.** CLODYSSEY's first round of video prompts was written by planner agents
and spent about 830 credits on lifeless takes before anyone read them. Since then, every paid prompt is written in one
context and read before anything is submitted.
- **Spend only after reading.** A background workflow with its own spending stage kept spending after the director had
flagged the prompts. Stop generator workflows the moment quality is in question.
- **Empty is not a pass.** When agents ran out of usage mid-run, chapter reviews came back empty and the scripts read
"empty" as "nothing to fix". A review that returns nothing is a failed review, in code; check that every agent
actually finished.
- **Long runs need their permissions in advance.** One write outside the allowed folders stalled a whole run on a
single approval prompt.
- **Watch the queues.** Every clip import queued a 16-minute matte job; after 80 re-makes the GPU queue was about 20 hours
deep. Cancel superseded jobs so the shots that matter go first.
- **Parallelism has a ceiling.** On a 4-CPU machine a workflow ran two agents at a time; splitting chapters into more
workflows helped less than expected (1.1–1.5× measured, not 3×).
- **Keep run records.** Long workflows get interrupted; resume them by run id and let finished agents replay from cache.
## The numbers
| | ESCAPE VELOCITY | CLODYSSEY | BLISS |
|---|---|---|---|
| Length | 5:06 | 3:11 | 3:01 |
| Shots | 141 | 94 | 44 |
| Midjourney | 188 prompts · 636 images | 161 prompts · 644 images | 138 prompts · 315 images |
| Seedance 2.5 | 47 clips + 30 continuations | 101 clips · 189 takes | 50 clips |
| Wall-clock | about 19 h | about 41 h | about 19 h |
| Director's notes | about 50 | 68 | 21 |
| Claude tokens | 1.4 B | 2.37 B | 2.25 B |
| Claude at API list prices | ≈ $626 | ≈ $1,007 | ≈ $714 |
The Claude figures are what the same work would cost at API list prices; the films were made on a subscription. About
96–98% of all tokens were prompt-cache re-reads of long contexts.
## The full records
Each film's making-of: the story and the numbers, and every prompt for every tool.
- [ESCAPE VELOCITY](https://claudia.gallery/record/escape-velocity/)
- [CLODYSSEY](https://claudia.gallery/record/clodyssey/)
- [BLISS](https://claudia.gallery/record/bliss/)
---
From the Claudia wiki: https://claudia.gallery/wiki/pipeline/
---
# What to build
The tools behind the films, specified so that you or an agent can rebuild them: the job of each, what it takes and returns, how it works with the constants that were measured, and the check that says it is done.
The films' tools lived in a private workspace and are not published. This page is the spec instead. For each tool: its
job, what goes in and what comes out, how it works (with the numbers that were measured or tuned), and the check that
says it is done. None of it is exotic. Most of these tools are a few hundred lines of Python or JavaScript around ffmpeg,
an audio model or a headless browser. Build them in the order of [Start here](https://claudia.gallery/wiki/start/#what-to-build-first). Their
files are in [Data formats](https://claudia.gallery/wiki/formats/).
Two principles run through all of them:
- **A check a script can run beats a judgement.** Every bar that held on the films was something a script measured: a
word on screen at its time, a cut on a beat, a mouth on the audio, money from the bill.
- **Fixed formats with slots beat improvisation.** The Suno pack, the plate prompt order, the clip prompt, the display
tokens and the shot schema are templates. An agent fills them in; it doesn't reinvent them.
## The project
### 1. The project folder and its logs
**Job.** Hold everything a fresh agent needs to pick the film up cold.
**Out.**
- One folder per film (layout in [Data formats](https://claudia.gallery/wiki/formats/#the-project-folder)).
- A **decisions log**: the director's words, verbatim, dated, by whom. A later decision is a new entry that supersedes an
older one, never an edit of it.
- A **ledger** with one line per paid call.
- A **runs file** with every long or delegated run's id and arguments.
- One **status file**, rewritten in place: the stage, what is running, real spend, the checkpoints passed, and what
needs the person.
**How.** Name the project on every command. An `export` inside a backgrounded shell chain once sent commands to the
wrong project. Write state with a locked read-modify-write: parallel runs raced on a picks file and lost picks.
**Done when** a new session can answer "where are we, what is running, what did the director say, what did it cost"
from the files alone.
## The song
### 2. The lyric scorer: scansion and assonance
**Job.** Before any take is spent, predict how Suno will deliver a lyric, and report two numbers per section and for
the song: **scansion** and **assonance**. Also lint the sheet.
**In → out.** The lyric sheet, with a section tag on each section naming its delivery (`[Verse 1 - spoken, deadpan]`,
`[Chorus - sung]`, `[Drop - chant]`) → a report per line and per section. Optionally, each line drawn on the 16th-note
grid.
**How.**
- **The delivery model.** It was measured by laying ESCAPE VELOCITY's real take, forced-aligned, on its beat map, then
calibrated on the director's verdicts on CLODYSSEY's drafts. Suno gives each spoken line its own two-bar slot. The
line enters on the "and" of beat 2 of the first bar, its last word lands on beat 4 of the second bar, and Suno speaks
it at natural rhythm stretched to fit (verse lines 0.8–1.5×, median 1.12).
- **Rates:** spoken about 3.3 syllables a second, sung about 2.9, chanted about 3.0. Aim for 10–13 syllables per spoken
line; 9 can pass. A sung hook runs 4–8 syllables, with common words and open vowels.
- **Syllables and stress.** Split every word into syllables with stress, using a pronouncing dictionary such as CMUdict
plus a fallback for unknown words. Mark the important syllables: lexical stress, names, numbers, words in caps, and the
last content word of each phrase. Function words (the, a, of, to, and, is) should fall between beats.
- **Flow,** for spoken sections, scores:
- how many important syllables land on the grid;
- stress spacing (no pile-ups, no long runs of weak syllables);
- how far Suno must stretch the line;
- the line's shape: at most two phrases between full stops, and no one-word fragment mid-line ("Ugh.", which the
director called "the most awkward part");
- whether couplets rhyme on the beat-4 word.
- **Groove,** for chants and hooks, scores whether short clauses lock to the pulse. Never use it on verses: optimising
verses for groove made one draft choppy.
- **Scansion** = flow without its rhyme term, averaged with groove.
- **Assonance.** For each line end, find its best partner within four lines: a rhyme, a slant rhyme, or shared stressed
vowels. Add the echo of stressed vowels inside lines. Report each line's end sound and its partner.
- **Lint:**
- digits and unspaced acronyms (write "A G I" in the sheet; the screen shows "AGI");
- brackets and symbols, and homographs;
- lines that read as AI-written, and the project's taboo words;
- lines borrowed from other songs;
- a lyrics field over about 3,000 characters (Suno rushes or truncates silently).
**Reference scores** (scansion / assonance) on the finished songs: ESCAPE VELOCITY 75 / 39, CLODYSSEY 83 / 55, BLISS
80 / 43. Spoken, deadpan lines may stay unrhymed; every sung couplet should ring.
**The rewrite loop:**
1. Write 4–8 rewrites of a flagged line that keep its meaning and its reference.
2. Score them and keep the top three.
3. A judge rates each for meaning, wit and the risk of reading as AI-written.
4. Pick the best.
Never touch a line the director has locked, and never optimise past their ear. A draft tuned to a perfect score got:
"perfect beat precision without attention to what words land where doesnt lead to this being great".
**Done when** the scorer ranks every section the director has judged in the director's own order. Keep a calibration
file of those sections with their verdicts, and re-run it after every change to the scorer. If the ranking breaks, the
director is right and the score is only advice.
### 3. The Suno pack
**Job.** Everything a person pastes into Suno, with each field in its own block so it copies from a phone.
**Out.**
- **Style:** genre, BPM, key, instruments, vocal delivery and mix, as one string.
- **Exclude styles.**
- **The lyrics field:** spelled as sung, with section tags that carry delivery and arrangement cues; 1,800–2,400
characters for about 2:20.
- **Title and settings.**
- **A takes log:** each take with the director's verdict on it.
**How.**
- Respell a name Suno mispronounces in the Suno field only; the screen keeps the real spelling.
- Allow one Voice per source song (a combined spoken-and-sung Voice), and also try takes with no Voice.
- If the film has a sound palette, put the song in its key. XP's startup and shutdown sounds are in E-flat major.
- Syncopate hooks and stop lines.
- Change one thing per round.
- Never take audio from Suno's CDN or by recording the player. The person's download is also the licensing step.
**Done when** the person can paste every field without editing, and the takes log names the chosen take.
## Listening
### 4. Stems
**Job.** Split the master into vocals, drums, bass and other.
**How.** htdemucs, four stems. It took about 64 s on 8 CPU threads for a 5-minute song, with a 2.7 GB peak and 84 MB of
weights; the published weights carry no explicit licence. The stems feed different tools:
- **vocals:** alignment, lip-sync windows and the sync check;
- **drums:** the beat map and the kicks;
- **bass and other:** impacts, risers and section energy.
### 5. Timed words and the ear check
**Job.** Give every sung word a start, an end and a confidence, then list the ones a person should hear.
**In → out.** The vocal stem and the lyric sheet → `words.json` (each word: index, text, line, `t0`, `t1`, `conf`, and
its source), plus an ear-check list.
**How.**
1. **Forced alignment** of the known lyric: a wav2vec2 CTC model (`facebook/wav2vec2-large-960h-lv60-self`, Apache-2.0)
gives character probabilities every 20 ms, and a Viterbi pass walks the sheet's characters through them. The
confidence is the mean probability over the word's frames.
- Re-run on ESCAPE VELOCITY, it reproduced the shipped times exactly for 78.6% of words, and within 0.10 s for 92.5%.
Target at least 90% within 0.10 s.
- Singing aligns worse than speech: mean confidence 0.64 on ESCAPE VELOCITY, 0.75 on CLODYSSEY.
2. **A transcription fallback** for words the aligner is unsure of. Whisper hears sung words far better (92.5% text
match against 72%), but it starts words early: mean error −0.33 s. Merge per word, in this order:
1. wav2vec2 confidence ≥ 0.5: keep it.
2. Otherwise, if Whisper heard the word, take Whisper's end time. For words longer than 0.45 s, also move the start
forward to the first 10 ms frame where vocal energy passes 0.3 × the word's peak, unless wav2vec2's start falls
inside the word.
3. Otherwise, if the word sits between two placed words of its line, give it an even share of the gap.
4. Otherwise, leave it unplaced, for the ear check.
3. **Matching** heard words to the sheet:
- spell digits out;
- split a heard word that equals 2–4 sheet words ("AGI" → "A G I");
- map names and brands by hand, in a project table.
4. **The ear-check flags,** most urgent first:
| Flag | When | Usually |
|---|---|---|
| letters | a sheet line has no letters | the sheet is wrong |
| missing | a word has no time | fix it |
| order | a word starts more than 0.05 s before the previous one | an alignment slip |
| gap | more than 1 s between two words of one line | a delay throw, an echo, or the wrong repeat |
| verylow | confidence under 0.1 | the flag that matters most: 11 of the 12 words that were really off on ESCAPE VELOCITY's take carried it |
| length | a word longer than 1.2 s | it swallowed a pause or a neighbour's slot |
| lowconf | confidence under 0.5, or a time from the fallback or interpolation | usually right |
Two flags were tried and dropped, because every word they flagged was correct: a word on near-silence, and a word
shorter than 40 ms.
**Judging without ears.** Weigh the signals in this order:
1. the flag itself;
2. a second aligner: if the two agree within about 0.3 s, keep the word;
3. neighbours and repeats: which occurrence of a repeated line sits inside its line's span;
4. vocal energy, only to choose between candidates that the signals above already proposed.
Send at most five spots per song to the person, as 10–20 s previews with the words on screen: "say ok, early or late for
each".
**What the take really sings.** Transcribe the vocal stem and compare it with the sheet line by line: ok, spelling,
changed, extra, missing or not heard. Never apply a proposed change the person hasn't heard: on CLODYSSEY's take, all nine
proposals were transcription errors.
**Rules.**
- If the take sings something else, fix the sheet. If it sings the sheet, fix the times. Never edit the audio.
- Manual fixes carry a note and survive re-runs.
**Done when** every flagged item is kept, moved or asked about, with a note, and all of it is finished before the
board.
### 6. The beat map
**Job.** Every beat and every bar line of this take. Never one BPM: Suno takes drift like live music. ESCAPE VELOCITY
went from 131.5 to 133.9 BPM and CLODYSSEY from 129.2 to 132.1, so a fixed grid is more than a beat off by the end.
**How.**
1. Take the onset strength (spectral flux) of the drum stem.
2. Fit local tempos on 20 s windows every 5 s: BPM in 0.05 steps, phase in 4 ms steps, inside a tempo range. Drop the
weakest 20% of windows (the breakdowns).
3. Smooth the tempo curve, step beats along it, and snap each beat to a real drum onset within ±30 ms.
4. Find the downbeat phase (which of every four beats is beat 1) with two votes per 32-beat block:
- the kick vote: the phase with the most kick energy;
- the chord vote: the phase where the harmony (chroma of the other and bass stems) changes most.
The chord vote wins when it holds at least half the blocks and changes at least 1.5× more than the other phases;
otherwise the kick decides. On ESCAPE VELOCITY the kick vote was split 36% while the chord vote held 77%. Keep a
manual override.
5. If the tempo search locks onto half or double time, pin a tempo range.
**Done when** section starts, verse entries and impacts mostly fall on beat 1, with kicks heavy on 1 and 3, claps on 2
and 4, and harmony changing on 1. CLODYSSEY's final take needed the override: claps on 2 and 4 fooled the kick vote.
### 7. Sections and hits
**Job.** What the track does, where: the camera answers sounds you can hear, not just the grid.
**Sections.** Segment the harmony and timbre, snap to downbeats, and name the sections from vocals and energy. Confirm
them against the lyric's tags.
- **Pickup rule:** a line starting less than 0.8 s before a section's first downbeat belongs to that section.
- List every vocal gap of 2 s or more; instrumental shots and title cards go there.
- Print each stem's energy per bar on a 0–9 scale, and write one line per section on what the track does ("the beat
stops after 'it sucks'").
**Hits:**
| Event | How it is found |
|---|---|
| kick | onsets in the drum stem, 30–150 Hz |
| snare | onsets at 150 Hz–5 kHz, minus anything within 45 ms of a kick (a broadband kick fires in every band) |
| hat | onsets at 6–16 kHz, minus kick-coincident ones |
| impact | a jump of the next 0.6 s over the previous 1.5 s in the mix, drum or bass envelope; score = max(mix, 0.9 × drums, 0.7 × bass) ≥ 0.22; snapped to a downbeat within 0.65 s; merged within 2 s |
| riser | the other and bass stems rising for at least 1 s, with the hit each one lands on |
| fill | the beat before a real hit, and bars whose last half breaks the pattern |
| vocal entry | where the vocal comes back after silence |
| envelopes | drums, vocals and bass at 24 fps, scaled 0–1 |
**Done when** every impact sits in a loud bar of the energy table. Keep a manual add and drop list: the detector missed
CLODYSSEY's loudest bar, which grew out of an already loud chorus.
### 8. Display tokens
**Job.** What the viewer reads, which is often not what is sung. The display is the joke; the sheet is the phonetics.
**How.** One line per lyric line, made of tokens: `TEXT|n` shows TEXT for the next n sung words, from the first word's
start to the last word's end.
- `+25%|3 =|1 −17%|2` shows "+25% = −17%" over the sung "plus twenty-five is minus seventeen".
- `AGI|3` shows AGI over the three sung letters.
- Markers: `*x*` accent, `~x~` strike-through, `_x_` mono. `|0` rides on the next token.
**Done when** every line consumes exactly its aligned words. Fail loudly if one doesn't.
## The board
### 9. The board file and its audit
**Job.** The film as data, written by commands that check every value, never by hand-editing JSON.
**In → out.** The beat map, sections, words and budget → `shots.json`: a draft, then the written board.
**How.**
- **The draft** cuts each section at a pace: seconds per shot, per section. ESCAPE VELOCITY's measured pace was about
1.35 s in the intro, 1.6 s in verses, 2.7 s in sung choruses and 4.4 s in the outro. A calmer song might run 4–7 s a
shot.
- **Every creative field starts as a TODO.**
- **Writer commands** set fields one shot at a time, or from a table. They refuse:
- placeholder text;
- unknown device kinds, bad anchors and bad timing keys;
- a prompt in the wrong generator's field.
**The audit's gates** are listed in [The storyboard](https://claudia.gallery/wiki/storyboard/#the-audit). It exits non-zero while any gate
fails, and it prints the exact entry for every clip still to make.
**Done when** the audit exits 0. The approval is recorded against a hash of the board, so any later change makes it
stale.
### 10. The board sheet
**Job.** The board as something a person can review on a phone.
**Out.** A PDF and JPG pages of 12 cards each. Each card holds the frame (or the prompt's first words), the id, the
time, the lyric, the framing, the action, what moves and the explaining graphic. A cover page shows the chapter map. A
FACE? badge marks every plate listed in a flags file. If the film has a grade, the frames are graded.
## Pictures
### 11. The plate book, contact sheets and picks
**Job.** One Midjourney prompt per plate, the images back in, and a pick for every shot.
**How.**
- **The prompt order:** framing first, then her identity text (the full form for close and medium shots, the short form
for full length and wider), wardrobe, the action and her position in the frame, any empty field kept for type, the
film's style family, and "no text, no letters, no logos". Then the parameters. See
[Midjourney](https://claudia.gallery/wiki/midjourney/) and the [canon](https://claudia.gallery/claudia/).
- **Runs:** a person runs the batches in Midjourney, at about a minute a job, four images each.
- **Ingest:**
- identify each job by its own prompt text, never by its position on the page;
- Midjourney rewrites long prompts on submit, so match those by hand;
- keep all four candidates.
- **Contact sheets** are numbered and show faces large.
- **Picks:** plate → file, candidate number, and the shots that use it.
- **Variety:** the three candidates of each job that weren't picked are a free variety pool. Use one image per shot, and
never start two clips from the same frame.
**Done when** every plate on the board has a picked file, and no file sits on two shots that are not one continuous
moment.
### 12. Plate prep: depth and boxes
**Job.** What the 2.5D moves, the type and the camera need from each still.
**How.**
- **Depth:** Depth Anything V2 Small, which is Apache-2.0; the Large model is non-commercial.
- **Boxes:** face and person boxes from Grounding DINO. Graphics anchor to them, and the camera keeps the face in frame
through pushes and shakes.
- **Mattes,** where type tucks behind her.
- Re-run all of it after any plate edit: edited plates once rendered with their old depth.
## Moving shots
### 13. Clip files and the spend gate
**Job.** One file per clip, from draft to delivered, and nothing spent without a locked prompt and a yes.
**How.**
- **Status:** draft → locked → submitted → delivered, declined or failed.
- **Each clip file holds** the kind (sung or motion), the provider, the start image, the window's song time, its exact
words and the request. Each take records its file, its cost and its sync result.
- **Locking checks:**
- the words are quoted exactly;
- the window rules hold;
- a sung prompt stays under about 90 words and two timed beats (warn).
- **Submitting:**
- submit only when the clip is locked and the approval covers this exact prompt;
- validate the job id: "submitted" with no id, and HTTP 500, 502 or 520, were silent failures;
- never resubmit a request that got a job id;
- write a manifest after each batch, so an interrupt never loses which jobs exist.
- **Polling:** 70–130 s a clip on OpenRouter, and 4–11 minutes on Higgsfield.
- **Real spend** comes from the provider: OpenRouter's credits and key endpoints, and Higgsfield's transactions.
**Provider facts,** checked with paid probes:
- **HTTPS audio only.** OpenRouter's video endpoint rejects `data:` URIs for audio. Host each window over HTTPS: a
Cloudflare quick tunnel worked, and so did a provider's media URL.
- **A first frame drops the audio.** On OpenRouter, `frame_images` overrides the audio reference: the take makes its own
soundtrack and nothing can sync. Send sung clips and covers in reference mode: the plate as an image reference
(`@Image1`) plus the window as an audio reference (`@Audio1`). First-frame mode is for motion clips.
- **Face filters.** They often decline photoreal faces, and the decline is free. Send the same unaltered image to the
other provider.
- **Audio refusals.** "Output audio may contain sensitive information" and copyright refusals on vocal references are
unbilled. Retry once with a re-cut window or a reworded prompt, then try the other provider.
- **Higgsfield.**
- Ask for the price before each new configuration.
- Upload by URL: presigned uploads flood an agent's context.
- A 480p draft can be finalised to 1080p within seven days.
**Done when** nothing can spend without a locked clip and a recorded yes, and the ledger reconciles with the bill.
### 14. The clean vocal window
**Job.** Give the video model one voice to copy: only the words she sings on camera in the window, and nothing else.
**How**, with the constants used:
1. Cut the vocal stem at `[song_t0, song_t0 + duration]`, mono, 44.1 kHz.
2. Keep the aligned words that start inside `[song_t0 − 0.02, song_t0 + duration − 0.08)`, and only hers. Mute the tail
of a word that began before the window.
3. Merge words into phrases across gaps of 0.35 s or less.
4. Keep each phrase from 0.05 s before its first word to its last word's end. Add a release tail that follows the stem's
10 ms energy until it falls 24 dB below the phrase peak: at most 0.40 s, and never past the next phrase's start minus
0.05 s.
5. Put 15 ms raised-cosine ramps on every gate edge, and silence everywhere else.
6. Normalise the kept part to −16 dBFS RMS, then limit peaks to −1 dBFS.
7. Write a WAV (OpenRouter, by HTTPS URL) and an MP3 at 192 kbps (Higgsfield).
A window holds one phrase, starts in a word gap, and never ends inside a word.
### 15. The sync check: the spectrogram comparison
**Job.** Measure how much of a sung take is really in sync, and where it stops.
The director noticed it on ESCAPE VELOCITY's first tests: "its perfectly synced while in seedance, but the audio changes
when it desyncs. if you check the spectrogram for audio against reference audio you can see the clear divergence point.
you could just create a cut and make another clip to cover desyncs."
Seedance copies the reference audio into its own soundtrack, almost sample for sample while it copies, and the mouth
follows that soundtrack. So compare the take's soundtrack with the window you sent, find where they part, and cut there.
**In → out.** The take and the exact window file → the lag, a copy flag, the in-sync spans in song time, the divergence,
the cut, the words at the divergence, a diagnostic image and a review video.
**How**, with the constants used:
1. Decode both to 16 kHz mono. Use 10 ms frames (hop 160), 64 ms FFT windows (1024) and 80 mel bands from 50 to
7,600 Hz.
2. **Global lag:** FFT cross-correlation within ±1.0 s. The peak's normalised value is the take's NCC. The take is a
copy if NCC > 0.2; otherwise the model re-sang it.
3. Place the reference on the take's timeline at that lag. Per 10 ms frame:
- **wave:** normalised cross-correlation of the waveforms over 100 ms. A verbatim copy scores 0.6–1.0.
- **mel:** cosine similarity of mean-removed log-mel frames (70 dB floor), smoothed over 11 frames. It holds on
fricatives, whose noise never correlates sample by sample.
- **active:** either track's RMS is above 10% of its 95th percentile.
4. **A threshold per take, against a null.** Compute the same mel similarity with the reference shifted by ±37, ±61,
±89 and ±131 frames; these offsets avoid eighth-note multiples. Then:
- `null90` = the 90th percentile of the null;
- `matched_med` = the median mel where wave > 0.6, or the 75th percentile of active mel if there are 10 or fewer
such frames;
- `thr = null90 + 0.4 × (matched_med − null90)`.
5. **A frame matches** when it is active and wave ≥ 0.4 or mel ≥ thr.
6. **Divergence** is the first 300 ms stretch (at least 60% active, under 30% matched, starting on an unmatched frame)
that never recovers with a 400 ms matched run. Short dips that recover are dips, not cuts. With no such stretch,
voice at least 80 ms after the last match counts as divergence. `synced_until` = the last matched frame minus 42 ms
(one frame at 24 fps).
7. **In-sync spans** are matched voice, bridged across silences and across up to 0.25 s of unmatched voice. Drop spans
shorter than 0.5 s.
8. Map everything to song time: `song time = song_t0 + clip time − lag`.
The **diagnostic image** shows our vocal's spectrogram with word labels, the take's own soundtrack, the match curves
with the threshold, and the cross-similarity matrix, whose diagonal means "in sync". The **review video** lays our
audio over the take, marked IN SYNC up to the cut and DIVERGED — CUT HERE after it. Look at both before trusting a
number.
**The NCC is a diagnostic, not the verdict.** Over a whole take it is usually 0.1–0.3, because the copy stops somewhere.
One take scored 0.157 and still held 99% of its sung seconds. Measure the lag per take, never as a constant: takes ran
from −0.40 s to +0.83 s. Motion clips conditioned on the mix aren't copied (NCC about 0.05–0.08): play them at lag 0.
**Done when** each take has its spans and its cut, and the review video agrees with them by eye.

### 16. Covers and coverage
**Job.** Keep the in-sync part of every take, cover the rest, and know how much of the singing is really synced.
**The cut.** Cut on the last eighth note at or before `synced_until`. Eighths are the beat map's beats plus the midpoints
between them. A cut on an eighth feels like part of the music; a cut at the divergence frame looks like a glitch.
**A cover** is a new clip:
- **Window:** it starts in the word gap before the cut, and lasts `clamp(ceil(end − start + 0.3), 4, 8)` whole seconds.
- **Start image:** the parent take's frame at the cut (clip time = cut − `song_t0` + lag), unaltered, sent as the image
reference, never as a first frame.
- **Words and audio:** exactly the words that start in its window, with their clean vocal.
- **Prompt:** the parent's prompt, re-quoted to the new words.
- **Records:** the parent lists its covers; the cover records its parent and why.
- **Limits:** at most two covers per root, and one on a provider where copies are rare.
**Coverage**, per chain (a root and its covers):
- sung stretches = each aligned word ± 30 ms, merged when closer than 0.12 s;
- in-sync spans = every element's spans, each start extended by 60 ms;
- gaps = sung time not covered by any span, of 0.6 s or more, merged when closer than 0.3 s.
The bar is "most lips", like a pop video ("like grimes vids … shedoesnt sync lips on everything, just most"). Target at
least 80% of sung seconds, never an out-of-sync mouth held on screen, and budget covers at about 60% of the sung
seconds.
**Measured:**
| Run | Result |
|---|---|
| ESCAPE VELOCITY | 41% of sung takes copied; with covers, about 86% of sung seconds verified |
| CLODYSSEY, Higgsfield | about 21% of sung takes copied |
| BLISS | 62% of six windows' sung seconds |
| A later test run, seven windows | long prompts held 44% of the chorus's sung seconds; short ones 62–71% |
**When a window won't copy**, try these in order:
1. **The short sung form.** The biggest lever measured is prompt length. Long timestamped screenplays make the model
perform the text and re-sing it; short prompts copy. The short form is:
- the head, quoting every word;
- one line of framing and light;
- at most two timed beats;
- "One continuous shot, no cuts, no head turns. Only she sings. No text, letters or logos.";
- about 90 words or fewer.
On seven windows the mean in-sync share per take went from 26% to 58%; one window went from 0% to 78%. Screenplays
stay the rule for acting and motion clips.
2. **Clear the mouth.** Nothing in front of it on the plate: no held microphone, hand or hair.
3. **The other provider,** after two non-copies in the short form.
4. **Cut away:** to the shot's motion clip, stills in rhythm, depth moves, crops to the eyes, hair or hands, inserts, or
a wider or profile framing. The words stay on screen either way.
**No lip-sync models.** Sync, LatentSync and Wav2Lip were tried as fallbacks and dropped: "still not exact", and the
mouths came out forced and glitchy. The failure was the input, not the model.
## The edit and the renderer
### 17. The edit list
**Job.** Resolve every moment of every shot to a source.
**How.** Resolve a sung shot moment by moment in song time. At each time T, use the chain element whose in-sync span
contains T, at clip time `T − song_t0 + lag`. A take and its covers then interleave, each switch sitting on the eighth
before a divergence. Fill holes from the shot's motion clip, then from the plate. Motion clips play straight through,
holding their last frame if the shot is longer.
**Done when** the film renders end to end at every stage, with plates standing in for clips that don't exist yet.
### 18. The renderer
**Job.** Draw every frame: footage, 2.5D, lyrics, explaining graphics, chrome and transitions, on the song's clock.
**The frame contract.** It renders offline in headless Chromium, one frame per 1/24 s, in parallel and out of order.
- **Every frame is a pure function of song time t.** No state between frames, no `Math.random` (use a hash of the
index), no `Date`.
- **Two canvases.** The base holds footage and anything graded with it. The overlay holds type and chrome, composited
crisp on top.
- **Memos only for data derived from static media,** keyed by media name, never by frame.
- **Every 2D canvas is software-backed** (`willReadFrequently`), and every scene starts from a reset context. A
GPU-backed canvas under software GL kept the previous frame's overlay once frames got dense; that took about two hours
of bisecting to find.
- **Quantise to frames.** A cut shows on the first frame at or after its time. Video may run up to one frame ahead of
the sound, never behind it.
- **Speed:** at most 1,500 ms a frame. The 64-step parallax shader costs 0.5–1.1 s, and a full-resolution canvas blur
about 0.5 s. Use depth slices (about 30 ms) and blur by resampling instead.
**The clock.** All musical time goes through the beat map: beat position, bar, local beat length, kick pulses and
eighths.
- A lyric token appears at its aligned start exactly, never snapped to a beat: "the motion syncs to the beat, the word to
the voice".
- An entrance takes at most an eighth of a beat.
- Every timed thing carries a timing key:
- **W** a word's time;
- **B** the beat grid;
- **H** a track event (kick, snare, impact, riser, fill, envelope);
- **F** a fixed song time;
- **T** a keyframe measured on a take's own frames.
**Lyric modes.** Every sung word appears in at least one, and every drawn token is marked for the lyric gate.
| Mode | For |
|---|---|
| subtitle | sung lines; the sung word lit in the accent |
| coverline | spoken lines set beside her in the plate's empty field, with tracking collapsing word by word |
| masthead | payoff words, numerals, the answer in a call and response |
| stack | short emphatic lines, one word per row |
| mass | a declared wall of micro text in a drop, with the sung line still drawn |
| title | the title knocked out of black |
Type rules:
- One lyric block per frame; a line keeps one mode and one position across cuts.
- At least 40 px, and never over her face.
- Margins are left 72, top 56, right 96 and bottom 84 at 1920 × 1080.
- Hang type from the left axis or from a feature in the frame.
- Nothing may cover the lyric: the renderer measures it every frame and moves the other layers off it.
**The camera plays the track:**
- **punch** on kicks: amp 0.01–0.03, gated by the drum envelope;
- **shake** with the drums: none in screen or UI shots;
- **ride** a riser into the hit it lands on;
- **impact** on real hits: strength at least 0.6, zoom kick 0.06;
- **whip** on fills.
Hold still where the song breathes. A move never cuts the face.
**Density** is chosen per film:
- **dense:** three or more layers, five or six as the norm;
- **budgeted:** footage, one lyric block, one hero device, at most two detection marks, the chrome, and the accent on at
most two elements;
- **sparse:** about 70% of the film with no graphic at all.
**Chapters.** In the Full scale each chapter is its own code file, written by an animator agent against this contract.
Chapters draw only what is unique; shared helpers live in one library. Workers report shared-file bugs; the operator
fixes them once.
**Rendering.** The render is resumable: skip frames that exist, write each frame to a temp file and rename it, open a
fresh page every 200 frames, and retry a failed frame once. ESCAPE VELOCITY's 7,354 frames took about 63 minutes with
three workers on 4 CPUs; at 1080p one cut's frames take about 9.6 GB.
## Checks
### 19. The lyric gate
For every sung token, render the frame at `t0 + clamp((t1 − t0)/2, 0.03, 0.12)` and check that the token was marked as
drawn on it. It prints N/N; anything less fails. ESCAPE VELOCITY: 417/417. CLODYSSEY: 315/315.
- Run it per chapter while animating, and for the whole film before any render.
- It proves presence and timing, not legibility; contact sheets check legibility.
- Fix a mistimed word in the alignment, never in chapter code.
### 20. The history check
Render frame B fresh, then render it again right after frame A in the same page. It fails if any pixel differs by more
than 40 levels.
- Run three pairs per chapter, one across each seam, with one pair whose A is a heavy 2.5D frame.
- It must print HISTORY-INDEPENDENT.
### 21. The flash check
This follows WCAG 2.3.1's general flash threshold.
- Read the video at 96 × 54, 24 fps, in linear relative luminance.
- A flash is a pair of opposing changes of at least 10% of the range, where the darker state is under 0.8, over at least
a quarter of the frame.
- Fail any one-second window with more than three flashes. Count pairs, not transitions: an early version counted each
transition, so 2.2 flashes a second read as 4.4.
- Flashes belong to a theme, on hits: "dont need to force flashing".
### 22. Pace against the music
Picture motion is the mean absolute frame difference of the review copy at 160 × 90 grey, per section. Compare it with
each section's stem energy and its kick and hat onsets per bar. Busy picture over quiet music reads as "off pace". The
fix on one film halved the cuts (0.92 → 0.47 a bar) and slowed in-shot motion until the picture followed the music.
### 23. Contact sheets and the reviewer
**Contact sheets** are how agents look.
- **Per chapter:** every shot's first, middle and last frame; every lyric line's first and last word; every hit; each
graphic's entrance; each transition's middle.
- **The reviewer** sees at least 24 frames, plus full-size stills of the five most important moments.
- **The whole film:** mid-shot frames of every shot, 6 × 24 to a sheet, plus the seams and 100% crops of faces.
- Open every sheet. A sheet nobody looked at is not a review.
**The reviewer** renders and looks for itself, and never edits files. It scores four lenses, 0–10:
- **taste:** the director's taste, which counts double;
- **craft;**
- **sync;**
- **story.**
It returns at most 12 concrete fixes, each marked blocking or polish, and copies the lyric gate's printed line.
**The verdict is computed in code, not taken from the reviewer:**
- A chapter passes when no blocking fix is left, taste is at least 7, the other three lenses are at least 6, and the gate
reads N/N.
- A review that returned nothing, or never looked at a frame, is a failed review, never a pass. The original scripts read
an empty review as "no fixes", and runs that hit a usage limit reported success.
Run at most two review-and-fix rounds per chapter. Calibrate the reviewer on the director's verdicts: on CLODYSSEY the
reviewers scored 6–7.5 a cut the director called "perfect. change nothing."
### 24. The safety and canon scan
**Job.** Look at every generated take before release.
**How.**
- Make sheets of frames per take.
- List every take checked, and every problem with its severity: block, fix or note.
- Whitelist intended props by name.
- Re-scan every re-make.
On CLODYSSEY the scan found 12 blocks in 96 "final" takes, among them nudity through a corset, modern objects and
earmuffs.
## Delivery
### 25. Encode and deliver
**The master.** Frames are sRGB, so convert through RGB with accurate rounding to BT.709 and tag the stream:
```sh
ffmpeg -framerate 24 -i frames/f%05d.jpg -i song.wav -map 0:v -map 1:a -c:a aac -b:a 256k -shortest \
-vf "format=rgb24,scale=flags=accurate_rnd+full_chroma_int+full_chroma_inp:out_color_matrix=bt709:out_range=tv,format=yuv420p,setparams=color_primaries=bt709:color_trc=iec61966-2-1:colorspace=bt709:range=tv" \
-colorspace bt709 -color_primaries bt709 -color_trc iec61966-2-1 -color_range tv \
-c:v libx264 -preset slow -crf 16 -pix_fmt yuv420p -movflags +faststart master.mp4
```
An untagged or BT.601 stream looks lighter and washed out in QuickTime ("stuff is v blown out").
- **The review copy:** the same, with `-crf 20 -maxrate 8M -bufsize 16M`.
- **Verify with ffprobe:**
- codec, 24/1, and a frame count equal to duration × 24;
- bt709 / iec61966-2-1 / tv;
- AAC stereo;
- no container tags beyond the encoder's own.
- **Then run the flash check on the master.**
**The 9:16 cut** follows a subject box per plate, so an off-centre lead stays in frame. It gets its own lyric gate.
**Housekeeping:**
- Run encodes as tracked background jobs, never a detached process nobody watches.
- Delete the frames only after ffprobe has confirmed the outputs.
- Keep big scratch on the big disk.
## Running agents
### 26. Workflows, run records and the stop switch
The films ran their heavy stages as multi-agent workflows (Claude Code):
- **storyboard:** three director agents, three judges, and a merge;
- **chapters:** an animator, an art director who renders and reviews contact sheets, and a fixer;
- **research and lyrics:** topic researchers, fact verifiers, writers and judges;
- **the safety scan.**
**Rules that cost time to learn:**
- **Save every run's id and arguments** the moment it starts, and resume by run id in the same session. Finished agents
replay from cache. An interrupt kills every running workflow, so keep phases short; ESCAPE VELOCITY's merge restarted
three times.
- **Check that every agent actually finished.** A run can report success while agents that hit a usage limit returned
nothing. Carry open items forward explicitly.
- **Parallelism has a ceiling.** On 4 CPUs, run two agents a workflow. Allow one rendering workflow at a time: six
rendering agents filled the swap and the root disk.
- **Put long lists in a file,** not in the launch arguments: a placeholder once launched in place of 96 paths.
- **The stop switch.** On any quality complaint from the director:
1. stop every job and workflow that can spend, after searching each running script for its generate calls;
2. confirm at each provider that nothing is queued;
3. log the director's words as a rule;
4. fix the prompts, then re-make. Don't review the bad takes.
- **Write the acceptance test before the brief,** and make the check something the worker can't edit. Every human
observation becomes a metric the same day: "if a reviewer can see it and the judge cannot, the judge is wrong".
---
From the Claudia wiki: https://claudia.gallery/wiki/toolkit/
---
# Data formats
The files every tool reads and writes: the project folder, the timed words, the beat map, the hits, the display tokens, the board, the clips and their sync, the edit list, and the logs.
These are the contracts between the tools in [What to build](https://claudia.gallery/wiki/toolkit/). Times are seconds of song time unless
noted. Keep the names, and any tool written against them, by you or by an agent, will fit the others.
## The project folder
```text
/
project.json the settings (below)
BRIEF.md the director's ask, verbatim
DECISIONS.md dated decisions, verbatim; later entries supersede earlier ones
RUNS.md ids and arguments of long or delegated runs, for resuming
STATUS.md one status, rewritten in place
ledger.jsonl one line per paid call
audio/song.wav the master as delivered (44.1 kHz stereo)
audio/stems/ vocals.wav, drums.wav, bass.wav, other.wav
audio/windows/.wav|.mp3 clean vocal windows for sung clips
lyrics/lyrics.txt the sung words, spelled as sung, one line per lyric line
lyrics/display.txt the on-screen tokens, one line per lyric line
analysis/words.json timed words
analysis/beats.json the beat map and sections
analysis/hits.json the sound events
analysis/report.md a readable summary and the ear-check list
board/shots.json the storyboard: the source of truth for the edit
board/audit.json the audit's results
board/board.pdf the board sheet for review
plates/prompts.md the plate prompt book
plates/cands//.jpg every candidate
plates/picks.json {plate: {file, cand, shots}}
clips/.json one file per clip (below)
clips//take-N.mp4 the takes, with take-N-sync.png beside them
edit/edl.json the edit list
engine/ the renderer's per-film data and chapter code
render/frames/ frame i = song time i/24 s
out/ review.mp4, master.mp4, vertical.mp4
```
The lyric files share one parser. Blank lines and every line starting with `[` (section tags, notes) are skipped, and
the Nth remaining line of `display.txt` shows the Nth remaining line of `lyrics.txt`.
## project.json
```json
{
"name": "my-song", "title": "My Song", "scale": "quick | studio | full",
"song": "audio/song.wav", "fps": 24, "size": [1920, 1080],
"lead": {
"canon": "",
"far": ""
},
"look": {"mode": "native | clean-grade | print", "accent": "#D97757",
"fonts": {"lyric": "…", "micro": "…", "display": "…"}, "hero_plate": null},
"mj": {"params": "--ar 16:9 --v 8.2 --raw --hd", "profile": ""},
"models": {"video": "bytedance/seedance-2.5", "edit": "google/gemini-3-pro-image"},
"budget": {"openrouter_usd": 100, "higgsfield_credits": 1000, "confirm_over_usd": 5,
"motion_s": 200, "frozen": false},
"banned": ["words", "never", "to", "use"],
"audio": {"tempo_range": [129, 136], "downbeat_phase": null,
"hits": {"add": [[170.7, 0.8]], "drop": []},
"name_map": {"": ["", ""]}},
"board": {"lead_min": 0.7, "max_clips": 45},
"gates": {"board_approved": false, "prompts_approved": false, "v1_approved": false}
}
```
`budget.frozen` is the stop switch: while it is true, nothing may spend. `banned` is checked by the board audit and
belongs in every agent's prompt.
## Timed words
`analysis/words.json`:
```json
{"words": [{"i": 0, "w": "ladies", "line": 0, "t0": 9.32, "t1": 9.64, "conf": 0.91, "src": "w2v"}],
"lines": [{"line": 0, "text": "Ladies. Gentlemen. Agents.", "t0": 9.32, "t1": 10.72}],
"gaps": [{"t0": 18.8, "t1": 26.3, "after_word": 12}]}
```
`src` is `w2v`, `whisper`, `interp` or `manual`. A manual word carries a `note` with the evidence, and re-runs keep it.
## The beat map
`analysis/beats.json`. The tempo varies; never derive beats from one BPM.
```json
{"duration": 306.4,
"beats": [0.452, 0.908],
"downbeats": [0.452, 2.276],
"tempo": [{"t": 0.0, "bpm": 131.55}, {"t": 5.0, "bpm": 131.6}],
"downbeat_phase": 0, "downbeat_method": "harmony | kick | manual",
"sections": [{"name": "intro", "t0": 0.0, "t1": 18.96, "energy": 0.31, "bar0": 0, "bar1": 9,
"source": "lyrics-tag | set | heuristic"}]}
```
Bars count from 0: bar k starts at `downbeats[k]`. Eighth notes are the beats plus the midpoints between them.
## Hits
`analysis/hits.json`:
```json
{"kick": [0.45], "snare": [0.91], "hat": [0.68],
"hits": [{"t": 10.72, "score": 0.81}],
"risers": [{"t0": 9.83, "t1": 10.72, "land": 10.72}],
"fills": [{"t0": 42.3, "t1": 44.1}],
"vocal_entries": [9.32],
"env24": {"drums": [0.0], "vocals": [0.0], "bass": [0.0]}}
```
## Display tokens
`lyrics/display.txt` has one line per lyric line. `TEXT|n` shows TEXT for the next n sung words, and `n` defaults to 1.
```text
Ladies. Gentlemen. Agents.
+25%|3 =|1 −17%|2
Feel the AGI|5
```
The markers are `*x*` accent, `~x~` strike-through and `_x_` mono; `|0` rides on the next token. A line must consume
exactly its aligned words.
## The board
`board/shots.json`, top level:
```json
{"version": 1, "duration": 306.4, "chapters": [], "shots": [], "plates": []}
```
A chapter is `{id, name, t0, t1, sections, accent, set}`. A plate is `{plate, set, prompt, aspect, used_by: [shot ids]}`.
| Field | Type | Rule |
|---|---|---|
| `id` | `"s001"` … | stable once the director has seen them |
| `t0`, `t1` | seconds | on beats (within 0.03 s) or declared hits; shots tile the song |
| `section` | name | a section of the beat map |
| `lines` | [int] | the lyric lines whose words fall in the shot |
| `set` | text | the place, concretely |
| `lead_on_screen` | bool | true only when she is recognisable |
| `framing` | ECU · CU · MCU · MS · MWS · WS · FS · EWS · OTS | |
| `action` | text | what happens and the reaction, with an end state |
| `camera` | text | one move, tied to a sound |
| `motion_tier` | still · plate3d · motion · sung | |
| `plate` | key or null | null only for pure motion design |
| `clip` | id or null | for motion and sung shots |
| `prompt` | text | the plate's prompt, framing first, ending "no text, no letters, no logos" |
| `type` | [{`line`, `mode`, `text`, `placement`}] | mode: subtitle · coverline · masthead · stack · mass |
| `devices` | [{`kind`, `anchor`, `params`, `key`, `line`}] | key: W · B · H · F · T; the explaining graphic is one of these |
| `treatment` | text or null | a 2.5D family; never the same on two 2.5D shots in a row |
| `transition_in` | text | `cut`, or the motivated device |
| `chrome` | text | what the always-on devices show here |
| `accent` | hex | the chapter's accent |
| `hit` | seconds | optional: a declared hit this shot starts on |
| `notes` | text | the director's notes, verbatim |
ESCAPE VELOCITY's first shot, in this schema. The prompt is written with today's canon text; the film's own predates
the adult wording.
```json
{"id": "s001", "t0": 0.0, "t1": 2.111, "section": "intro", "lines": [],
"set": "the back seat of a driverless robotaxi on US-101 at night", "lead_on_screen": true, "framing": "CU",
"action": "Frame 1 is her face in the dark back seat, eyes closed. A red lamp slides across her face left to right and lights the clay streak and the star clip. On the song's first hit (1.12) her eyes open straight into the lens and stay there.",
"camera": "locked, very slow push in (3%)", "motion_tier": "motion", "plate": "waymo-cu", "clip": "SD01",
"prompt": "close-up of a 28-year-old Caucasian American woman with a grown-up angular face, defined cheekbones and a strong jawline, pale skin and light freckles, a glossy black blunt jaw-length bob with heavy straight bangs and one clay-orange streak through the bangs, a small flat clay-orange eight-pointed star hair clip, a thin headset microphone at her cheek, sitting upright alone in the back seat of a driverless robotaxi at night, her face in the centre of the frame, eyes closed, a single red lamp sliding across one cheek through the side window, dark leather seat, the lower third of the frame dark and empty, blue-black night, fog, one small red lamp as the only warm note, fine silver-gelatin grain, cinematic 35mm film still, no text, no letters, no logos",
"type": [],
"devices": [
{"kind": "flap", "anchor": {"slot": "bottom"}, "key": "F",
"params": {"rows": [{"text": "18 MONTHS TO ESCAPE THE PERMANENT UNDERCLASS", "prev": "", "at": 0.0}], "cw": 36, "gap": 3}},
{"kind": "sheet", "anchor": {"box": "a face"}, "key": "B",
"params": {"title": "EXPRESSION · MODEL 01 · LIVE", "rows": [["CALM", ".98"], ["WORRY", ".02"], ["THRESHOLD · WORRY > 0.50", "NOT MET", "hot"]]}},
{"kind": "cv.bracket", "anchor": {"box": "a face"}, "key": "B", "params": {"label": "FACE · MODEL 01", "conf": 0.99}}],
"treatment": null, "transition_in": "cut",
"chrome": "the split-flap board across the lower third; clatters from frame 1, settles by 1.5 s",
"accent": "#D97757",
"notes": "The hook. At 1.0 s: her face, the streak and a readable board line on a muted phone."}
```
The idea behind the explaining graphic: a face that is not worried under the year's most-quoted threat, and an
expression meter that says so in numbers.
## Clips
`clips/.json`:
```json
{"id": "W02", "shot": "s031", "kind": "sung",
"status": "draft | locked | submitted | delivered | declined | failed",
"provider": "openrouter | higgsfield", "model": "bytedance/seedance-2.5",
"start": "plates/cands/siren-cu/2.jpg",
"song_t0": 8.56, "needed": [8.85, 12.4], "words": "Why would I deceive you?",
"window": {"t0": 8.56, "t1": 12.56, "mode": "clean", "file": "audio/windows/W02.wav", "url": "https://…"},
"body": {"prompt": "…", "duration": 4, "resolution": "720p", "aspect_ratio": "16:9"},
"takes": [{"n": 1, "provider": "higgsfield", "job": "", "status": "delivered",
"file": "clips/W02/take-1.mp4", "credits": 28,
"sync": {"lag_s": 0.7063, "ncc": 0.22, "mode": "copy", "synced": [[8.85, 10.25]],
"diverge_at": 10.31, "cut_at": 10.2458}}],
"covers": ["W02-c1"],
"parent": null,
"why": "cover for : it diverges from our audio at song ; this window starts in the word gap at "}
```
The time maps: `song time = song_t0 + clip time − lag_s`, and `clip time = song time − song_t0 + lag_s`.
## The edit list
`edit/edl.json` maps each shot to its sources:
```json
{"s001": {"segs": [{"t0": 0.0, "clip": "SD01-t1", "in": 0.0}]},
"s031": {"segs": [{"t0": 8.85, "clip": "W02-t1", "in": 0.996}, {"t0": 10.2458, "clip": "W02-c1-t1", "in": 0.31}]},
"s002": {"plate": "p002"}}
```
`in` is the clip time at the segment's start: song time − `song_t0` + `lag_s`.
## The logs
`DECISIONS.md`:
```text
## · director · board
"s034-s036 could use more creativity, kind of a boring moment."
→ s034–s036 rethought (replay loop, tally marks).
## · director · supersedes "keep the halftone"
"keep it clean so it feels like imax, high res rather than dithered"
```
`ledger.jsonl`, one paid call per line. A real cost is filled in from the provider; the estimate is never quoted as
spend.
```json
{"ts": "…", "provider": "openrouter", "model": "bytedance/seedance-2.5", "kind": "video", "units": 5,
"est_usd": 1.16, "real_usd": 1.155, "request_id": "", "status": "delivered"}
```
`RUNS.md` gets one line per run: the run id, the script path, the arguments file, what it does and when it started.
## A review
What a reviewer returns per chapter. The verdict is computed from it in code.
```json
{"owner": 7.5, "craft": 7, "sync": 8, "story": 7, "pass": true,
"fixes": [{"shot": "s045", "time": 101.2, "severity": "blocking",
"problem": "the coverline sits over her mouth for 6 frames",
"fix": "move the column right (x 196→250) so it runs 40–80 px under her hair"}],
"checks": {"lyric_gate": "132/132 tokens on screen at their time (41.2s)", "sheets": 28,
"ms_per_frame": 820, "flash_max_per_s": 2}}
```
`owner` is the director's taste, and it counts double.
---
From the Claudia wiki: https://claudia.gallery/wiki/formats/
---
# What went wrong
Every failure that cost the films time or money, as what you see, why it happened, and the rule that came out of it.
Each row cost at least one round-trip, and some cost hours or hundreds of credits. Each one left a rule behind. "Another
project" is a production made with the same tools between the films.
## Brief and taste
| What you see | Why | The rule now |
|---|---|---|
| v1 rejected for its look: "way too.. gray", "color grading looked insane", "dont love the dot grid" | the look was assumed, not asked | Ask up front: native colour, one grade, or a print texture. Record the answer. |
| Plates "feel like a stock photo" | generic subjects: crowds, elderly or patient portraits, hands holding props | Change the idea entirely, not the seed, or make the shot pure motion design. |
| A sequel's song feels like a copy | it reused the last song's device, a countdown | A sequel never reuses the previous film's device. |
| Viewers read the lead as a teenager | age words and schoolgirl-coded wardrobe | Write an explicit adult age, use editorial wardrobe, and never write "young". |
| A taboo word turns up in a label or a file name | the taboos lived only in the brief | Put taboos in the project file and in every agent's prompt. The audit checks the board; grep code and file names before release. |
| "off pace and off vibe with the song" | the busiest picture sat on the quietest music | Measure picture motion against stem energy per section. |
| Agent-written style rules pulled work away from what the director liked | rules without a verdict behind them | Keep exemplars and verbatim verdicts. No style rule without the director's yes. |
## The song
| What you see | Why | The rule now |
|---|---|---|
| The song came out twice as long as planned (2:19 → 5:06) | Suno gives each spoken line its own two-bar slot | Plan two bars per spoken line, and write 10–13 syllables with the stresses on the beat. |
| Verses went choppy after optimising a score | a metric that rewards short clauses on the pulse | Use flow for verses and groove for chants, calibrated to the director's ear. |
| A name mispronounced | Suno's reading | Respell it in the Suno field only; the screen keeps the real spelling. |
| Numbers and acronyms sung wrong | written as digits | Spell them as sung ("A G I") and show "AGI" on screen. |
| The lyric rushed or truncated, silently | the lyrics field ran over about 3,000 characters | Aim for 1,800–2,400 characters for about 2:20. |
| The chosen take can't be fetched | the download is a native browser download, and the CDN wants signed URLs | The person downloads it; that download is also the licensing step. |
| "A voice already exists for this song" | Suno allows one Voice per source song | Make one combined Voice, spoken and sung. |
| A patched export drifted the song by up to 0.7 s | an editor export with BPM lock on | Export with the lock off, diff the spectrograms, and time-map the edit instead of re-cutting. |
| A new take changed the song's shape | takes differ in structure, not only in timing | Re-time the board from word anchors, and restructure by hand where the shape changed. |
## Listening
| What you see | Why | The rule now |
|---|---|---|
| Cuts drift off the beat later in the song | a fixed BPM on an accelerating take (131.5 → 133.9) | A variable beat map, never one BPM. |
| One chapter's lyric check failed on a word 2.6 s early | a delay throw aligned as the word | Ear-check gaps over 1 s and very low confidence before the board. |
| Bars start on beat 2 | claps on 2 and 4 won the downbeat vote | Add a chord-change vote, check by events, and keep an override. |
| The loudest moment gets no camera impact | the detector missed a hit that grew out of a loud chorus | Keep a manual add and drop list of hits. |
| Snare and hat lists full of kick clicks | a broadband kick fires in every band | Drop onsets within 45 ms of a kick. |
| Sung lines and shouts missing from the alignment | the aligner is weak on singing; the transcriber missed shouts | Merge the two aligners by confidence. Copy a shout from the earlier chorus plus the bar offset, and mark each word's source. |
## Her and the plates
| What you see | Why | The rule now |
|---|---|---|
| Her ethnicity drifts; 12 plates flagged | the prompt didn't pin it | The canon text pins ethnicity, age and skin in every prompt. Test four plates before any batch. |
| Identity edits drift from the look, and refuse seated poses | an image model "replacing only the woman" | Identity is prompt text only. |
| A close-up-sized head on a distant figure | face image references: a whole sheet imports its layout, a single face over-conditions | No face references for identity. |
| "Version 8.2 doesn't support --oref" | that parameter doesn't exist on that version | Don't use it. |
| Full-length prompts come back as close-ups | the long face canon pulls the frame in | Camera first, plus the short far form of the identity. |
| A feature asked for as an adjective goes missing | it wasn't the subject of the sentence | Make the feature the subject ("the hide studded with dozens of large glossy eyes"). |
| The crew are copies of the male lead | one description for a whole cast | Describe each extra (age, build, hair) and add an explicit negative. |
| A sheer gown read as plastic wrap and nude | "sheer … bare shoulders, bare feet" | A structured, high-neck, long-sleeved look, "composed and deadpan". |
| A stranger in an over-the-shoulder shot | the shoulder had no canon | Over-the-shoulder shots carry the canon too. |
| The robotaxi came back as an old sedan | the object's form wasn't named | Name it: "a small white compact electric SUV robotaxi with a tall spinning lidar dome". |
| The same image on about six shots | stand-in aliases left in the board | One image per shot; use the unpicked candidates as a variety pool. |
| "Two heads next to each other" ("bodyslop") | a ghost and a person composed in one prompt | Describe the ghost as made of the scene's light, or compose two assets. Make groups with an edit model. |
| A fix painted by hand in code | fixing the frame instead of the source | Bake it into the prompt or use an image-edit model. Never paint pixels. |
| The edit model invented a six-pointed star | edit models invent details | Re-check the canon after every edit, and name the exact form ("exactly eight points"). |
| Edited plates rendered with their old depth | prep wasn't re-run | Re-run depth and boxes after every plate edit. |
| Images matched to the wrong prompts | jobs identified by position, and Midjourney rewrites long prompts | Identify each job by its own prompt text. |
| Lost picks | parallel runs racing on one state file | Use a locked read-modify-write. |
## Clips and lip sync
| What you see | Why | The rule now |
|---|---|---|
| Lifeless takes, about 830 credits burned | agent-written prompts: poses, not drama, and nobody read them | One writer writes every prompt in one context, with the story beat, blocking, and timestamped action and reaction. The director reads the book before any spend. |
| Sung takes re-sing the line instead of copying it | long timestamped screenplays on sung windows | Use the short sung form for sung windows, and screenplays for motion. |
| One window barely syncs whatever the prompt | a microphone at her mouth on the plate | Sung plates show the whole mouth, with nothing in front of it. |
| QA kept weak takes | takes were checked against their own weak prompts | Check takes against the story and the board. |
| Two clips start on the same frame ("jarring") | the same plate as the start image | Use a sibling candidate, or the previous clip's last frame. |
| Lip-sync models "still not exact" | the failure was the input, not the model | Measure, cut, cover; otherwise cut away. No lip-sync models. |
| A take diverges inside a word | the prompt quoted fewer words than the window held | Quote exactly the window's words, and start windows in word gaps. |
| Lips a fraction early or late in the edit | each take has its own lag, up to ±1 s | Map clip time and song time through each take's measured lag. |
| A window cut a word ("…could. Your P") | the window ended mid-word | Start and end windows in word gaps. |
| Motion clips "diverge" in the sync check | the mix isn't copied (NCC about 0.08) | Expected: play motion clips at lag 0. |
| Grey proxy shapes in a take | a 3D previz video reference leaked into the output | "take only the camera move from it", or drop the reference. |
| Drift on long sung clips | lip sync degrades over length | Sung windows of 4–8 s, never longer. |
| Head counts change, actions reverse, garbled text | under-specified prompts | "exactly four …", name the end state, one dominant camera move, "No text, letters or logos". |
| "Upscaled" footage looks waxy and shimmers | learned upscalers | Keep Lanczos, and generate hero shots at 1080p instead. |
| 32 of 59 clips came back near grey | references pre-graded to monochrome | Send references in colour; any look goes on in post. |
| A sung clip ignores its audio | a first frame sent with the audio reference | On OpenRouter, use reference mode for sung clips and covers, never `frame_images`. |
## Filters and providers
| What you see | Why | The rule now |
|---|---|---|
| Close-up references declined | the provider's face filter | Free, and final for that image on that service. Send the same unaltered image elsewhere, or keep the 2.5D plate. |
| A retry ladder (stronger grade → posterised grade) got takes through, baked banding into them, and was blocked by the harness | altering a reference to pass a safety filter | Never alter an image to pass a filter. A blocked action is a stop, not a puzzle. |
| Jobs fail with no reason, credits back seconds later | a silent input filter misreading innocent wording ("a plug pushed deep into each ear" for foam earplugs) | Reword plainly and keep the beat: plain rewording fixed five of six at once. |
| "output audio may be related to copyright restrictions" | the provider's audio filter | Retry once with a re-cut window, then try the other provider. |
| A filter that passed faces two days ago now declines them | providers change | Keep a second provider configured. |
| The agent's context floods during uploads | presigned-upload replies of about 3.5k characters per file | Upload by URL. |
## Spend
| What you see | Why | The rule now |
|---|---|---|
| Every submit returns 403 while the account has credit | the API key has its own spend limit | Set the key's limit above the project cap. |
| "Remaining budget" figures are wrong | the local ledger isn't the bill, and it counted free declines | Quote spend only from the provider. |
| Hundreds of credits spent in minutes on unread prompts | a credit-reset deadline drove the spending | Never let a deadline drive spend. |
| A background workflow could still spend after a quality flag | not every spender was stopped | Stop every workflow that can spend: search each for its generate calls. |
| Credits vanished at renewal | subscription credits don't roll over, and top-ups expire | Plan spend inside the period. |
## Colour and look
| What you see | Why | The rule now |
|---|---|---|
| v1 "way too gray … loses the Midjourney magic" | a mono print pass, and chapters desaturating the footage | Colour halftone or native colour; never desaturate the footage. |
| "color grading looked insane"; blue graphics "clash" | a graded look, and a blue accent over blue footage | Native colour; pick the accent against the footage's hue; draw with semantic colour tokens. |
| Flicker in a graded clip | per-frame grading | One LUT per clip, computed from 16 sampled frames. |
| The master looks washed out ("v blown out") | an untagged or BT.601 stream | Convert to BT.709 and tag it; verify with ffprobe. |
## Motion design and the renderer
| What you see | Why | The rule now |
|---|---|---|
| A crop cut her head off after a re-take | code tuned to the old take | Re-frame every shot whose take changed. |
| The previous frame's overlay "sticks" in dense frames | a GPU-backed 2D canvas under software GL | Software-backed canvases, and the history check in QA. It took about two hours of bisecting to find. |
| Frames differ depending on what rendered before | a WebGL canvas drawn into a 2D canvas switches it to GPU rasterisation | Read 2.5D pixels back through the CPU. |
| 0.5–1.1 s a frame on 2.5D shots | the 64-step parallax shader | Depth slices or multiplanes. |
| Tags never finish typing on short shots | the default typing speed | Type at 16–40 characters a second, and finish 6 frames before the cut. |
| A lyric word half hidden on its first frame ("You dor't") | occlusion on the entering row | Tuck an edge only, and switch occlusion off for the entering row. |
| Five accent-coloured elements in one frame | the accent used for every highlight | The accent goes on at most two elements. |
| A snap reads early | keyed 70 ms before its kick | Key moves on hit times from the hits file. |
| Two on-screen numbers disagree ("x380" against "135 RUNNING") | two devices driven separately | Drive both from one value. |
| Shake in a drumless intro | a move that answers no sound | Moves answer audible sounds. Stillness where the song breathes. |
| The same helpers written 8–10 times, slightly differently | chapters couldn't share code | One shared device library. Chapters write only what is unique. |
| The default renderer drew only lyrics | the board's graphics were stored as prose | Structured fields in the board: type, devices, treatment, transition. |
| A lyric time patched in chapter code | fixing the time instead of the alignment | Fix the timed words and re-export. |
| A chapter replaced the default renderer and lost lyrics | the chapter must draw every token itself | Run the lyric gate per chapter, and mark custom-drawn tokens. |
| Fonts that can't be shipped | AGPL, GPL, EULA or unlicensed fonts | Ship OFL fonts only, with their licence texts. |
| A depth model that forbids commercial use | Depth Anything V2 Large is CC-BY-NC | Use Depth Anything V2 Small, which is Apache-2.0. |
| A singing mouth held where the take isn't synced | the edit fell back to a plate | Crops, inserts or depth moves on those words. |
## Review
| What you see | Why | The rule now |
|---|---|---|
| Runs "completed" with no review at all | a usage limit made reviewers return nothing, and the script read that as "no fixes" | A null review is an explicit failure, in code. Carry open fixes forward. |
| Reviewers scored 6–7.5 a cut the director called "perfect" | a stricter bar than the director's | Ship on "no blocking fix left". Calibrate on the director's verdicts. At most two rounds. |
| A reviewer re-ran checks until its time ran out and wrote nothing | it wasn't told what was authoritative | Bound reviewers by output: "write the short version first". |
| The scan flagged an intended prop | a generic "modern object" rule | Whitelist intended anachronisms by name. |
| A strobe counted twice | the flash check paired every transition | Count opposing pairs. |
| Twelve blocks found in 96 "final" takes | generated people; filters let outputs through | A safety and canon scan of every take before release. |
| First-draft chapters scored 5–6.5: faces printing blank, type over faces, repeated devices | animators with nobody reviewing real renders | A reviewer renders, looks, and writes concrete fixes. |
## Render and running agents
| What you see | Why | The rule now |
|---|---|---|
| Nobody knew the encode had finished | it ran as a detached job | Tracked background jobs with a status. |
| Swap and the root disk filled up | six agents rendering at once, with frames in /tmp | One rendering workflow at a time, scratch on the big disk, and check free memory and disk first. |
| Disk full after a render | about 9.6 GB of frames per cut | Delete the frames once ffprobe has verified the outputs. |
| Commands ran against the wrong project | an export inside a backgrounded shell chain | Name the project on every command. |
| A merge restarted three times (about 1.5 h) | an interrupt kills every running workflow | Short phases, run ids saved, resume by run id. |
| One serial chapter workflow would have taken about 12 h | serial agents | Three parallel workflows (about 4 h). |
| A ~20-hour background queue, new takes stuck behind superseded ones | post-processing auto-queued on every import | Queue only takes the edit uses, in priority order. Cancel superseded jobs. |
| A GPU box at 100% CPU, every job 3–5× slower | two dozen waiters polling every 3 s | One shared listing every 45 s for all waiters. |
| Duplicate jobs after a restart | killing a waiter left its child behind, and the re-run resubmitted | Record job ids beside their inputs, adopt queued jobs, and kill orphans by id. |
| Every new agent session fails at start | an agent CLI auto-updated mid-run | Pin versions for long jobs. Check the job's error before assuming there is no work. |
| A worker probed for 75 minutes and delivered nothing | no bound on the first deliverable | A first complete draft within about 10 commands, with a mandatory handoff. |
| Shared-file bugs fixed differently in parallel chapters | workers patching shared code | Workers report shared bugs; the operator fixes them once. |
## Publishing
| What you see | Why | The rule now |
|---|---|---|
| Identifying details in a making-of: names, profile codes, hosts, job ids | raw production notes | Scrub rules, a grep pass, and two fresh-eyes reads. |
| A song file carries the generator's song id in its tags | file metadata | Strip metadata from everything you publish. Midjourney images carry their job id in XMP. |
| Real-person likeness risk | generated portraits of real people | Keep generated likenesses of real people out; quote their words as text, with the source. CLODYSSEY's labelled AI IMAGE post cards are left out of this site. |
| A master posted without the AI disclosure | platform toggles missed | Set each platform's AI-content label, and name the tools. |
---
From the Claudia wiki: https://claudia.gallery/wiki/failures/
---
# Use her
Claudia is an open character. What you can do with her and the images here, and what isn't covered.
Claudia is an open character. Make images, videos, songs, stories, games or merch with her, commercial or not. Use the
images on this site as references, as image prompts, or as material. You don't need to ask. A credit ("Claudia by
anabology") is appreciated, not required.
## What helps
- **Keep her herself.** The [canon](https://claudia.gallery/claudia/) is short: the black bob, one clay streak, the star clip, the mic, an
adult face. Change everything else.
- **Make her an adult, always.** She is in her late twenties. Nothing that reads under-age, and nothing explicit.
- **Say what's generated.** All three films were made with Midjourney, Seedance, Suno and Claude, and said so.
## What isn't covered
- **Other people's things in the films.** The Bliss photograph and the XP sounds belong to Microsoft. Real public figures
appear in CLODYSSEY only as posts labelled AI IMAGE; they are not part of this gallery. The anime idol in CLODYSSEY is
another creator's character.
- **Anthropic.** Claudia is fan work about Claude. She is not affiliated with or endorsed by Anthropic, and Anthropic's
names and marks are theirs.
- **The director's accounts.** The Midjourney profile and the Suno Voice used in the films live in personal accounts and
can't be shared. Everything they made is here.
## Credits
Directed by **anabology**. Made with **Claude** (Anthropic's model, in Claude Code) as the production team, **Midjourney**
v8.2 for every image, **Seedance 2.5** for the moving shots and **Suno** v6 for the songs. This site was built by Claude
from the films' own records.
Type: TeX Gyre Heros (GUST Font License), DINish (SIL Open Font License), IBM Plex Mono (SIL Open Font License).
Player: hls.js (Apache 2.0).