Operator guide
Six steps, in order, because the order is the method — a seed is chosen before a set is rendered, and a set is approved before a threshold can be measured. Each step permanently constrains the next.
Every field note here is something that actually happened while building this, with the number it cost. They are the only parts of this guide that could not have been written in advance.
The six steps
Traits, tiers, and the words the model actually listens to.
STEP 02One frame decides everything downstream. Judge it on two things, not one.
STEP 03Twelve frames across four lights. Why the lighting spread is the point.
STEP 04The one step no measurement can do for you.
STEP 05Measured against strangers, never chosen.
STEP 06Scoring every scene, and reading the ledger afterwards.
Finish a step before starting the next. Every one of them writes something the following steps read, and several cannot be undone without discarding work you have already paid for.
A character starts as a written declaration: a descriptor in prose, and a set of traits each placed in one of three tiers. The descriptor is frozen at creation — later prompts append to it and never rewrite it, because prompt drift is the second way a character slips. Reference images hold the face while the text quietly changes the build.
The tier belongs to the trait, not to you. An immutable trait defines the character and cannot be overridden by any scene. A controlled trait can be moved for one render, and only by an explicit declaration. A free trait — wardrobe, makeup, accessories — varies scene to scene and is never checked. Putting a trait in the wrong tier is the difference between a gate and a wall: judge wardrobe and every scene that did its job fails.
A character declared with “a full, rounded build” rendered slim, repeatedly, even at full prompt length. Rewritten as “a plus-size body, heavy through the hips and midsection”, she rendered correctly on the first attempt. The model was not ignoring the instruction; it was not recognising the phrasing. Describe observable geometry, not a label.
Name what a photograph would show: height in feet, where weight sits, hair length and texture, the shape of the jaw.
Use category words a person would use in conversation. Some are ignored; others are refused outright by the provider and return blank frames you are still billed for.
Candidates are rendered from the descriptor alone, with no reference image, so each is an independent draw. They are not meant to look like each other and measurably do not: across our cast, candidate frames of the same character score as low as 0.18 similarity to one another. Identity-locking has not started yet.
You pick one. Every canonical frame is then conditioned on it, the centroid that judges every future render is computed from those frames, and nothing downstream can tell a good seed from a bad one. This is the least reversible decision in the process.
Judge it on two things. The first is obvious: does it depict the character you declared. The second is not, and it costs money to learn.
We picked the frame that best matched the declaration: a short, heavily muscled man in his thirties. Its skin-tone measurement carried a confidence of 0.457 against a floor of 0.6, so every frame conditioned on it produced readings nothing could trust, and the set failed. The seed we should have picked measured 0.794. The best-looking candidate and the most measurable candidate were not the same frame, and only one of those facts is visible by eye.
A seed that renders a beautiful frame the measurer cannot read is a seed that produces a character you cannot prove anything about.
Twelve frames, conditioned on the seed, spread across four declared lighting conditions and several angles and framings. The lighting spread is not variety for its own sake: a set shot under one light bounds no tolerance. The spread is what the tolerance is later derived from.
Frames here are measured, not judged. There is no calibrated tolerance yet — the set is what produces it — so judging the set against a provisional guess means failing frames on a number those same frames are about to disprove.
Our provisional skin-tone tolerance was 12.0 ITA. Calibration, run on the very frames that tolerance had just failed, derived 50.83. The first character survived only because her lighting happened to move the measurement by 9 points; the second character's moved by 101, and every frame in his set was marked failed. The canonical set is no longer judged this way.
Some frames will render fine and still be unreadable — a face turned into shadow gives a skin-tone measurement nothing should trust. Those are set aside rather than discarded: they are kept in their own section with the measurements that did succeed, because the render worked and the picture is usable even though its record is unfinished. They are not part of the canonical set and cannot become references.
Every frame of one character under dim low-key indoor lighting was too dark to read a skin tone from. That is a fact about that character under that light and re-rendering does not fix it. A condition that was attempted and yielded nothing measurable now drops out, provided enough remain to bound a tolerance at all.
This is the step no measurement can do for you, and the reason the system has a human gate at all. You accept or reject each frame individually. Accepted frames become the reference set — the face the character now is — and every later render is compared against them.
Judge the picture first. The numbers under each frame are evidence for that call, not a substitute for it: a frame can clear every threshold and still be the wrong person, and the whole reason a human sits here is that the machine cannot yet tell you which.
A rejected frame is informative, not wasted. Rejections become negative controls, and the gap between what you accepted and what you rejected is part of what calibration reads.
A similarity score on its own means nothing. Without knowing what a stranger scores, 0.6116 could be a strong match or a coincidence. So the character is scored against a cohort of 90 people who are not her, and separately against her own approved frames, and the threshold is placed in the gap between those two distributions.
It is reported with the error rates it produces — how many strangers it wrongly accepts, how many of her own frames it wrongly turns away. A threshold quoted without its error rates is a number, not a measurement.
Thresholds are not comparable between characters. Each is measured on one person against one cohort. Borrowing another character's number produces a verdict with nothing behind it.
From here every scene is scored on the way in, against the threshold measured for that character. Frames below it never become assets, and the reason is written to the ledger beside the cost — so a rejection is auditable months later, by somebody who was not there.
Captions are checked against a declared voice and a single never-say list before anything is approved for publishing. One list, every character, one place.
Spending is bounded before it happens: a daily cap checked inside the transaction that queues the job, a per-run ceiling on frames, and a price shown and confirmed before any batch is submitted. The cap bounds the mistake; the confirmation is what stops it being made.
Four canonical-set attempts failed while we were building this, on four different faults. Each released its reservation and settled only what it had actually rendered — a total of 0.66 USD across all four. A ledger that survives being wrong is the one worth having.
Motion, publishing to a platform, and reproducible seeds do not exist. Calibration has been run on one character. The system page carries the full inventory, including which identity signals gate a frame today and which are only recorded.