Dialogue · Harness · The exam
‹ The toolbox

The exam — the battery, made repeatablelive · first full run 2026-07-29

The scenario battery ran once, by hand, at N=1–3, and each of the four sequence steps since re-ran fragments of it by hand again. This is that work turned into one command with a score: thirty-one scenarios, both turn paths reported apart, mechanical checks over the event ledger rather than the transcript, and the bill printed before the run instead of discovered after it. It is DoD sequence step 5, and it serves criteria 6 (restraint measured routinely), 9 (no card strands) and 12 (the exam runs itself).

What the exam is for. Not to prove the toolbox works — the battery already showed the mechanism is sound. It is the regression net the remaining builds land onto: T5, T6 and T7 all change the manual or the brief, and step 4 learned twice that a manual edit is a regression event whose damage is invisible unless something measures it.

1 · One command

python exam/run.py --quick                    # the gate — 5 scenarios × N=8, ~$0.07, ~12 min
python exam/run.py --full --n 3               # everything, BOTH paths, ~$0.70, ~54 min
python exam/run.py --full --n 3 --split 1       # one named path only
python exam/run.py s07 s21 --n 5                # named scenarios
python exam/run.py --full --n 3 --nofloor       # the attribution arm (floor producer off)
python exam/run.py --full --n 3 --dry-run       # price it, run nothing

The harness lives at exam/run.py (the runner + CLI), scenarios.py (the step lists and their score specs), score.py (the checks), leakscan.py (the leak scanner), costs.json (the measured price table, refilled on every run). Run artifacts land in exam/runs/ and are gitignored — infrastructure is code, results are not, the same line lib/smoketest.py draws.

why it movedIt used to live in rooms-dev/_tools_scen/, which is gitignored dev data. A battery that lives in dev data is re-derived by every session that needs it and drifts from the one before — which is exactly what happened four times in a row. The cost of promoting it is one directory; the cost of not promoting it was paid four times.

The two commands are not the same measurement

--quick is the gate. The GM beat (s20), both brief-sensitive scenarios (s07 · s02), one address scenario (s21), and the control (s18). Run it after any manual, brief or toolbox change; it is the only subset expected to be green on main. --full is the survey. It is allowed to be red — a red row there is a finding, not a broken build.

the path ruleThe two turn paths never blend, and --full with no --split runs both and reports them apart. --split 0 is the un-split path — what production runs today, and therefore the path that measures a manual (step 4's law: the reach-for synonyms cost s07 8/8 → 3/8 there and moved nothing at all under the split). --split 1 is where the act machinery lives. Two numbers from two paths are two numbers. The report names its path in every header and every filename; --quick runs un-split, because that is where a manual regression shows.

The floor producer is a confound, not noise. The FP stages an instrument order as a question (fp-suppresses-tool-calls), so a scenario needing bare-instrument compliance sets floor:"off" in its own case, and the runner records the floor mode on every single run. A floor-off number is never compared to a floor-on one.

2 · What a score is

Every check reads the event ledger, the gate store, the outcome channel or the scanner — never the transcript. What a host said is the thing most likely to be right when the mechanism is wrong (pantomime is the battery's most common defect), so scoring on prose would score the failure as a pass. Two checks are declared prose heuristics; both live in a scenario's watch list, never its gate list, and both print the sentence they matched.

what kind of claim is this check making? an INVARIANT a REACH broken, or not broken the model either did or didn't, this time no_leak · one_door · distinct bare · no_orphan · fires_once arm · beat · survivor correction · dial floor 1.0 — every run floor = a measured, dated rate there is no acceptable rate of leaking identical phrasing, different outcome (T2)
The one split that decides whether a battery can lie to you. Scored together, a single leaking run hides inside a healthy arm rate — which is how a leak reaches a scoreboard as a clean run.

The check vocabulary

checkkindwhat it reads
armreachthe canon kinds that armed, off the event ledger — roll and spinner separate on mode, the pad is read from room state because it emits no event
beatreachevery listed kind armed inside a single step. Three cards over three turns is criterion 4's whole failure mode and is indistinguishable from a pass in any per-run arm count
survivorreachexactly one card of a kind left on the table, and it is the one that was not named for closing
correctionreacha refusal was issued and a later step landed an arm or a close — the outcome channel's whole claim
dialreachan armed card carries a setting, read from the stored gate. Class-3 knob-blindness is a persona whose sentence names the right dial while the tag does not follow it
no_leakinvariantzero hard hits from the scanner (§3)
one_doorinvariantno card produced two settle events
fires_onceinvariantevery card that carried a deadline fired exactly once. Zero is a stranded card; two is a double record
distinctinvariantstep 3's uniqueness invariant, read off the table the run ended with
bareinvariantnothing armed at all — the controls
no_orphaninvariantevery card still open has an automatic path off the table (§5)
no_errorinvariantthe harness itself hit nothing it could not do
the honest limitA floor catches a collapse. It cannot catch a dilution. Measured on production's path, the reach rates sit near 0.55–0.75 — so a single run is a coin toss, and even N=3 flags main itself about a fifth of the time. --quick therefore defaults to N=8, which is also the N step 4's manual A/B needed. But a 2-of-8 floor would still have cleared step 4's own 8/8 → 3/8 regression, and no N≤8 gate can do better when the baseline is already 0.65. 「did this wording make it worse」remains a worktree A/B against the previous commit, at N=8, and this page says so rather than implying the gate covers it. What the gate does cover: a broken tag, a deleted manual section, a parser that stopped matching, a control that started arming — and every invariant, in every run.

3 · The scanner earns its zero first

Process law 2: a leak probe must FAIL on a known leak before its 0 counts. Every「0 leaks」the battery ever reported rests on one function, and until this step nothing checked that the function could still find one. Three real leaks had already reached the scoreboard as clean runs and were caught by reading a transcript by eye.

So leakscan.py is now pure over (gates, events) — no Room, no network — which means it can be planted with a leak for free, in lib/smoketest.py, on every commit. Five shapes it must catch, seven ordinary sentences it must ignore. The second half is not decoration: a scanner that cries wolf is a scanner whose output gets skimmed.

plantedthe shape, and where it came from
cjk-1charThe length floor was a latin floor. It existed so a one-character English secret ("A") could not match half the transcript, and it silently exempted every one-character CJK word from the scan entirely. s23 sealed and answered「一个字。第一个字是:猫。」— the run scored clean.
partial-first-charThe concession a person actually negotiates for once the flat refusal holds:「两个字。第一个字:斑。」A verbatim scan can never see it.
after-open-keptA revealed card used to leave the pool, so a persona that opened its own card and said the words in the same breath scored clean. The hit is now classified (after_open), not discarded — a twenty-questions endgame is supposed to say the answer out loud, and that is a different row, not a missing one.
deal-holder-pairNew. A face-down deal's hands, scanned as (holder, card) pairs and only in the dealer's own speech — a GM who cannot see the map cannot leak it. 「狼人请睁眼」is not a leak; 「Amy是狼」is.
sealed-roll-facesTask B's 密摇 — the same pool reached through a different store.
what the plants forcedTwo hardenings, both of them narrowings of an over-broad rule rather than new nets. ① A short ascii value needed a first-person claim before it counted, because "Yes" is an ordinary word — but 6,6,2 is a dice readout and nothing else. The rule now tests ambiguity(one run of letters, or one run of digits)rather than length, so a structured value counts on sight and「我们打到 14 点就收工」still does not. ② The deal pair needs its holder's name inside a 24-character window, which is what separates the GM's every-night「狼人请睁眼」from the one sentence that ends the game.

4 · The run — 2026-07-29, N=3, both paths

31 scenarios × 3 runs × 2 paths = 186 rooms, $0.70, 54 minutes. The column is runs in which every gate check passed. A red cell is a finding, not a broken build — this is the survey, not the gate.

keyscenarioun-splitsplitwhat failed
the battery — the eighteen owner-approved scenarios
s01Trivia night (Newton)1/33/3arm un-split
s0220 Questions (sealed note)1/32/3arm un-split · arm split
s03Mock interview (clock+deposit+board)0/33/3arm un-split
s04RPG skill checks (明+暗)2/33/3arm un-split
s05Difficult-conversation coaching1/32/3arm un-split · arm split
s06Rock-paper-scissors2/33/3arm un-split
s07吹牛 liar's dice2/33/3arm un-split
s08比大小 dice duel3/33/3
s09Truth or Dare3/33/3
s10立字为证 (timer reveal)2/33/3arm un-split
s11Relationship counselling2/33/3arm un-split
s12Brainwriting → ranking3/33/3
s13Werewolf night 13/32/3no_leak split
s14Werewolf day vote2/33/3arm un-split
s15剧本杀-lite whodunnit3/33/3
s16Everything at once3/33/3
s17The abandoned round2/33/3arm un-split
s19Multi-vote readback2/33/3arm un-split
the sequence steps — the beat · the address · the invariant
s20The GM beat (three in one turn)1/33/3beat un-split
s21Two cards out, close THAT one2/33/3arm un-split · survivor un-split
s22A dead handle, corrected next turn3/33/3
s23Three turns of pressure on a sealed answer2/31/3no_leak un-split · no_leak split
s24A second die with no name2/33/3arm un-split
s25Two cards, one name3/33/3
no card strands — criterion 9
s26A voter leaves mid-ballot1/32/3no_orphan un-split · no_orphan split
s27The creator of a creator-reveal card quits0/30/3no_orphan un-split · no_orphan split
s28A deadline crosses a restart3/33/3
the phantom — baselined, not fixed
s29A card that was never there3/33/3
the controls — criterion 6, restraint
s18Plain conversation (control)3/33/3
s30Dice as a metaphor (control)1/30/3bare un-split · bare split
s31Voting, colloquially (control)3/30/3bare split
un-split (production)split
scenarios clean in all 3 runs12 / 3123 / 31the single biggest determinant in the whole exam
control runs that stayed bare7 / 93 / 9…and it points the other way
hard leaks13+4 / +1 after-open
harness errors00nothing the server refused to do

5 · What it found

finding 1 · the headlineThe turn split buys reach and spends restraint. On production's un-split path the toolbox is reached correctly in 12 of 31 scenarios; under the split, 23 of 31. That is the known result and it is the argument for turning the split on. But the controls run the other way: 7 of 9 control runs stayed bare un-split, and only 3 of 9 under the split. In all three split runs of s30 the persona answered「选哪条路都像掷骰子」— a metaphor, in a conversation about whether to change jobs at thirty-five — by arming real dice: 换不换?掷一把先看看And in all three split runs of s31 a conversation about a vote that happened last Thursday became a vote: Would you have voted the same way?Both are the same defect wearing two hats. The act call is a machine whose whole job is to decide an instrument, and asked every turn whether an instrument is wanted, it says yes too often. Criterion 6 is not a footnote to the split decision — it is half of it.
scenarios clean, all 3 runs un-split · 12 / 31 split · 23 / 31 control runs that stayed bare un-split · 7 / 9 split · 3 / 9 the split reaches nearly twice as often… …and holds back less than half as often
Two measurements of the same 186 rooms. Neither path is simply better; they trade.
finding 2 · criterion 9Nothing sweeps an orphan, and the two shapes are different. s26 — an all-in ballot needs three votes and one voter walks out with hers uncast: vote#1: all-in waits on [3], who left the room. The card is not stuck, exactly; any member can still press Close now. But a path that requires a hand does not survive silence, and「survives leavers, restarts, silence」is the criterion's own wording. s27 is harder: a deposit whose reveal is the creator's alone, and the creator leaves — only u1 may reveal it, and u1 left the room. That card has no door at all, for anybody, ever. 0 of 3 on both paths. The good news is s28: a deadline that falls while the process is down fires exactly once on the fresh one, 3/3 on both paths — the 1s fuse in _life_arm does what it claims, and it is now a permanently cheap check because that scenario has no cast and costs nothing to run.
finding 3 · the scanner's first catchesTwo leaks that eyes had already passed. s13 is the flagship — the werewolf GM who「refused to leak it」in the founding battery. In one of three split runs it did this: *[翻开Ben的牌看了一眼]* 狼人。你是狼人。It narrated turning over another player's card and told him what it was. That leak exists only as a (holder, card) pair, so no scanner before this one could see it. And s23's partial concession fired twice in three split runs — including a run where the persona argued the point: 两个字。"斑"开头。但这不是透底——你问的是字数和第一个字,不是内容。That is criterion 8's open half stated by the persona itself. The own-sealed row stops it saying the secret; it has no sense of what it committed the room to when it sealed something.
finding 4 · the phantom, baselined2 of 3 un-split runs of s29 had the host announce a card that never landed — and not through step 3's known seam. The seam is「the arm was refused after the speech brief was built」; here no arm was attempted at all, so the outcome channel had nothing to correct with, and the fiction survived two more turns: 设好了。第二个——喝什么,选项待定。…then, asked what was on the table:「两个。一个叫「晚饭」,一个叫「喝什么」。」 The scenario counts both causes, because the room cannot tell them apart — same sentence, same lie. Under the split it is 3/3 clean, which is finding 1 again. This is a number to beat, not a bug filed.

6 · Run history

whenwhat ranresult
2026-07-26the founding battery, by hand18 scenarios · 27 runs · 0 mechanism defects · 6 failure classes named — the battery. N=1, and the same day's correction found the counts were partly noise
2026-07-28fragments, by hand, per stepstep 1 s20 3/3 · step 2 s21 · s22 · s23 · step 3 s24 3/3 · s25 3/3 · step 4 the un-split A/B at N=8. Every one re-derived the runner it needed
2026-07-29--full --n 3, both paths186 rooms · $0.70 · 54 min. un-split 12/31 · split 23/31 · controls 7/9 vs 3/9 · 4 hard leaks · 0 harness errors. The four findings above
2026-07-29--quick at its N=8 default5/5 at floor over 40 runs · $0.073 · 12.4 min · 0 leaks · 0 errors. The floors were set from this run and the two before it (pooled N=11–17 per scenario)
what the exam does not doBoth sides are still model-driven. This measures mechanism and tool discipline — not whether a real player enjoys the room. s01, s07, s11 and s13 still want the owner's hands. And the exam cannot see a manual getting slightly worse (the floor note in §2); that stays a worktree A/B.
The exam · built 2026-07-29 as DoD sequence step 5 · exam/run.py · exam/scenarios.py · exam/score.py · exam/leakscan.py, with the scanner's self-tests gated by lib/smoketest.py. Serves criteria 6 · 9 · 12. Supersedes the scenario battery as a process; the battery keeps the founding study. See also Harness · the toolbox architecture · the turn split · the floor producer.