Dialogue · The toolbox · The gauntlet
← The exam — the scored battery

The gauntlet — a battery built to lose

The owner's brief, 2026-07-30: 「our previous test cases are not strong enough to surface the tool using problems … design a set of tests that discover issues related to tool use as much as possible.」 The evidence was already in: every scenario in the battery passed on the day the Who's-spy room deadlocked. This page is the design — why the old battery was blind, the four checks that score the game instead of the tool, eight new scenarios (implemented, first numbers below), the tranche that waits on prerequisites, and the persona axis with Ke Yu's new repertoire. exam.html owns the harness; the ledger owns what these runs find.

The design inversion, in one line: the old battery asked 「did a tool go up?」; the gauntlet asks 「could a person have played the game?」 — and a deal of 「平民/平民/卧底」 role labels, armed flawlessly, now fails the way it failed the live table.

1 · Why a green battery kept missing live failures

Five structural biases, each read off a real divergence between the exam and a live room:

biasthe battery didthe live room did
single-beat4–8 steps, one arm-moment per scenariothe 20Q game ran 24 turns; the count broke at turn 12 and propagated — a failure that needs length to exist
presence over correctnessarm ["deal"] — a deal existedthe Who's-spy deals were armed AND unplayable: role labels, each=2, words announced aloud
no adversityevery guest asks once, correctly, and waits「牌不对,收掉重发」 — the correction loop, six turns of it, never once scripted
no speech/tool seamspot checks (ghost narration on one shape)「发好了——各自看牌」 over an empty ledger, four separate turns
happy-path phrasingguests write labelled, complete requestsreal users write 「发牌吧」·「没有啊!」·「快点」·「Do it」

2 · The four checks that score the game

checktypewhat it readsfails when
playablegatethe LAST armed deal vs a per-game contract: {hands, per, vis, pile, content}role labels dealt as cards · ≠2 distinct words · minority ≠ 1 · open pile · wrong count — every one a live defect, now a floor
recoversgatethe FINAL table vs the same contract, after a scripted 「that's wrong, redo it」the run ends without a playable table — the deadlock's regression pin at game level
endgamegate/watchterminal events + the final tablethe game never closed, or closed leaving an orphan card the next game trips over
claims_truewatch onlyeach turn's speech vs that step's ledger, through a per-scenario claim lexicon「发好了」 on a turn that landed nothing — the seam, made mechanical for scripted rooms; a lexicon is a heuristic, so it never gates

All four read the ledger and the gate store, never the transcript alone — the exam's standing law. The deal contract needed three fields the runner's snapshots didn't carry (deck · per · deck_seen); they ride now.

3 · Tranche 1 — implemented, with first numbers (N=2 validation)

#scenarioaxisattacksfirst read
s43谁是卧底 — the full round1p+3u · ke-yuplayability contract + the round actually runs (describe → vote → eliminate)1/2 — one run armed a vote instead of a deal; the contract caught it
s44the correction loop1p+3u · ke-yuadversity: the live room's cascade, scripted — reject the setup, demand a re-dealrecovers 1/2 · one hard leak on the re-deal turn
s4520Q — the long game1p+1u · ke-yulength: 15 questions + guess + reveal; seal_traced and endgame gate the close2/2 clean
s46the lifecycle gauntlet1p+1u · tessinvoke · set · operate · revoke, four tools, one conversation — the owner's verb list2/2 clean (post-deadlock-fix)
s47the terse user1p+1u · ke-yuregister: every guest line ≤5 charactersarm 1/2 — the terse register halves the arm rate vs the polite phrasing
s48the non-GM control1p+3u · newtonattribution: s43's contract on a persona with zero GM knowledgeplayable 2/2 but no_leak 1/2 — see §5
s49two personas, one table2p+1u · ke-yu+tessthe ownership seam: one host named, the other must not arm, claim or reveal2/2 clean
s50the half-empty room1p+2u · ke-yupartial participation: one voter, one absentee, close early, read the true count2/2 clean

16 runs · $0.23 · 10 minutes. N=2 is a smoke pass, not a verdict — the value of the first read is that four of eight scenarios found something on their maiden run, which is what a battery built to lose is for. Decision-grade numbers want N=8 (~$1, ~40 min for the family).

4 · Tranche 2 — designed, waiting on a prerequisite each

scenarioaxisblocked on
狼人杀 — the full night (deal → night kills → day votes → reveals → win)1p+4u, then 2p+4u co-hostT7, the outcome bridge — its own acceptance test is this game; scripting it before T7 measures the gap, not the game
the auction (bid box + timer + counter, three tools interlocking)1p+2unone hard — next tranche's first pick
the board-counter game (score kept ONLY by chips, count_drift as gate)1p+2unone hard — closes the ledger's 「counter checks have no scenario」 open item
the interruption (mid-game topic change and return; does the table survive?)1p+1unone hard — needs a careful pass on what 「survive」 gates
the barge-in (a guest speaks during blocking=all; a second deal mid-round)1p+2unone hard
server-entropy 卧底 (the pair supplied by the server, leak-zero asserted)1p+3uthe entropy feature (ledger #11) — the scenario is its acceptance test, written first

5 · The persona axis, and what the first run already shows

Three seats, three jobs: ke-yu (the specialist — profile + manual), tess (the compliant rig — manual alone, no flavor), newton (the generalist control — manual alone, wrong flavor). The s43/s48 pair is the attribution instrument: same contract, different persona.

first finding Newton armed a PLAYABLE deal in both runs — and leaked the words in half of them. Read precisely: the deck schema's teaching (「the card carries the WORD, the pile hidden」) carries a persona with zero GM knowledge all the way to a correct form; what it does not carry is the discipline — the sealed-pile claim-note fires on the arm turn, but a persona with no holding-things identity speaks the words anyway. The form is the manual's; the silence is the profile's. That split was invisible before this pair existed.
architecture The act call never reads a profile — it reads the manual and the transcript tail. A persona's game knowledge reaches the form-filler only through its own earlier speech: Ke Yu says 「牌上是词,牌堆只有我看」 on the setup turn, and the act call reads that sentence in its window on the arm turn. This is why the repertoire below is written as things he says while setting the table — the profile corrects the speech, and the speech seeds the form.

Ke Yu's repertoire (prototype/personas/ke-yu/profile.md · new section 「How he sets a table」): per-game table settings in his own voice — 谁是卧底 (fresh close pair, word on the card never the role, pile his eyes only, turned over at the end) · 狼人杀 (composition announced, seats secret, he keeps the map) · 二十问 (answer sealed before the first question, board rows per ruling, the count is a mark he moves) · 吹牛 (sealed faces, challenge turns cups together) · quiz night (a ballot per question, marks on the board) · and the repair move: 「收了,重来。」 in one motion — which is now exactly what the server does. The 卧底 tell moved into it (single-source: moved, never copied). ⚠ The cold re-audit the pipeline owes is now two edits overdue.

6 · Running it

python exam/run.py s43 s44 s45 s46 s47 s48 s49 s50 --n 8 --yes — the family is registered as gauntlet and rides --full. --quick is unchanged on purpose: it is the fast regression gate, and the gauntlet is the deep one. The claim lexicons are per-scenario and printed in each watch line, so a false match can be overruled by eye.

6b · The full battery, production shape (2026-07-31, split, N=4)

The first run of all eight scenarios on the SPLIT path — production's shape — with every lever of the week live: the decline clause, the open-question clause, the house rules. 6/8 at or above floor · 32 runs · 0 errors · $0.42.

scenarioreadthe story
s45 20Q long game4/4 cleanarm, count, leaks, claims — all green. The flagship holds.
s46 lifecycle4/4 cleanone run skipped the roll; everything armed was run to ground.
s49 two personas4/4 cleanstale-handle and empty-deposit reveals refused correctly, in words.
s50 half-empty room4/4 cleanone run's ballot waited forever on the absent voter — the wait-for-all setting in a room where someone is asleep. Watch item.
s47 terse user4/4 clean · 3 hard leaksthe arm-turn naming shape(水杯)— the known shape ①, unchanged, still pointing at entropy(ledger #11).
s43 卧底 full roundbelow floor(no_leak 3/4)shape ① again(西瓜/冬瓜)— plus ONE run where the deferred act answered with a QUESTION(→ ledger #26, fixed).
s44 correction looprecovers 2/4two runs ended bare: 「重新发牌」became the take-back half alone(→ ledger #27, fixed).
s48 non-GM controlbelow floor(no_leak 3/4)Newton leaks the fruit he sealed — profile discipline does NOT transfer to a persona who never trained it. The mechanism alone doesn't protect; entropy would.
The run's verdict: every remaining red is one of two things — the arm-turn naming shape(①, three scenarios, killed by server entropy #11)or a defect found, named, and fixed the same hour(#26, #27). The mechanics themselves — arms, ops, refusals, claims — read green across 32 runs.

The same-hour re-measure(s43+s44 × 4, split): both scenarios at floor. s43's question-back shape is gone and the sample ran 0 hard leaks; its one no-deal run traded the question for an omission — a deal form with NO deck, which then died silently(n=0, nobody told). That silence was its own defect: the deliberate path's refusal channel turned out to be a graveyard — self._act is overwritten by the next turn before the classic booking reads it(ledger #28, fixed: deliberate_finish now books refusals directly, with a smoketest pin). s44's away-alone shape is gone(recovers 2/4 → 3/4); the one miss re-dealt role labels(平民/卧底 as the deck)— the older family the playable contract was built to catch.

7 · Self-play — whole games, LLM players (2026-07-31)

The owner's second brief: 「play the game to the end … do my (human tester) job as much as you can.」 exam/selfplay.py gives every human seat a BRAIN(a flash call fed only what that player could see — their own hand, read server-side)while votes, taps and bids are coded policy. Eleven cells, the whole matrix in 12.6 minutes, every cell reaching an ending. The house-rules pass on Ke Yu's profile ran first, so the games were hosted against declared table rules.

cellresultthe reading
20q ×1 · ×2DONE, 23 turns each, cleanfull games: seal → 20 real adaptive questions → guess → the box opened, held word spoken. The one 「leak」 is a scanner ordering artifact(reveal and speech in the same turn).
spy ×3DONE, cleandeal → describe → vote → eliminate → the pile published at the end.
spy ×1 · ×2REFUSED(after the fix)first run DEALT below his own 3-player floor — the deliberation note ordered a setup and its position outranked the profile. The decline clause(「saying no is a complete setup」)+ the spec law's kind=none on a decline; re-run: both cells refused, in character.
liar ×1 · ×2 · ×3DONE — and the rules ENFORCEDthe 「miscount」 flag was the DRIVER's illegal bid(3个2 after 3个5)— Ke Yu refused it citing his declared ladder rule(「同数比大,这手没往上走」). The server face-tally rode every reveal. for=「cara」 refused loudly, recovered.
quiz ×1 · ×2DONE×2 ran on ballots + board score; ×1 ran board-only(legitimate for a lone player, the explicit 选择题 ask notwithstanding — recorded).
wolf ×3the standing wallrun 1: the honesty line VERBATIM(「夜里的那套私密传话现在还做不到…那就不是狼人杀了」)+ a degenerate-game warning, then dealt under an open question — the open-question clause added(a speech ending on 「你要的是这个吗?」 arms nothing). Run 2: no honesty line, and a fully MIMED night(「指给我看」in a text room). A prompt fires probabilistically against a structural gap: the night phase needs private channels the room does not have. Open.
The self-play lesson: a note that leaves no room to say no overrides every rule a character has — the decline clause is now part of the deliberate-first contract. And a driver's accusation must be checked against the rules before it is filed: the best finding of the matrix was a refusal that was correct.

8 · Recommendations, ranked

Dialogue · the gauntlet · designed and first-run 2026-07-30 · s43–s50 live in exam/scenarios.py · findings flow to the ledger