Dialogue · The toolbox · The robustness track
← The gauntlet — the battery that finds these

The robustness track — levers toward the 90% table

The owner's brief, 2026-07-31: 「think: what other levers we can pull to make the tool use more robust so that a new user gets a positive surprise 90% of the times. One possible lever: let the persona decide if he wants to act/speak…」 This page is the working plan and the scoreboard. The measure: the gauntlet's run-clean rate on the production shape — 26/32 ≈ 81% when the track opened.

1 · The levers — built, measured, live

leverwhat it doesmeasureddial
The house's draw
(server entropy, #11)
draw="words" — the house fills a deposit of your own; draw="pair" + a ?1/?1/?2 deck — the house draws the confusable pair at arm. The secret is chosen AFTER the speech published, by physics: nobody can leak what nobody chose, and no host converges on 斑马. The binding note and the dealer's map deliver the words; the history holds only the pool name.hard leaks on the three leak scenarios 8 → 1 · the non-GM control(Newton)below-floor → 4/4 clean · the act call reached for it UNPROMPTED in self-playlib/wordpools.py — always on
The second attemptan act call whose EVERY form refused re-runs ONCE with the refusal reasons in front of it(「fix exactly what the refusal names — or kind=none」); a missing on a deliberate turn retries as a decision. Never retried: a clean none, a partial arm, the gate.converts the s43-r2 shape(a deckless deal → a whole round over a bare table)same-turn; both attempts billed; 3 smoke pinsalways on
The silence door
(the owner's lever, half 1)
on the floor producer's [MODE] hold the panel call is SKIPPED — structural, free silence; the humans' lines ride the next drain. Hard guards in code: ≥2 humans · never over a PROP verdict · never on a dispatch. Every hold books a coordinator event.live probe: two humans bantering → hold, intent 「let the humans finish their thread」; the direct ask answered in charactersilence_door · ON
The coda
(the owner's lever, half 2)
the third fire of a GM's rhythm: explain → the cards land → ONE short line hands the first move(「Amy 先来」). Only after a deliberate setup whose cards wait on players; skipped if anyone spoke meanwhile; its own act call pinned off.live probe:「Cards are down… Amy, you go first.」— the full rhythm end to end; the exam's _turn now mirrors itcoda · ON
The pattern the four share: move a decision OFF the prompt and INTO the mechanism. The draw removes the choice, the retry removes the wait, the door removes the obligation to speak, the coda removes the dead air — and every one of them books its decision as an event, so the exam can score it.

2 · Day 2 — the owner's four moves (2026-07-31 evening)

The owner's brief: 「1) let the persona act/speak as he sees fit — the CC way … 2) study and reorg the FP/Prop/Act/Speech topology … 3) examples per tool … 4) play kits, same idea as skills.」 Built kits-first by the owner's priority call.

movewhat shippedmeasured
4 · Play kits
(the kit shelf — its own page, with the five-beat flow diagram)
a game's researched RULEBOOK as a file(prototype/kits/, six kits: 卧底 · 大话骰 · 二十问 · 问答之夜 · 两真一假 · 真心话大冒险): a HOST BRIEF the panel reads(appended once, cache-safe)+ a TABLE SPEC the act call reads while active. The shelf is a ninth tile on the tool sheet; a message naming a game gets a one-tap OFFER chip(server substring, no model call). Variants are DIALS the host must declare — canonical rules researched, two profile errors corrected(1s wild by default in 大话骰 · spy-wins-at-three in 卧底).the control experiment: isaac-newton(zero GM training)hosted 卧底 BY THE BOOK with the kit — drawn pair, real ballot, whole pile turned at the end; the same persona without it pantomimed reveals, contradicted its own state and hand-wrote an unusable ballot
1 · The loop
(think→act→speak→verify)
the missing tail: deliberate_verify() — after phase 2's cards land, ONE look back(table vs speech, the landed handles in front of it); expected verdict none, exceptional verdict ONE fix form through the same door, form-level capped, never re-arming a live card. The full turn is now prop(think)→ speech(explain)→ act(set)→ verify(check)→ coda(hand the move)— five beats, all behind the first bubble.stubbed pins green(stash · none-verdict · consume-once · one-fix cap); live run measured on the selfplay cells — the none verdict is the norm, the event ledger books every look
3 · Examplesfolded into the kits(per-game worked settings ARE the examples, loaded only when relevant)+ the act schema's existing draw/deck teachings — a standing examples block on every act call was rejected: the dilution law(s07: folding clauses into lists = 8/8→3/8)prices prompt weight paid on every turn, and a kit pays only when its game is on the table.by design
2 · Topologystudied, kept: prop(cheap think)→ FP ∥ act → speech → drain → phase 2 → verify → coda. The one merge considered — prop+FP as one call — was rejected: the prop's fail-open property is load-bearing(a dead judge must not silence staging)and the two run concurrently anyway, so the merge buys ~1s and spends a safety property.decision recorded
The loop's cost honesty: a game-setup turn now runs prop ~1s → speech ~3s(publishes)→ act ~3s + verify ~2s + coda ~3s behind the bubble — worst case ~12s of wall time of which the user WAITS only ~4s before reading; the owner's ~10s-masked budget holds. Ops turns are untouched.

3 · The constitution(ratified by the owner, 2026-07-31)

Code is the world; the LLM is the mind. The world records everything faithfully and is selectively visible to each character — your own hand, the dealer's map, your private pad, never omniscience. It executes whatever act a mind commits, enforces physics(arithmetic, caps, secrecy walls — a refusal is a wall, not an opinion), and supplies what fairness requires the world to own: shuffles, dice, the drawn secret. The minds are layered — the persona is the actor(judges in-fiction, owns its plan, may depart from it); the floor producer and the prop master are the coordinators(stage the production, route the calls)— and they interact with the world spontaneously. Code never breaks a tie, never picks a next move, never decides that anyone failed.
the review criterion Every injected line the world writes into a mind's context must be a fact or a standing stage direction — never a per-moment verdict.「The counters last moved: never」passes.「Step ④ due: close the ballot」does not. This is the test each future lever is reviewed against — the auto-advancing rundown pointer was rejected under it before it was built. The constitution's full consequence — every world change waking the minds, wake→do as the one loop — is designed on the world-event room.

4 · The ownership experiment(run under the constitution, same day)

The question: can the persona OWN its game plan — write its own run sheet, check its own boxes — or does the world need a pointer? Replayed the three mary-test live failures on the SAME non-GM personas(Jensen's triple seal · Kim's frozen tally · Stan's 2T1F role flip), three rounds:

roundwhat changedresult
1 — the sheet taught in prose alonethe kits' HOST BRIEF told the host to write a run sheet into its pad; the mirror surfaced settles + chip stalenessall three failures REPRODUCED, and run-sheet=no everywhere — the speech has no pad grammar under the split, and the hands(the act call)were never told. A teaching that reaches nobody who can write is not a teaching
2 — identity keep-rules + the sheet reaches the handsa re-declared chip LABEL keeps its address(a board rewrite can no longer reset a count)· one DRAWN envelope per author(a new draw replaces the old, close-with-release)· the kits' TABLE SPEC gains the pad stepnudge CLOSED(one envelope)· tally CLOSED(4 ticks, the sheet landed and updated)· flip still open — but the pad showed a CORRECT, self-updated plan: the failure had moved below the plan
3 — one kit branch + one mirror fact2T1F's spec gains the host-as-storyteller branch(who="me"+my_answer — the hands had copied the literal who="uN")· the mirror states every envelope's contents-state(rows in · answered by)flip CLOSED: the host sealed his own lie at arm, revealed before the verdict, the sheet tracked the whole round. 3/3
The experiment's verdict on plan→do: the persona CAN own the plan at this model tier — Stan's pad rewrote itself on the role flip and tracked the round's state unprompted — provided the world does its two constitutional jobs: keep identity straight(labels keep addresses; one drawn secret at a time)and put the record beside the plan(the mirror). No pointer, no per-step commands, no code deciding a game. The failures were never intelligence failures; they were absent-context and broken-identity failures wearing intelligence's clothes.

5 · Candidate levers — not built, ranked

candidatethe ideawhy it waits
verify-then-speak on opson OPERATE turns the speech runs parallel to the act and cannot know what landed — the C8 assembly fix narrows it, but a speech that READS the ledger before claiming would close claims_true structurally.costs serializing the two calls(~3s latency)the parallel design bought on purpose; the claims seam is currently 4/4 green
the self-wake clocka table armed and nobody moving for N seconds → the host nudges once(the clock instrument already exists as machinery).needs idle detection the worker doesn't have; the coda covers the commonest case(right after setup)
multi-fire beyond the codaa persona loops speak/act freely inside one turn until done — the general form of the owner's lever.unbounded cost and a new turn grammar; the deliberate turn + coda already give three fires on the beat that needs them
an EN word pool per room languagethe pools carry zh + en; the room picks by its own language. A room in a third language falls back to zh.third-language rooms are rare today; the fallback is safe

6 · The scoreboard

The track's measure is the gauntlet, production shape(split · deliberate ON · N=4). Opening state 2026-07-31 morning: 6/8 scenarios at floor, 26/32 runs clean, 9 hard leaks. After the day's levers: the three leak scenarios re-ran at 3/3 at floor, hard leaks 8 → 1, and the two mechanism defects the run surfaced(#26 the question into the void · #27 the half-move · #28 the graveyard channel)are closed with pins. The next full-battery run prices the 90% claim honestly — run it after the next change, not before.

Dialogue · the robustness track · opened 2026-07-31 on the owner's brief · the ledger: tool-defects.html · the battery: tool-gauntlet.html