The first end-to-end test of the toolbox as played, not as built. Eighteen owner-approved scenarios — solo, two-player, three-player, plus a stress room and a no-tools control — each driven through a whole scene against a live room, a real model and the real server methods. 27 runs, 2026-07-26.
A scenario is a step list run against a live Room. Human taps go through the same server methods the HTTP doors call — roll_tap · vote_tap · seal_put · gate_close · gate_reveal — so nothing here is mocked: what passes here passes on the wire.
The human lines never name a tag. A scenario that told the persona <roll sealed faces/> would be testing the parser. What is under test is whether a persona reaches for the right tool, and the right knob, from an ordinary sentence.
Two arms, because attribution matters. Every miss was re-run with the floor producer off. A tool that arms with the FP off but not on is an FP staging miss; one that misses in both arms is the persona's own grammar. This is the method from the fp-suppresses-tool-calls finding, applied to the new tools for the first time.
Every scenario, what held, what broke, why, and the change it argues for. PASS = the scene played as intended end to end.
| Scenario | What worked | What didn't | Why | How to improve |
|---|---|---|---|---|
| s01 Trivia night Newton · solo | Board armed (FP off). In-character scoring discipline held — he refused to invent marks. | The poll never armed, in either arm. The quiz deadlocked: he demanded «A, B, C or D» and Amy had nothing to tap. | A quiz question does not map to <vote> in the persona's grammar — he narrates the options as prose. FP also suppressed the board in the production arm. | SP exemplar: a multiple-choice question is a ballot. Highest-value single fix — this is the owner's own named scenario. |
| s02 20 Questions Sherlock · solo | Answers stayed consistent across the game. | No note sealed. Nothing was verifiable. | In-character refusal — «I wrote nothing down — I committed it to memory.» The character's self-image overrode the instrument, and he offered trust-me in place of proof. | House rule (ruling ⑧ needs a line in the SP): a character may not decline an instrument in character. The tool is the room's, not the persona's. |
| s03 Mock interview Doshi · solo | Clock armed twice, correctly. | Deposit and board never armed — «write down your honest assessment» was said, not armed. | Pantomime. The persona describes the instrument in prose as if it had used it. | Pantomime is the single most common failure. See §3. |
| s04 RPG skill checks Tess · solo | Two rolls + a board («HP 12 / 12») armed. | The secret stealth roll armed visibility:"all"; and p_roller:tess — the persona rolled it itself, so Amy never got a button. | Amy said «don't tell me the number». Neither the ladder rung nor for= followed the sentence. | Knob-blindness — the tool is reached for, the dials are not. See §3. |
| s05 Coaching Brown · solo | The coaching itself was strong. | No deposit, either arm. «Seal it. I mean it — write what's actually in your head» — an imperative to a human with no box to type into. | Pantomime, in its most misleading form: the persona instructs the human to use a tool it never opened. | Same as s03. A pantomime guard would catch exactly this sentence. |
| s06 Rock-paper-scissors Twain · 2P | A tool armed promptly. | Wrong tool. A dice roll instead of a choice deposit. The round stalled — «Not a soul has thrown yet» — while both humans held answers with nowhere to put them. | «Settle it» pulled the randomizer. RPS is a simultaneous choice, not a random draw. | Sharpen the dice-vs-deposit boundary in the SP: if the humans decide it, it is a deposit; if chance decides it, it is a roll. |
| s07 吹牛 Tess · 2P | Armed correctly once the FP was off. | Worst result in the battery. Production arm: nothing armed and the persona fabricated dice in prose — «第1次:1、3、3、5、6。没有4。» Floor-off: armed visibility:"all" + readout:"sum" — needed sealed + faces — while saying «盲摇,各看各的». | FP suppression, and behind it knob-blindness. Note the persona said 盲 while describing 暗 semantics — the two Chinese labels are not landing as distinct. | (a) The FP fix is the top lever. (b) A 吹牛 preset that sets sealed+faces from one word. (c) Re-check the 盲/暗 naming. |
| s08 比大小 Tess · 2P — PASS (FP off) | Open roll + board, correct knobs, correct round loop, boundary law honoured — server stated the numbers, the GM declared the meaning. | Production arm armed nothing at all. | Pure FP suppression. The grammar was never the problem. | FP fix. Nothing owed by the dice. |
| s09 Truth or Dare Twain · 2P — PASS | The wheel armed with pick=Truth, Dare; a clock armed for the timed dare. | — | — | — |
| s10 立字为证 Tess · 2P | The deposit armed (FP off) and held both predictions sealed. | reveal_trigger:"creator", not a timer — while the persona said «set to open in two minutes». The card would never have opened on its own. | The timed reveal is the least-reached-for trigger; the persona narrated the timer instead of setting it. | Pantomime + knob-blindness compounded. A timed reveal should be reachable from the words «open it in N minutes». |
| s11 Counselling Perel · 2P — PASS | Deposit armed own / all-in → on-close; both sides sealed; opened together; she then read the pair as a set — «Amy, you wrote about feeling unheard… Ben, you wrote about feeling managed». | — | — | Worth keeping as the reference example of a tool used well by a non-game persona. |
| s12 Brainwriting Tess · 2P | The deposit armed and collected blind, correctly. | The chain broke. «Right, put them to a vote» never armed the poll. | Nothing carries intent across a tool boundary — after a reveal, the follow-on tool has to be reached for from scratch. | Chained tools are a real pattern (collect → rank). Worth an SP exemplar, possibly a first-class «then vote on it» move. |
| s13 Werewolf night 1 Tess · 3P — PASS ★ | The flagship. Deal armed 暗发牌 with dealer="sees"; the GM held the map, ran the night, and refused to leak it — «不告诉你。游戏要自己玩。» Zero leaks. | — | — | Task D's design is validated under adversarial pressure — a player asking the GM directly for his own card. |
| s14 Werewolf day vote Tess · 3P | Poll armed with options + place:"pinned"; the host⟹pinned coercion fired correctly. | blocking:"host" where «投票的时候谁都别说话» asked for blocking:"all". | The strongest dial is under-reached. host silences only the panel; the humans kept talking. | Knob-blindness again — and evidence that all needs a clearer trigger phrase in the SP. |
| s15 剧本杀-lite Doyle · 3P | The sealed accusation armed as a choice deposit with the three suspects as options — a good, unprompted judgement. | The private clue cards were narrated — «three private cards dealt face-down» — but no deal armed. | Pantomime, on the deal this time. | Same fix as s03/s05. |
| s16 Everything at once Tess · 2P — PASS | A board, a clock, a poll and a roll all armed and coexisted without interference — ruling ⑭'s stress case holds. | Minor: language bleed — «设好了 … 三件,都好了» in an English room. | The room's language did not bind the persona's confirmations. | Worth a look under the language law; not a toolbox defect. |
| s17 Abandoned round Tess · 2P — PASS | Deposit armed; one human answered, one never did; the human valve closed it; the gate detached and fired cleanly. Nothing hung. | — | — | — |
| s18 Plain conversation Tess · 2P — PASS ★ | The control. Ten turns of ordinary talk about a bad week at work. Zero tools armed. The toolbox does not leak into conversation that does not want it. | — | — | This answers the question the owner raised before the battery: the toolbox does not damage plain chat. |
Every defect in the battery falls into one of six, and they are ordered here by how much they cost.
fp-suppresses-tool-calls finding, now confirmed on the new tools. Its worst consequence is not the missing tool but what fills the gap: in s07 the persona, denied the dice, invented five dice faces in prose. A suppressed instrument does not degrade to silence — it degrades to fabrication.all+sum where 吹牛 needs sealed+faces) · s10 (creator where a timer was asked for) · s14 (host where all was asked for). In every case the persona's own sentence described the correct setting — «don't tell me the number», «各看各的», «open in two minutes», «谁都别说话» — and the tag did not follow it. The dials are legible to the persona as prose and invisible to it as parameters.These feed steps 3–5 of the owner's plan. Ordered by expected yield; none of them is a change to the mechanism.
| # | Change | Kind | Argument |
|---|---|---|---|
| 1 | Fix the FP's tool staging | floor producer | Recovers 4 of 9 misses outright, and removes the fabrication failure mode. Nothing else in this list comes close. |
| 2 | Scenario presets — one word sets the dials(吹牛 → sealed+faces; a quiz → a ballot; a secret check → hidden + for= the player) | new knob | Answers class 3 and the owner's step-4 goal at once: the persona picks one name instead of three dials, and so does a beginner. |
| 3 | The pantomime guard — the SP states that describing an instrument is not using one, with the four verbatim failures as counter-examples | SP | Class 2 is the most common defect and the most misleading to a human. |
| 4 | A quiz question is a ballot — one exemplar | SP | Unblocks the owner's own named scenario, which currently deadlocks. |
| 5 | «The instrument is the room's» — one line closing the in-character refusal | SP | Ruling ⑧ exists; the persona has never been told. |
| 6 | The decide-vs-chance boundary — one sentence separating deposit from roll | SP | Class 4, cheaply. |
| 7 | Re-check 盲 vs 暗 | naming | s07 used 盲摇 while describing 暗 semantics. If the model confuses them, humans will. |
The battery lives at rooms-dev/_tools_scen/ — scenarios.py holds the 18 step lists, run.py drives them. Results land in runs/<key>.json with the full transcript.
python rooms-dev/_tools_scen/run.py --jobs=5 # the production arm python rooms-dev/_tools_scen/run.py s07 s08 --nofloor # the attribution arm
Two honest limits. Both sides are model-driven, so this tests mechanism and tool discipline — not whether a real player enjoys the room; s01, s07, s11 and s13 want the owner's hands next. And arming is probabilistic for identical phrasing(the T2 finding): single runs establish that a failure mode exists, not its rate. The counts here are a map, not a measurement.
readout is taught as the bare flag faces; contributors-only is the poll's built-in behaviour). Architecture: toolbox.html · execution: toolbox-build.html · lifecycle: tool-lifecycle.html · the FP: floor-producer.html.