Designed 2026-08-06 · BUILT 2026-08-07. Engine lib/interrogate.py · kit prototype/kits/interrogation.md · bench exam/interrogate_dryrun.py(gated in the smoketest)· the board, the length selector, the fold and 先到这 in the room. The first kit on the shelf that is not a game. You put an idea on the table; every seat attacks it from its own life, and keeps attacking until it is satisfied or you fold. You get three folds. What you leave with is a document of the holes.
The design is not speculative. The owner ran an interrogation of his own — a 董事会 on an AI estate-planning idea, seven seats, three rounds, in a coding session with a wiki behind it — and the transcript was read end to end. What it settles:
| what the run proved | what it cost, and the fix |
|---|---|
| The entry gate did more work than any question. The host refused to convene until three things existed: the question, a position(「不许写"我不确定"」), and what it rested on. | It also refused to start — fine for a private assistant, hostile in an app. Here the gate never blocks; whatever it can't fill becomes the first thread instead. |
| A standing rule separated the answers with content from the rhetoric — every answer had to carry first-hand experience, and violations were called by name(「硬规则违反」). | That rule is one genre's currency. Generalized below: an answer must contain something that could be wrong. |
| An unanswered question with no cost gets asked forever. One seat's block was skipped whole in round 2; it came back in round 3 louder, and again in round 4 —「第三次」. | This is the entire argument for the fold budget. An evasion has to spend something or it is free. |
| Some questions cannot be settled by talking at all.「去找 5 个真实用户聊」is an errand, not a question — and it was asked three times because the run had no way to mark it. | The 待查 exit. Biggest single addition over the run. |
| Three structural seats carried the run, none of them a「master」: the one that only finds contradictions, the one standing in for whoever bears the consequence, and the closing audit of what everyone collectively failed to ask. | The audit produced the best line in the transcript — the payer is the client, but the user is the heir, twenty years later. It only works from someone who wasn't asking. |
| Appetite is much higher than expected. Seven questions per round, answered in long form, three rounds deep. | But by round 3 the questions had grown(a)(b)(c)sub-parts. One question per seat, no sub-parts — a question that needs a paragraph to ask is a bad question. |
That run is a product interrogation, which is why it seats an imaginary customer and demands lived experience. Neither generalizes. The move that makes the kit work on a philosophical claim, a life decision or a draft is small: stop demanding personal experience, start asking what the claim rests on — and let the answer set the terms.
| in that run(product genre) | the generic form |
|---|---|
| 「你的亲身经历是什么」 | 这想法哪来的 — lived it / saw the numbers / worked it out / it just feels right. You name it once, and it sets what counts as an answer for the whole run. Say「纯推理」and nobody asks for anecdotes; they attack the reasoning. |
| every answer carries first-hand experience | every answer contains something that could be wrong — a case, a mechanism, a number, a commitment you'd accept. Restating the claim in fresh words is not an answer. |
| X(陈女士, the end customer) | the one who has to live with it — seated only if there is one. A product has users; a life decision has the people around you; a purely conceptual claim may have nobody, and the seat stays empty. |
| 「去找 5 个真人访谈」 | the ground seat — asks how we would check this, in whatever currency you declared |
| Plato(只找内部矛盾) | unchanged — already generic, always seated |
| the closing blind-spot audit | unchanged — always runs, at the end of a sitting |
The first draft of this was an interview: four questions to answer before anything happened. Average users do not fill forms — they answer badly, or they leave. So nobody is asked anything. You talk; the chair says back what it heard, wrong parts included. A wrong guess is far easier to fix than a blank is to fill.
Four guesses, no questions, ten seconds to read, one line to fix. Note what is not there: no labels. No「你的主张:」/「依据:」. The structure lives in the machine; the screen shows a person talking. The moment a label appears it is a form again.
| what came in | what the chair says |
|---|---|
| too wide — a topic, not a position | 「这题太大了,你到底想吵哪一块?」+ two or three concrete versions to pick from |
| a question, not a claim | 「你现在心里偏哪边?没有偏向我们没法追——有偏向就能追。」 |
| two claims tangled together | 「这里面其实是两件事,先打哪个?」— the real run did exactly this(诊断 vs 做不做)and picking one is why it went anywhere |
| nothing could change your mind | 「听着你已经想明白了。那要不我们改成追问你为什么这么信?」— still a good run, different target |
Two sharpening exchanges maximum, then it starts regardless. If the claim is still soft the interrogation will expose that faster than more sharpening will, and a kit that keeps asking you to rephrase before it does anything is a kit nobody fires twice.
Nobody thinks in turns. The budget bounds the sitting, not the run — see §7.
The first draft capped a thread at two attempts. Dropped, on the owner's ruling: the ping-pong is the product — it is what actually sharpens the thinking. A seat presses until it is satisfied or you fold. Fewer questions, deeper ones.
The board holds state, not history. The chat already has the transcript; the doc holds the archive. One row per thread, rewritten in place — and the row carries the thing that makes the whole kit answerable: what that seat is still waiting for.
A board's body was one markdown blob, and a blob cannot be touched: the room could paint it, but it could not colour one name, scroll to one line, or notice that one row had changed. The owner's report from room 6976 was the whole of it — when the status changes there is no visual reminder, and the user doesn't know he's making progress.
So the board — not the kit — takes an optional second view of the same body: the rows as data, beside the text. Fill them and you inherit all three things below. Leave them out and the board is the markdown path, unchanged, which is what every hand-written board still is. The interrogation is the first kit to fill them; any kit that writes rows gets the same for free.
| what the row carries | what the room does with it |
|---|---|
| the seat — the persona's slug | The short name wears that seat's colour, the same one its bubbles and its avatar wear. Five bold-ink names down a card is a list you have to read; five coloured ones is a card you can glance at. |
| the line — the seat's own checklist words | Rendered whole, clamped at two lines. Whole and unclamped are not the same promise: at 375px these run to four wrapped lines each, and four rows of that is a 400px wall where a board is supposed to be a glance. A clamp is only honest when there is a way to the rest — and the row goes to the line, so there is. |
the format — Name: question | And nothing else. A seat reads the board on every call(「Karpathy:pressing — still wants:…」)and then writes its own checklist line, and some hand that shape straight back — so the row arrived carrying the name it already prints and the state its glyph already says. The world peels its own scaffolding back off(_peel_row): the seat's name, the five state words and the wants-word, in the room's lexicon and both shipped defaults, stacked as deep as they come. Only ever when a separator follows — otherwise「Plato argued…」would lose its subject. In code, not in the prompt: a clause is a request, this is a guarantee. |
| the anchor — the bubble id of that seat's own ask | Tapping the row scrolls the chat to that line and flashes it. It reuses the room's one bubble jump — the same landing the quote-reply and the saved-note lines already use — rather than inventing a second one. A row with no anchor is a plain line, not a dead button. |
| the move stamp — a counter the world bumps on every state change | The client diffs it against what it last painted, and the collapsed chip carries how many rows moved since you last opened it. A counter and never a clock — the world has no wall time a phone will agree with. The row itself does not move: v672 washed a changed row green for 1.8s and the owner's verdict was that a board should sit still(「no need of the flash effect or the update wash effect」). An open board says what changed by saying it — the glyphs are right there. |
Measured, after the owner called the first cut bulky. A row had a 56px floor — a 40px minimum written against the CONTENT box with 8px of padding on top of it, so a one-line row was as tall as a three-line one. A touch minimum is a promise about the outer box. With that fixed and the line clamped, at 375×812: a row goes 76–97px → 49px(38 for a one-liner), the four-thread body 396 → 243px, the whole panel 476 → 322px.
Two rules the build turned on. The count is computed on the client, because「since you last opened it」is a fact about one window and the board's push is viewer-less by design — one payload for every screen. And the badge is deliberately exempt from the top bar's compression ladder: the rungs drop a chip's name and then its value, and a chip squeezed down to its picture is exactly the moment a count is the only news left on it.
The kit's own write also moved to after the turn speaks. It used to run first — one step before the bubbles existed — so a thread opened on this turn went out with a null anchor and its row could not be tapped until something else moved the board. The row's whole new trick is「take me to the line」, and the line does not exist until the turn has spoken.
From the owner's own trial: it is a tiring process, the user will burn many organic tokens; there is only so much a person can process before he needs a rest. So the run is bounded by how much you have to answer, not by rounds or question count — and an interrogation you return to over a week beats one you exhaust in an hour.
Every ruling is perspectival:「did that satisfy me」is nothing but lens, so the seat's full profile rides every call, and(owner's ruling)so does the whole interrogation so far — a seat should be able to catch a contradiction that surfaced in someone else's thread.
That is affordable because the history is append-only: a seat's next call is a strict extension of its previous one, so everything it read last time is still there, unchanged, in the same order. That is exactly what a prefix cache wants.
Written continuously into the wildcard pane, not only at the end, so it survives a closed tab. Five sections:
It is a play kit and belongs on the shelf, but it is not a cartridge. The device's vocabulary is built for symmetric games — a living set, ballots over the living, a press for the whole room, an elimination, a round. This kit is one person against N seats, with per-seat perspectival rulings and threads that run at different speeds. What it can borrow: the sealed card, the board and its counters, the shelf, the capsule, the take-off. What it needs new: the thread runner, the per-seat ruling call, and a press only the pitcher can make.
| question | ruling |
|---|---|
| fold budget | 3 for the whole run(owner) |
| who rules a thread | the asker rules its own, from its own profile(owner) |
| does the chair also ask | yes(owner)— the blind-spot audit loses some independence; accepted, in exchange for not spending a seat in small rooms |
| how long a seat may press | until satisfied or you fold(owner)— the ping-pong is the product |
| what bounds the run | turns, not questions(owner)— and it bounds the sitting, not the run |
| threads per seat | ≤ 1 live, always(owner) |
| setup shape | free text + the chair's read-back, four guesses, no labels(owner) |
| the「cannot be settled by talking」state | called 待查 — not「out of room」; most users do not think of a chat as a room(owner) |
| the kit's Chinese name | open —「诘问」is the plain candidate |
| which seat chairs by default | open — proposal: user picks at mount, first seat pre-selected |
| seat floor | open — proposal: 3 seats minimum, 4–6 the sweet spot |
Two benches, and they caught different classes of defect. The dry run(scripted minds, no API)found the process bugs; the live runs(real calls, ~$0.004 a sweep)found the prompt bugs, which no stub can reach.
| found by | the defect | the fix |
|---|---|---|
| dry run | The turn budget hung off thread endings, so a seat pressing forever never ended one — a sitting ran twelve answers past its length. A deep thread is the kit working, and it is also what burns the most turns. | the budget is read after any ruling of any verdict, and after a fold |
| dry run | With folds spent AND the board quiet it offered another sweep instead of closing. | folds outrank the budget; both outrank a quiet board |
| dry run(room) | The sweep opens on the length tap, not on a message — so with publishing hung off the turn path alone, five questions were computed, stored, and never spoken. A full board and an empty chat. | one publisher, four doors |
| live run 1 | Five seats asked the same question. All five question calls were armed at once, so every seat read an identical, empty history — four of them asked「why would anyone trust an AI over a lawyer」. A seat cannot avoid an angle it cannot see. | the sweep serves one call at a time: it costs the parallelism and buys the only thing a sweep is for |
| live run 1 | Five English questions in a Chinese room. The prompts are English by convention and nothing ever said which language to answer in. | an explicit language rule… |
| live run 3 | …which one seat still ignored, with an English-dominant profile sitting between the rule and the task. | the rule rides the tail as well as the head — position is the dose, again |
| live run 1 | A press said the same thing twice: the reaction carried the re-ask, and the room published both. | 「`say` and `question` are DIFFERENT SENTENCES」 |
| live run 2 | The chair called a perfectly interrogable pitch「too wide」and asked for a rewrite — the front door becoming a form, which is the thing §3 exists to prevent. | clear is the DEFAULT; the bar for sharpening is high |
| paint check | The fold button measured 22px tall — a mouse target on a control that spends one of three. And the board took 47% of a phone viewport. | 36px on a phone; the board starts collapsed there at 11%, keeping the claim and the fold counter |