Dialogue · The toolbox · The interrogation
← The play kits — the shelf this sits on

The interrogation — the kit that presses

Designed 2026-08-06 · BUILT 2026-08-07. Engine lib/interrogate.py · kit prototype/kits/interrogation.md · bench exam/interrogate_dryrun.py(gated in the smoketest)· the board, the length selector, the fold and 先到这 in the room. The first kit on the shelf that is not a game. You put an idea on the table; every seat attacks it from its own life, and keeps attacking until it is satisfied or you fold. You get three folds. What you leave with is a document of the holes.

why this one is worth building The other utility kits proposed alongside it (a decide-together vote, a blind estimate, a draw of lots) failed the same test: a person could run them with a coin, a spreadsheet or a group chat. This one needs the cast — five perspectives you do not have — and it needs the referee, because a budget of three evasions means nothing unless a machine is counting. And it leaves something behind.

1 · The evidence — one real run, reviewed

The design is not speculative. The owner ran an interrogation of his own — a 董事会 on an AI estate-planning idea, seven seats, three rounds, in a coding session with a wiki behind it — and the transcript was read end to end. What it settles:

what the run provedwhat it cost, and the fix
The entry gate did more work than any question. The host refused to convene until three things existed: the question, a position(「不许写"我不确定"」), and what it rested on.It also refused to start — fine for a private assistant, hostile in an app. Here the gate never blocks; whatever it can't fill becomes the first thread instead.
A standing rule separated the answers with content from the rhetoric — every answer had to carry first-hand experience, and violations were called by name(「硬规则违反」).That rule is one genre's currency. Generalized below: an answer must contain something that could be wrong.
An unanswered question with no cost gets asked forever. One seat's block was skipped whole in round 2; it came back in round 3 louder, and again in round 4 —「第三次」.This is the entire argument for the fold budget. An evasion has to spend something or it is free.
Some questions cannot be settled by talking at all.「去找 5 个真实用户聊」is an errand, not a question — and it was asked three times because the run had no way to mark it.The 待查 exit. Biggest single addition over the run.
Three structural seats carried the run, none of them a「master」: the one that only finds contradictions, the one standing in for whoever bears the consequence, and the closing audit of what everyone collectively failed to ask.The audit produced the best line in the transcript — the payer is the client, but the user is the heir, twenty years later. It only works from someone who wasn't asking.
Appetite is much higher than expected. Seven questions per round, answered in long form, three rounds deep.But by round 3 the questions had grown(a)(b)(c)sub-parts. One question per seat, no sub-parts — a question that needs a paragraph to ask is a bad question.

2 · Making it generic

That run is a product interrogation, which is why it seats an imaginary customer and demands lived experience. Neither generalizes. The move that makes the kit work on a philosophical claim, a life decision or a draft is small: stop demanding personal experience, start asking what the claim rests on — and let the answer set the terms.

in that run(product genre)the generic form
「你的亲身经历是什么」这想法哪来的 — lived it / saw the numbers / worked it out / it just feels right. You name it once, and it sets what counts as an answer for the whole run. Say「纯推理」and nobody asks for anecdotes; they attack the reasoning.
every answer carries first-hand experienceevery answer contains something that could be wrong — a case, a mechanism, a number, a commitment you'd accept. Restating the claim in fresh words is not an answer.
X(陈女士, the end customer)the one who has to live with it — seated only if there is one. A product has users; a life decision has the people around you; a purely conceptual claim may have nobody, and the seat stays empty.
「去找 5 个真人访谈」the ground seat — asks how we would check this, in whatever currency you declared
Plato(只找内部矛盾)unchanged — already generic, always seated
the closing blind-spot auditunchanged — always runs, at the end of a sitting

3 · Setup — the chair guesses, you correct

The first draft of this was an interview: four questions to answer before anything happened. Average users do not fill forms — they answer badly, or they leave. So nobody is asked anything. You talk; the chair says back what it heard, wrong parts included. A wrong guess is far easier to fix than a blank is to fill.

you ▸ 我想辞职去做自由职业 张小龙 明白了,我先说说我理解的,不对你就纠正我。 你想辞职去做自由职业,而且你觉得这是对的。 听着像是最近工作里憋出来的。 要让你打消这个念头,大概得是钱接不上。 这事儿最先砸到的是家里人。 对吗?哪句不对改哪句,或者直接说「开始」。

Four guesses, no questions, ten seconds to read, one line to fix. Note what is not there: no labels. No「你的主张:」/「依据:」. The structure lives in the machine; the screen shows a person talking. The moment a label appears it is a form again.

The gate never blocks. Anything the card can't fill becomes the first thread on the board, with a seat's name on it. You don't know what would make you drop it? That is a live question from turn one. Refusal becomes material instead of a locked door.

When what came in isn't debatable yet — four moves, and only four

what came inwhat the chair says
too wide — a topic, not a position「这题太大了,你到底想吵哪一块?」+ two or three concrete versions to pick from
a question, not a claim「你现在心里偏哪边?没有偏向我们没法追——有偏向就能追。」
two claims tangled together「这里面其实是两件事,先打哪个?」— the real run did exactly this(诊断 vs 做不做)and picking one is why it went anywhere
nothing could change your mind「听着你已经想明白了。那要不我们改成追问你为什么这么信?」— still a good run, different target

Two sharpening exchanges maximum, then it starts regardless. If the claim is still soft the interrogation will expose that faster than more sharpening will, and a kit that keeps asking you to rephrase before it does anything is a kit nobody fires twice.

Then the length selector — in effort, not turns

快 · 大概答十来次 中 · 半小时上下 长 · 慢慢来,可以分几次

Nobody thinks in turns. The budget bounds the sitting, not the run — see §7.

4 · The whole run

you talk — free text THE CHAIR reads back what it heard four guesses, no questions · no labels on screen you fix a line, or say「开始」 the gate never blocks — whatever it cannot fill becomes thread ① length:快 · 中 · 长 THE OPENING SWEEP every seat asks ONE — each its own bubble, so it can be replied to ① 张小龙 press ×2 他还想要:这三个月 实际接到几单 ② the contradiction seat press ×1 他还想要:你说想自由, 又说要稳定收入 ③ the ground seat waiting — you have not answered it, so it is silent a seat speaks ONLY when its own thread is answered — you set your own pace ✓ 它满意了 ✗ 我答不上(3 次) 待查 — 聊不出来 the sitting ends the board is quiet · the turns run out · or you say「先到这」 THE BLIND-SPOT AUDIT what did every one of us fail to ask? the doc → the wildcard pane next sitting:the board IS the save file Three folds for the whole run, not per sitting. A seat holds at most ONE live thread; a closed thread reopens only next sitting.
Setup is a conversation, not a form. The sweep is the only synchronized moment — after it, threads run at whatever pace you set by choosing which question to pick up.

5 · One thread — press until satisfied

The first draft capped a thread at two attempts. Dropped, on the owner's ruling: the ping-pong is the product — it is what actually sharpens the thinking. A seat presses until it is satisfied or you fold. Fewer questions, deeper ones.

the seat asks one question, no sub-parts you answer the ASKER rules its own thread — from its own profile, its own lens ✓ closed it presses — and MUST name what is still missing, in its own words back to you 待查 what it wants cannot be produced by talking ✗ 我答不上 spends one of three the two guards that make unlimited pressing safe ① a press names what would CLOSE the thread — so the goalposts cannot move, and you know what to aim at. ② if what is missing cannot be produced by talking, the thread becomes 待查 and the seat stops. That is the release valve.
The seat rules from its own profile —「did that satisfy me」has nothing but perspective, so this is never a generic judge call.

6 · The board — the checklist

The board holds state, not history. The chat already has the transcript; the doc holds the archive. One row per thread, rewritten in place — and the row carries the thing that makes the whole kit answerable: what that seat is still waiting for.

CLAIM 辞职去做自由职业,这是对的选择 v1 张小龙 追问 ×2 他还想要:你这三个月实际接到几单 Jobs ✓ 已解决 Plato 追问 ×1 他还想要:你说想自由,又说要稳定收入 Karpathy 待查 要真实报价——房间里问不出来 Dorsey ⏭ 跳过 这条进「破绽」,收在文档里 还能说「我答不上」: 2 / 3 v1 → v2 的改写都留在下面
Five states, each with its own glyph on the room’s board: ⏳ 等你 · 🔥 追问中 · ✅ 已解决 · ⏭ 跳过 · 🔎 待查. The open-ask column is what stops a seat moving its goalposts — and tells you exactly what to aim at.
An accepted reframe rewrites the card, and the old one stays underneath. The v1 → v2 trail is free, and on a philosophical run it is the output: what moved your claim, and what moved it.

6.1 · The row is an object — and that belongs to the board, not to this kit

A board's body was one markdown blob, and a blob cannot be touched: the room could paint it, but it could not colour one name, scroll to one line, or notice that one row had changed. The owner's report from room 6976 was the whole of it — when the status changes there is no visual reminder, and the user doesn't know he's making progress.

So the board — not the kit — takes an optional second view of the same body: the rows as data, beside the text. Fill them and you inherit all three things below. Leave them out and the board is the markdown path, unchanged, which is what every hand-written board still is. The interrogation is the first kit to fill them; any kit that writes rows gets the same for free.

诘问 2 ④ collapsed, the chip counts what moved — the one thing the width squeeze may not take 诘问 辞职去做自由职业,这是对的选择 Moritz ×2 — 他还想要:这三个月实际接到几单 Jobs — 已解决,不再占位 🔎 Plato — 要真实报价,房间里问不出来 还能回答 6 ③ → 那一句 ① 名字 = 它自己的座位色 one seat, one colour, everywhere ② 上次看之后变过的行 washes once, then goes quiet ③ 点一行 → 跳到那个座位 的原话,用房间自己的那一个跳转
The three capabilities the row buys, and the fourth the chip buys. Nothing here is new server data — the checklist line, the seat, the anchor and the move stamp all come off the same projection §6 already described.
what the row carrieswhat the room does with it
the seat — the persona's slugThe short name wears that seat's colour, the same one its bubbles and its avatar wear. Five bold-ink names down a card is a list you have to read; five coloured ones is a card you can glance at.
the line — the seat's own checklist wordsRendered whole, clamped at two lines. Whole and unclamped are not the same promise: at 375px these run to four wrapped lines each, and four rows of that is a 400px wall where a board is supposed to be a glance. A clamp is only honest when there is a way to the rest — and the row goes to the line, so there is.
the formatName: questionAnd nothing else. A seat reads the board on every call(「Karpathy:pressing — still wants:…」)and then writes its own checklist line, and some hand that shape straight back — so the row arrived carrying the name it already prints and the state its glyph already says. The world peels its own scaffolding back off(_peel_row): the seat's name, the five state words and the wants-word, in the room's lexicon and both shipped defaults, stacked as deep as they come. Only ever when a separator follows — otherwise「Plato argued…」would lose its subject. In code, not in the prompt: a clause is a request, this is a guarantee.
the anchor — the bubble id of that seat's own askTapping the row scrolls the chat to that line and flashes it. It reuses the room's one bubble jump — the same landing the quote-reply and the saved-note lines already use — rather than inventing a second one. A row with no anchor is a plain line, not a dead button.
the move stamp — a counter the world bumps on every state changeThe client diffs it against what it last painted, and the collapsed chip carries how many rows moved since you last opened it. A counter and never a clock — the world has no wall time a phone will agree with. The row itself does not move: v672 washed a changed row green for 1.8s and the owner's verdict was that a board should sit still(「no need of the flash effect or the update wash effect」). An open board says what changed by saying it — the glyphs are right there.

Measured, after the owner called the first cut bulky. A row had a 56px floor — a 40px minimum written against the CONTENT box with 8px of padding on top of it, so a one-line row was as tall as a three-line one. A touch minimum is a promise about the outer box. With that fixed and the line clamped, at 375×812: a row goes 76–97px → 49px(38 for a one-liner), the four-thread body 396 → 243px, the whole panel 476 → 322px.

Two rules the build turned on. The count is computed on the client, because「since you last opened it」is a fact about one window and the board's push is viewer-less by design — one payload for every screen. And the badge is deliberately exempt from the top bar's compression ladder: the rungs drop a chip's name and then its value, and a chip squeezed down to its picture is exactly the moment a count is the only news left on it.

The kit's own write also moved to after the turn speaks. It used to run first — one step before the bubbles existed — so a thread opened on this turn went out with a null anchor and its row could not be tapped until something else moved the board. The row's whole new trick is「take me to the line」, and the line does not exist until the turn has spoken.

7 · Sittings — the turn budget bounds the sitting, not the run

From the owner's own trial: it is a tiring process, the user will burn many organic tokens; there is only so much a person can process before he needs a rest. So the run is bounded by how much you have to answer, not by rounds or question count — and an interrogation you return to over a week beats one you exhaust in an hour.

no live-thread cap needed A seat speaks only when its own thread is answered. Answer only 张小龙 and only he presses. The user sets their own load by choosing which question to pick up — the only fixed cost is reading the opening sweep once.

8 · What each call reads — and how it stays cheap

Every ruling is perspectival:「did that satisfy me」is nothing but lens, so the seat's full profile rides every call, and(owner's ruling)so does the whole interrogation so far — a seat should be able to catch a contradiction that surfaced in someone else's thread.

That is affordable because the history is append-only: a seat's next call is a strict extension of its previous one, so everything it read last time is still there, unchanged, in the same order. That is exactly what a prefix cache wants.

NEVER CHANGES the seat's own profile — in full the claim card v1 + the kit's rules written into cache ONCE, at this seat's first call the sweep is the warm-up APPEND-ONLY every message since the interrogation started, in order — nothing rewritten a seat's next call is a strict extension of its last only the NEW tail is ever paid for — one answer wakes ONE seat, so it is written once, not once per seat RE-RENDERED EVERY CALL(small) the board as it stands · this seat's thread what it is being asked to do now a few hundred tokens ⚠ never edit anything above the tail — a reframe APPENDS v2, it does not rewrite v1; the board changes every call, so it lives at the very END, never in the middle of the history. One cold start per SITTING(1h prefix life)— another reason the sitting is the unit. Rough order for a normal sitting: ~3–4× cheaper than paying full price per call.
The profile stays in the head, not the tail. Sharing one prefix across seats would force the profile into the tail — where you pay full price for it on every call, which costs more than duplicating history writes per seat.
implementation note This is a different call shape from a normal room turn, where the persona rides the system prompt. Here the profile must be the cached head of a per-seat prefix. A seat reads every other thread but rules and presses only its own — an obligation in the call, and the board makes a violation obvious.

9 · The document — the actual product

Written continuously into the wildcard pane, not only at the end, so it survives a closed tab. Five sections:

  1. 破绽 · the holes — the threads that cost you a fold, verbatim. This is what you came for.
  2. How the claim moved — v1 → v2 → v3, and what forced each change. On a philosophical run this is the whole output.
  3. What survived — the answers that landed.
  4. 待查 · the errands — the questions whose answers are not in the room.
  5. 「什么会让你打消念头」, revisited — is that condition closer or further than when you started?

10 · Where it runs — the shelf, not the flow schema

It is a play kit and belongs on the shelf, but it is not a cartridge. The device's vocabulary is built for symmetric games — a living set, ballots over the living, a press for the whole room, an elimination, a round. This kit is one person against N seats, with per-seat perspectival rulings and threads that run at different speeds. What it can borrow: the sealed card, the board and its counters, the shelf, the capsule, the take-off. What it needs new: the thread runner, the per-seat ruling call, and a press only the pitcher can make.

deferred One pitcher per run. Other humans may watch and pile on, but the fold budget belongs to the pitcher alone — two people sharing three folds is a fight, not a kit. A human chair is interesting and also deferred; the roles are confusing enough on the first build.

11 · Settled, and still open

questionruling
fold budget3 for the whole run(owner)
who rules a threadthe asker rules its own, from its own profile(owner)
does the chair also askyes(owner)— the blind-spot audit loses some independence; accepted, in exchange for not spending a seat in small rooms
how long a seat may pressuntil satisfied or you fold(owner)— the ping-pong is the product
what bounds the runturns, not questions(owner)— and it bounds the sitting, not the run
threads per seat≤ 1 live, always(owner)
setup shapefree text + the chair's read-back, four guesses, no labels(owner)
the「cannot be settled by talking」statecalled 待查 — not「out of room」; most users do not think of a chat as a room(owner)
the kit's Chinese nameopen —「诘问」is the plain candidate
which seat chairs by defaultopen — proposal: user picks at mount, first seat pre-selected
seat flooropen — proposal: 3 seats minimum, 4–6 the sweet spot

12 · What the benches found

Two benches, and they caught different classes of defect. The dry run(scripted minds, no API)found the process bugs; the live runs(real calls, ~$0.004 a sweep)found the prompt bugs, which no stub can reach.

found bythe defectthe fix
dry runThe turn budget hung off thread endings, so a seat pressing forever never ended one — a sitting ran twelve answers past its length. A deep thread is the kit working, and it is also what burns the most turns.the budget is read after any ruling of any verdict, and after a fold
dry runWith folds spent AND the board quiet it offered another sweep instead of closing.folds outrank the budget; both outrank a quiet board
dry run(room)The sweep opens on the length tap, not on a message — so with publishing hung off the turn path alone, five questions were computed, stored, and never spoken. A full board and an empty chat.one publisher, four doors
live run 1Five seats asked the same question. All five question calls were armed at once, so every seat read an identical, empty history — four of them asked「why would anyone trust an AI over a lawyer」. A seat cannot avoid an angle it cannot see.the sweep serves one call at a time: it costs the parallelism and buys the only thing a sweep is for
live run 1Five English questions in a Chinese room. The prompts are English by convention and nothing ever said which language to answer in.an explicit language rule…
live run 3…which one seat still ignored, with an English-dominant profile sitting between the rule and the task.the rule rides the tail as well as the head — position is the dose, again
live run 1A press said the same thing twice: the reaction carried the re-ask, and the room published both.「`say` and `question` are DIFFERENT SENTENCES」
live run 2The chair called a perfectly interrogable pitch「too wide」and asked for a rewrite — the front door becoming a form, which is the thing §3 exists to prevent.clear is the DEFAULT; the bar for sharpening is high
paint checkThe fold button measured 22px tall — a mouse target on a control that spends one of three. And the board took 47% of a phone viewport.36px on a phone; the board starts collapsed there at 11%, keeping the claim and the fold counter
what the fourth live run looks like One pitch(an AI estate-planning tool for the $1M+ middle class), five seats, five genuinely different angles: legal compliance across states · show me one concrete case and where it got the details wrong · capturing unstated intent versus recombining legal phrases · a structure the training data never held — first principles, or a statistical guess? · and a moral one, in its own voice. Then one answer, one press that named exactly what was still missing. $0.004.
the risk that decides whether it works Not the machinery — politeness. Our personas are trained friendly and the panel's manners will sand every question down to「有没有考虑过…」. The mandatory blind-spot line under each question is the guard(a question that cannot name what it is probing is not asked), and it is exactly the kind of change a green smoke test cannot verify: measure round-one questions against a baseline before shipping.
Dialogue · the interrogation kit · designed 2026-08-06, built 2026-08-07 · grounded in one real run(董事会 / AI estate planning, 7 seats, 3 rounds)reviewed end to end · shelf: prototype/kits/ · not a flow cartridge — see §10