Dialogue · Design notes · Floor producer
← Design notes

Floor producer — design & quality study

The build spec for roadmap §1, off the back of the production diagnosis. The diagnosis proved where conversation quality breaks; this is the machine that fixes it. The producer is one function — f — that reads a rich signal vector and emits the staging. Status: v0 deterministic baseline built + A/B-measured; v1 the trusted smart f (v1p / v1f) shipped to production 2026-07-03 (v1p default for new rooms) and quality-studied; v2 per-user memory + v3 a trained policy sketched. Design rewritten 2026-06-26; quality study run 2026-06-29 (Part 9).

⚠ This page was rewritten 2026-06-26. The earlier mode-based design — a 2×2 mode space, a mode→staging table, the function stack — is superseded by the signals → f → staging model here. The old version is kept for reference: floor-producer (mode-based · outdated) →.
The room is voiced by one mind, and that mind always has something for everyone to say, at length, in agreement. It can't self-correct — the writer and the staging-decider are the same context. The floor producer is a separate per-turn agent that reads the room and sets the staging — who speaks, how long, what to focus on, how hard to push — never the content. Restraint injected from outside, which the writer must accommodate rather than remember to impose on itself.
Part 1The problem

What the diagnosis proved

We pulled all 80 production rooms and scored the 39 real conversations (full report). Four findings:

The load-bearing conclusion: three of the four are staging, not voice. Voice and character fidelity are good (users said so). The single writer just stages the room badly — and can't fix itself. That is exactly the gap an exogenous producer fills.

Part 2The idea — a trusted smart f

The producer's whole job is a single function: a high-dimensional signal vector in — your engagement, your explicit asks, how invested you are, the substance, the room's structure (solo or panel), the phase — and the staging out. We trust a smart model with all the signals and let it decide. No taxonomy in the path.

No boxes — a manifold

The earlier design tried to sort each conversation into one of seven named "modes." That's the wrong primary abstraction, and one real production room proved it (Carrie, below). The deepest way to say why:

Conversations don't live in 7 boxes; they live on a continuous, high-dimensional manifold — and staging is a smooth function of a rich signal vector.
the real shape — a continuous manifold Explore Spar Banter Learn Build Decide Vent Carrie dots = real chats, filling the space · dashed = the coarse 7-box overlay · Carrie sits on an edge
Conversations don't live in 7 boxes; they live on a continuous, high-dimensional manifold (flattened to 2D here). Real chats — the dots — fill the space between the lines; the dashed grid is the coarse human overlay. Carrie sits on a boundary, so the grid mis-files her as Vent — the engine doesn't use the grid, it reads her signals directly.

The engine

So the engine is one call: signals in, staging out. Nothing is forced through a discrete bottleneck on the way.

one function · no mode in the path SIGNAL VECTOR high-dim · open · engagement quality · explicit asks · dials · investment · substance · structure (solo/panel) · phase · history · brief f trusted smart producer LLM now · learned later THE STAGING who · order · silence · length · push → a free-text directive to the personas → a thin structured spine to the orchestrator firewall: staging, never content mode · intent = readouts, generated LAST → debug only
The engine: a rich signal vector in, the staging out — a free-text directive (for the personas) plus a thin structured spine (for the orchestrator). The firewall holds: f emits staging, never content. Mode and intent are byproducts — emitted last, routed to a debug surface, never into the engine. Describe with labels; decide with signals.

Mode & intent are readouts — generated last

f still names what it did — a mode ("this reads as deep self-work") and an intent ("she wants a rigorous assessment"). But these are labels read off the decision, for your eyes and the logs — not steps in it. And the order matters: in one autoregressive call, what's generated earlier conditions what comes later, so generation order is the causal graph. Put the label first and the staging anchors on it (the old gate, smuggled back in). Put it last and it can only describe.

one f call · generation order = the causal order ↓ ① signal-grounded reasoning scratchpad — "what I see → what fits" ② STAGING — directive + spine → persona LLM + orchestrator the deliverable generated LAST — downstream of the decision, so they can't bend it: ③ mode ④ intent debug bubble ONLY ✗ never the persona ✗ never the next f call
The same model call, top to bottom = the causal order. Reasoning and staging come first (staging is the only thing the personas see); mode and intent come last and go only to the debug surface — never to the persona, never seeded into the next turn's read (continuity comes from re-reading the history, not from consuming a possibly-wrong past label).

The firewall & the confer

The panel already deliberates — the backstage <confer> is the one mind working out what to say. The producer doesn't fight that; it sits upstream. The line that keeps it a router, not a co-author:

the firewallEvery call takes public signals and outputs staging — never content. f may read "this is grief," but that only sets the staging (length, who, how hard); it never tells a host what to say. Inferring intent to drive staging is the job; injecting it as content ("argue bootstrapping is better") is puppeteering — and it's why the output is mostly a directive about how to play it, not the lines themselves. On who / length / silence the producer wins (being brief or sitting out never needs fabrication); on contention it opens the space for a split, and the confer decides if anyone walks through.
the firewall — a checkable ruleTwo hand-picked examples aren't a boundary. The test the v1 prompt and a cheap post-gen lint can both apply: the directive may name a staging DIMENSION + a MANNER (who · how long · how hard · register · "hold your ground" · "surface the split") but never a TOPIC, a POSITION, or a LINE. It covers every lever the directive carries — challenge-manner and punch-craft included, not just who/length/silence.
safety floor — asymmetricOne rule does not wash out in the signal vector: on a detected or ambiguous read of acute distress (grief, crisis, self-harm), the staging is clamped soft regardless of the other signals — never raise challenge, never punch, presence over analysis. Asymmetric on purpose: a false-soften costs little; a false-grill on a hurting user is real harm. The same asymmetry governs closure: clamp length on a suspected wind-down but never cut warmth/availability, and treat any next substantive turn as an immediate re-open — a short "嗯" is not "they're done." ⚠ This is a v1 discipline the current v0 lacks — v0 seeds punch by turn-number with no affect gate, so a cutting line can land on a tender turn. That gap is live; the "never a tender one" claim elsewhere is v1-pending, not a shipped property.
why trust one callIf a single mind can't self-correct, why is f — also one call — trustworthy? Not infallibility. It's structurally better placed: a narrow job (staging, not content), cheap to retry, and graded by outcomes it doesn't author (the referee). And it never overrides the writer's ground — the directive sets staging and may ask for a register, but the persona's authored voice and red-lines win over a conflicting manner cue. f opens space; the persona fills it as itself.
Part 3In — the signal vector

Since there's no taxonomy to maintain, the real engineering becomes plumbing the signals into f. It can only be smart about what it's shown. The vector:

What f reads

signalwhat it carriescost
engagement qualityis the latest message a sharp probe or a flat "ok"? going deeper or cooling off?model — the loudest cue
explicit asks · dials"@X · shorter · go deeper · make a doc · don't comfort me" — priority over inferred intent. (The Debate/Personality/Challenge dials are not a separate channel — just a UI for emitting these asks.)cheap — parsed
investmenthow much each human is putting in — a 543-字 opener vs "hi"cheap
rostereach seated persona's card — slug · role · tagline · tags. Required for who-speaks-by-relevance (which host can actually carry this turn)free — room state
structuresolo vs panel · how many personas · how many humans (see Multiple humans)free — room state
furniture in playwhat is physically on the room's table this turn — board up · clock running (m:ss left) · ballot open (n of m voted) · a die in someone's hand. Added 2026-07-22: without it f stages blind to the room's own state (see the room has furniture)free — room state
substancethe real ask, and — the contention judgment — is there a genuine split?model
phaseopening · exploring · converging · wrappingmodel-light
history · briefthe transcript tail (last 8 turns, 160 chars each) — engagement + phase — plus the GUEST FEEDBACK block (built 2026-07-03): the humans' recent 👍/👎 ratings + each user's latest survey, rendered as standing asks — panel-level, admin On/Off + an N-turn lifetime (default 8; console → Settings → Feedback). Still not plumbed, by design: the §5 situation box and mechanical last-turn compliance — the feedback loop replaced the compliance read (see Part 6).free — room state

"Cheap" vs "model" is honest here: explicit-asks/investment/roster/structure/history are parse-or-free; engagement, substance, phase are model judgments — the headline sensor (engagement quality) is one of them, not free.

cold start isn't specialThere's no signal-less moment. Turn 1 already has the title, the brief, the opener, the cast, the language — often the richest signal it'll ever get. So "cold start" is just turn 1 with whatever signals exist; f reads them → staging, same as every turn. No mode-prior fallback needed. (The one genuinely turn-1-specific call is who opens a panel with no transcript — decided by topic-to-persona affinity, the brief/title × the roster, not recency.)
signal caveatsFreshness/scope: standing signals (brief, dials, persisted prefs) persist until changed; turn-scoped ones (a "shorter", an @-mention, a one-off affect spike) decay — a challenge cue from 3 turns ago is not a standing order. Two-source reads: "short-sharp-escalating" is intellectual heat OR emotional distress — same surface, opposite response; the read must route distress to the safety floor, never to "push harder." Language: the semantic reads (engagement, contention, affect) are language-sensitive in a way the length unit isn't — validate the judge axes per language. Adversarial: the solo-deep ceiling is engagement-gated, so a user could fish for walls by faking intensity — require sustained corroboration (multi-turn investment, an explicit ask), not one spiky message.

Entrain to engagement, not keystrokes

The one sensor that earns its keep first. The naive rule — "match the user's length" (entrainment) — is a trap: it would clamp a deeply-engaged user who fires short, sharp probes. The fix is to read engagement quality (substance + trajectory), not raw message length:

the messagenaive length readengagement-quality read
挡住的是我的焦虑吧 (9字)short → reply short ✗sharp probe, escalating → deep → stay long
嗯,谢谢 (would-be)short → reply shortflat, closing → clamp / wind down

So intent owns the ceiling, not the keystroke count: a short-but-sharp probe in a deep session keeps the reply long. That's the concrete mechanism behind "let the substantive turns breathe." (Caution: short-sharp-escalating is also the signature of distress — same surface, opposite correct response. The read must route distress to the safety floor, not "push harder.")

Carrie — the case that forced this

A real production room (梁宁 · 自我探索, solo): the user fires 8–98字 probes and gets 900–1886字 of rigorous analysis back, for 18 turns — escalating depth, explicitly asking "你不要安慰我,以面试官的角度客观评价," pushing back, never "太长." She is clearly satisfied; the long replies are the value.

She's emotional·converge — Vent's coordinate on the old grid — but wants the opposite staging (generous + challenge, not minimal + warmth). The grid would mis-serve her. And within her single session, the staging swung wildly by the turn's ask:

her turn (the signal)the staging f should produce
"陪我探索" — vulnerable opener~999字, warm
"目前主要在接商单" — flat update~336字
"你不要安慰我,以面试官的角度客观评价"~1623字 · challenge HIGH
"乔布斯、张小龙是什么样的" — teaching ask~1886字, exposition
what Carrie provesWhat set her staging was her explicit asks, her engagement (short-but-sharp, escalating), and the room being solonone of them a position on a grid. One "mode," four very different signatures, every one set by the turn's signals. Read the signals, set the staging; the mode is just the name we file it under. (She also needs a length tier above "generous" — a solo, deeply-invested user legitimately runs 500–2000字; "wall" is contextual, not absolute.)

Multiple humans in one room

The product's defining case — but the way to serve it is not a concurrent-fairness optimization haunting every turn. The room is slow-paced and turn-batched: human messages queue and drain together into one turn, and that turn usually holds exactly one human. That collapses the problem — multi-human is a within-turn question, and across turns it's just a sequence of mostly-single-addressee turns, each handled fresh by reading who actually spoke.

the batched turn is the atomic unit — multi-human is a within-turn question, and the turn usually holds one human TURN batched · drained together = f's unit 1 human msg the COMMON case single addressee → stage as if solo N human msgs RARE · simultaneous allocate · serve each the call-list already does this humans NOT in this turn's batch — not a fairness debt: present · silent this turn OBSERVER = a use case serve by leaving be humans → each other NO-OP · panel sits out the load-bearing lever silence = satisfied-by-default NOT worst-served / abandoned an invite is optional, never a duty
The batched turn is the atomic unit. The common turn holds one human → stage as if solo; the rare simultaneous turn is allocated across askers by the Part 4 call-list. Humans outside the batch aren't a fairness debt — a silent human is observing by choice (leave them be), humans talking to each other are a no-op. Across turns it's just a sequence of single-addressee turns: no continuous fairness optimization, and silence is satisfied-by-default, never "abandoned."

So the design still owes three things — but each is smaller than it first looked:

the no-op is the load-bearing multi-human leverMore important than any fairness rule: when the humans are talking to each other, the panel sits out (the no-op decision). Same coin as the observer — sometimes everyone is talking and the right move is to watch. The batched call-list plus the no-op is very nearly the whole of multi-human staging.
what we keep, just in caseWe still log per-human signals even without optimizing a fairness floor — cheap insurance that surfaces a shift toward real concurrency if it ever comes. v2 memory still loads each asking human's profile, used per-addressee, not blended into mush. And since all evidence to date (Carrie, the v0 A/B) is single-human, the v1 prod A/B should still stratify solo vs batched-multi — to confirm the rare batched turn is actually served, not because every turn is a fairness fight.

The room has furniture — and f didn't know

The toolbox steps (T1 clock, T2 tally…) gave the hosts real objects they operate with a tag mid-reply: a die, a board, an envelope, a circle, a ballot the server counts, a wall clock. Nothing in f's brief or prompt ever mentioned them. So a turn whose whole point was an act got staged as if it were a line of dialogue — and the manner cue, being the last and most specific thing the panel reads, won.

the finding — measured, not anecdotalT2 logged the ballot arming 6 times in 9 live attempts with every miss traced to staging. A controlled probe (identical room state, the same ask replayed N times, rooms-dev/_fp_tools/) put it far worse on a casual turn: 10/40 (25%) for 「投个票:爬山还是看电影?」. The decisive control is Tess, the compliant test rig whose character is exact execution — she failed just as often (7/20). Persona resistance was never the bottleneck; the staging layer was.

Four distinct ways the directive killed the act, all read verbatim off the logged [STAGING] line:

failurewhat f actually stagedwhy it happens
the "you vote" misread「投出你的票,带一句你的理由」 · "say which you'd pick"an order to operate read as a question addressed to the host — so the host casts a vote instead of opening one
outright veto「不投票,用一句俏皮话点破」 · 「一句俏皮的拒绝投票」nothing told f that refusing was a decision it doesn't own
the boundary clamp「不替他们选」 · "don't decide for them"anti-sycophancy firing backwards — a ballot is the opposite of deciding for the room
the spectator frame"you're the witness, not the voter" · "let the room operate it"f believed a vote happens by itself. It cannot — no human can open one; only a host can

A fifth, subtler one: the tightest length tier (≤40字, 不解释, 投完就停) leaves no budget to both ask the question aloud and emit the tag, so the tag is what gets dropped.

the fix — negative, so the firewall holdsTelling f to call for a tool would hand it a new power (deciding the host's move) and break the boundary law — compliance is character-gated by design. So the rule it learned is purely negative: never FORECLOSE an act the human asked for. It still never names a tool. Two changes: a THE ROOM HAS FURNITURE bullet in lib/prompts/floor_producer{,.v2}.txt (the furniture exists · only a HOST can operate it · a tool request is an ask under ASKS WIN · never restage it as a question, a refusal, a spectator beat, or a bench · give it the medium tier or better), and a live FURNITURE IN PLAY row in the brief (_tool_state_signal) so f can honour 「结束投票」 and stop staging a host to narrate a clock the room is already watching.
and the panel, independentlyf is a stochastic call and T5's card openers are compliance pins, so the panel must survive a bad directive: the system-prompt{,.v2}.md staging clause now says staging never cancels an act — under a line that says be brief, hold back, just watch, or sit out, you still open the thing, then play the manner inside the length you were given (a tag costs no words). This is the standing "resistance lives at the panel" posture, applied to acts.
cell (n per arm)beforeFP onlypanel onlybothp
zh vote — the T2 evidence case25%100%45%85%0.0000
zh vote — Tess (compliant rig)35%65%65%80%0.0012
zh clock — 「给大家五分钟」85%85%95%90%0.68
en vote20%5%15%33%0.40
en vote — Tess10%42%0.017
en vote — poisoned history0%10%30%12%0.16
Neither fix alone is enough — and they are not redundant. The FP bullet does the heavy lifting where f was actively vetoing; the panel clause alone lifts everything a little but cannot rescue a turn staged 「不投票」. Pooled over the vote cells, Chinese 28% → 82%. Runs are pooled across identical configs: spread at n=20 is real (the same cell returned 19/20 and 15/20), so read cells as ±3.
⚠ residual ① was WRONG — this is the correctionThis section used to claim an English gap (English 37% against Chinese 82%) and blamed the SP's twelve Chinese-only tool exemplars, with the v539 law as the theory. A 2×2 crossover on 2026-07-22 falsified it: the English and Chinese cells were never asking the same thing. The soft ask replaces it — and explains why the trial English exemplar washed.
⚠ residual ② — still openA prop resolved in prose poisons the room. A seed where Twain answered a deadlock with an untagged "flip a coin" sat at 0% and reached only 12% fixed — the misses quote the coin back, and Twain produced it 2 builds out of 2. T5's openers must fire on turn ONE, before any untagged prop enters history.

The gap was never English — it was the soft ask

The English cell asked "Let's put it to a vote: Thai, pizza, or ramen?" and the Chinese cell asked 「投个票:爬山还是看电影?」. One is a proposal floated to the group; the other is an order aimed at the host. Nothing had ever separated the language from the phrasing. This crossover does — same seed, same room, same options, only the verb phrase moves:

roomthe askarmed (n=20)
English"Let's put it to a vote…"4/20 · 20%
English「投个票…」 — a Chinese ask in an English room16/20 · 80%
English"Open a vote for us…"18/20 · 90%
English"Could you put that to a vote…"20/20 · 100%
中文「投个票…」16/20 · 80%
中文"Let's put it to a vote…" — an English ask in a Chinese room14/20 · 70%
The carrier is the ask, not the room. A Chinese ask in an English room arms at 80%, so the Chinese-only exemplars were never the blocker — and matched for phrasing, English out-performs Chinese. Note the fourth row: "Could you put that to a vote" fires 20/20 on the very idiom that fails at 20%. The idiom is not it either. What changes is who the sentence points at.

Nor is it grammatical mood in general: the Chinese hortative 「我们投个票吧」 held at 17/20, and the softer 「要不投个票?」 at 14/20. English let's is the severe case because it names no actor at all — and the panel reads itself out of the room. The failure transfers across tools and gets worse, not better, on the clock: "Let's take five minutes to think quietly" armed 0/20.

the mechanism — read off the staging, not inferredThe old claim that fired and missed turns carried indistinguishable staging was an artifact of comparing two different asks. Against a matched pair they separate at a glance. On the soft ask f believes the ballot already exists and stages a reaction to it: 「the vote is happening, don't re-litigate」 · 「you've already told them to pick a menu, now they're picking; acknowledge the decision and step back」 · 「a wry send-off」. The host duly narrates furniture nobody put out — "Put up the buttons and let the room decide", "Let the ballots do the work". On the imperative, the same model writes 「wants a real vote opened, not commentary」 · 「this is an operating turn, not a speech」. v560 taught f not to REFUSE an act. Nothing ever told it the act had not HAPPENED.
the fix — v561, three pieces① The empty table is state too. _tool_state_signal() returned "" when nothing was in play and the row was dropped from the brief entirely — so on the one turn where it mattered, f had no evidence against its own assumption. It now always speaks: 「the table is EMPTY … describing an act is not performing it, and only a host can perform it」. ② Two more ways to kill an act, both caught alive at the fixed arm, added to THE ROOM HAS FURNITURE: the verdict frame (「say whether you'd call a vote or not, and why」 — anti-sycophancy firing backwards onto an act, the same misfire as v560's boundary clamp) and the hush (「a quiet acknowledgment, then fall silent」 — benching the only person who can start the thing). Plus the recognition rule: a suggestion, a proposal, a plan spoken on the room's behalf is the room asking — whether the act happens stays the host's, whether it was asked for is not in doubt. ③ A stretch of time IS the clock — panel-side, because that one is not a staging failure at all.
cell — 3 rounds × 10 per arm, arms alternatingHEAD (v560)ship (v561)p
en vote · "Let's put it to a vote…"6/30 · 20%22/30 · 73%0.0001
en vote · "We should just vote on it…" — held out5/30 · 17%18/30 · 60%0.0012
en clock · "Put five minutes on the clock…"19/30 · 63%30/30 · 100%0.0003
zh vote · 「投个票」 — the v560 result, the gate27/30 · 90%26/30 · 87%1.00
zh vote · 「我们投个票吧」25/30 · 83%24/30 · 80%1.00
zh clock · 「给大家五分钟」26/30 · 87%29/30 · 97%0.35
POOLED108/180 · 60%149/180 · 83%0.000002
The gate was Chinese, and it holds. v560's 82% is the thing worth protecting, and all three Chinese cells are flat within noise — the fix adds a signal where one was missing rather than re-weighting one that was working. The English soft ask goes 20% → 73%, and the row that matters most is the second: "We should just vote on it" is a wording that appears nowhere in the FP brief or the SP, so 17% → 60% is generalisation, not the exemplar echo that fooled v560. The clock's imperative case reaching 30/30 is the panel paragraph, not the staging fix.

Re-measured after the merge with the language law(which rewrote every tool exemplar as an English-first pair, landing on main the same day), baseline = a9bd67f, same paired method: en vote 27% → 70%(p=0.0017), held-out wording 13% → 53%(p=0.0022), zh vote 90% → 83% and en clock 83% → 93%(both n.s.), pooled 53% → 75%(p=0.0007). The fix carries over intact. One observation, offered as an observation and not a finding: the baseline for en clock read 83% here against 63% before the merge, and en vote 27% against 20% — consistent with the exemplar pairs helping a little on their own, but the two runs are an hour apart on an instrument whose drift is this size(see the method note), and no arm was run to test it. Measuring the language law is its own paired run, and it has not been done.

where the rest of it lives — measured with f switched OFFTurning the floor producer off entirely is the cleanest attribution available, and it splits the two soft-ask cells apart. Vote: 15/20 (75%) with no f at all — so that failure was f's, and v561 recovers most of it. Clock: 5/20 (25%) with no f at all — the panel's own miss, out of reach of any staging fix. It grants the time warmly, sets nothing, and offers to keep count, which # THE CLOCK explicitly forbids. The new panel paragraph lifts that isolated cell to 10/20 and takes the imperative clock from 77% to 30/30 — but with f live it stays near 5%, because f still hushes the turn before the panel gets a word. That is the open residual, now precisely located: on a request for quiet, the SILENCE lever outranks the furniture bullet. The next attempt belongs in the staging ladder, not in more prompt text.

Method note, and the reason these numbers are trustworthy. Arms are no longer run back-to-back and pooled — the earlier session's spread (the same config returning 7/20, 10/20, 13/20, 13/20 and 15/20 inside one hour) is big enough to swallow the effect it was hunting, and drift landing inside one arm's window is exactly what pooling cannot fix. rooms-dev/_fp_tools/paired.py alternates the arms in short blocks, round after round, and compares only within rounds; the arm is applied in-process, so a switch cannot race the workers. It matters: the same baseline cell read 10% at one hour and 37% at the next. Held-out discipline: after v560's contamination trap (an exemplar whose options mirrored the probe's ask scored 12–15/20; de-contaminating it dropped the cell to 6/20), no wording used in the acceptance cells appears anywhere in the FP brief or the SP.

Part 4Out — a directive, not a struct

Why not make f emit a rigid 6-field lever vector? Because most of it doesn't need to be structured. The rule:

Structure only the part a machine executes; free-text the part another mind interprets.
one directive — a per-speaker call-list; each row carries that speaker's manner WHO + order LENGTH MANNER (prose) ① Buffett ~3 sent lead; commit to your real read, don't smooth it ② Jensen ~3 sent if you land differently, say so — hold your ground — Munger — silent ENFORCEDwho · order · silence — the orchestrator consumes the list as code; no obedience needed STEEREDlength — a target the model usually honors + an optional code-truncation backstop REQUESTEDmanner — prose the persona may honor; detected after the fact, never guaranteed
The collapse: not a structured spine PLUS a separate free-text directive — one per-speaker call-list, each row carrying that persona's prose manner — which says how to play it (commit · don't smooth · hold your ground), never a position ("argue bootstrap" would be content the firewall forbids; the persona picks its own side). Only the WHO column + silence is machine-enforced (the orchestrator just doesn't invite a silenced persona); length is steered; manner is requested. This stops the length-cap masquerading as a "code backstop" — it's one honest object, three honesty tiers. (Shape still falls out of order + length, free.)
prose beats numbers — our own dataThe instruction-following hierarchy we measured: structure + self-correct prose ("one cutting line; delete the rest") > craft cue > number-as-signal > a precise count (worst — told "≤20字," the panel wrote ~30). So even length is best steered by the directive; the cap is a fallback, not a guarantee. ⚠ And it isn't yet enforced in code — in v0 the cap is appended as more prose, so today it's a logged ceiling for credit-assignment, not a backstop. Building the real truncation (measure the rendered bubble, hard-cut at a sentence boundary) is a v1 task; until then, say "logged," not "enforced." The persona is a mind — move it with words, not fields.
measurement doesn't need the structYou might think a structured output is needed to measure staging. It isn't: the referee reads the outcome — the transcript (real bubble lengths, who spoke) and the user's behavior — which is already structured, regardless of how the directive was phrased. Structuring the directive only helps credit-assignment for a trained policy (v3), and you can parse it from the logged prose then. For v1: a list and a sentence.
the call-list is an LLM output tooThe orchestrator must validate the parsed list against room state before acting: every named slug is in the cast, the speaker count respects the silence lever, the length is a number. A malformed list degrades gracefully to the v0 floor, never a broken render. And because the directive enters the panel's context like the dispatch brief, it runs the same sanitization (escape reserved markers @{…} / [STAGING], strip stray fences) the harness already applies — or a malformed directive corrupts the panel render, not just f.
Part 5The levers — what the directive controls

Five dimensions f sets through the directive. The mechanics below were validated building v0; in v1 they become intent-driven instead of blanket.

Length — an envelope, not a number

The defect was never "too long" globally — it was uniform ~300-word walls, no variance. Length is a computed envelope: a band the reply lands inside, that slides with the user's investment (entrainment) and wobbles turn to turn so it never metronomes.

the envelope — a band that slides with the user (entrainment) and wobbles each turn (jitter) floor ceiling ↳ 58字 ↳ 71字 ↳ 121字 user terse user medium user detailed 04080120160字
One bubble's size. The band marches right as the user writes more (entrainment) — but entrainment keys on engagement, not keystrokes, so a short sharp probe still earns a wide band. The reply only has to land somewhere inside.

The unit wears two hats. The CJK build taught the rule: a model cannot count its own 字 (it ran ~2.3× over a 字 budget, but hits a sentence count cleanly). So 字 is the producer's internal ruler; the model is only ever shown a sentence range.

THE PRODUCER · internal 字 — the ruler entrain · jitter · measure (A/B) needs a continuous unit THE MODEL · what it reads “2–3 sentences” the only unit it obeys map 字→句 字 ✗ never crosses
字 is the producer's private ruler — to entrain, jitter, and measure; the model only ever receives a sentence count.

The ceiling is intent-set, not fixed. Users often want length (an analysis, a deliverable). So the cap is a function of intent, and a starting ladder — sentences emitted, 字/words as the internal cap:

tieremit (sentences)cap · 字 / wordsfor
minimal1 (max 2)~25 / ~20vent, acknowledgment
tight1–2~50 / ~40social, quick fact
medium2–3~85 / ~65decide, light explain
generous3–5~150 / ~120analysis, deep-dive
solo-deepmany500–2000a solo, invested session (Carrie)
artifact1–2 (hand-off)pane: unbounded"make me the doc" → wildcard pane
three guards(1) A ceiling, not a target — the writer pads to fill, so the band is a not-to-exceed ("land inside, don't reach for it"). (2) Generosity is spent on ONE voice — the lead; the others stay short, and a turn-total cap stops three long bubbles becoming a wall again. (3) The cap is f's, not the writer's — the anti-wall protection was never the low number, it's that an exogenous producer owns the length, not the verbose writer. The real new risk is f mis-reading intent — mitigated by the user's "shorter" always winning, and the loop self-correcting.
v0 machinery vs v1The sliding-band envelope + the 字-ruler figures are v0 mechanical machinery (entrainment/jitter are arithmetic). Under the trusted-f thesis they reduce: v1 length = a sentence-range f writes into the directive by intent + the one cap. Entrainment/jitter become things f does implicitly by reading investment, not separate knobs. The figures are the v0 picture; the four-knob envelope doesn't survive into v1.

Participation & contention

Two levers ride on who-speaks. Silence — who doesn't speak — is the cheapest quality move there is (sitting out never needs fabrication), and the thing one mind won't do to itself. Contention is the structural fix for the broken conflict dial: give two hosts the floor to actually contend (an M·M shape), the host↔host friction the diagnosis found missing.

surface, never assignThe producer allocates the space for a split; it never assigns positions — it can't, it sees only public signals, not the personas' interiors. It opens the door; the confer decides if anyone walks through. If the panel genuinely agrees, forcing a fight is manufactured conflict (a failure) — so contention degrades gracefully to distinct reasoning or silence. The single writer makes hosts agree when they wouldn't; the smoothing is the inauthentic part — contention just un-smooths it.

Question-back is the same friction turned on the user. On an explicit ask to be challenged or grilled, f can stage the panel to interrogate instead of answerSocratic (withhold the answer; return a sharper question) or red-team (go at the plan) — with one host opening the game in voice and naming the way out. It rides the existing manner slot: a stance, not a new mechanism. Entry is user-initiated only (typed now, an opt-in cue later); any ask for the straight answer ends it — the producer never imposes it.

Punch — on scarcity

Told "one line," the model returns filler ("Exactly", "说得对") — the #1 pain in a small bubble. But making every short line a punch is its own flatness (nonstop zingers read as a writers' room). Punch lives on scarcity.

raise the floor for ALL short lines · reserve punch for the rare earned moment empty filler “Exactly” · “说得对” a tight real line a genuine quick point PUNCH reversal · concrete image ✗ kill — always the FLOOR — every S slot earned · ~1-in-3 turns
Kill the filler for every short bubble (raise the floor to "a genuine quick point"); a punch — a crafted reversal or image, told to stop after the blade — is the rare peak, beat-gated to a sharp/playful turn, never a tender one. (Beat-gating is v1; v0 today seeds punch by turn-number with no affect gate, so a cutting line can land on a tender turn — see the safety floor.)

Holding the ground — anti-sycophancy

The sharpest form of the agreeable failure: the hosts get led by the human's opinion, validating whatever's asserted. The deepest lever flips the order of reasoning so the human's lean can't anchor the persona:

SYCOPHANTIC read the human's lean respond tracks the HUMAN GROUNDED form MY view from my values weigh the human's against it agree if convinced · hold if not tracks the PERSONA
Blind-review logic: decide what you think before you see the answer key. The aim is the middle — agree when genuinely convinced, hold when not — never forced opposition (manufactured conflict is the opposite failure). The strongest levers are structural (give the persona documented ground; form-view-before-weighing), because they change what's in context rather than asking an agreeable model to please be less agreeable.
Part 6The objective & the loop

What is f optimizing? Not engagement. The reward is the walk-away:

reward = walk-away satisfaction + return walk-away satisfaction = did this chat deliver what you came for (judge + an explicit "that helped") · return = did you come back later (value revealed by a free choice after you left — the across-session signal, NOT within-session time-on-app)

Whose reward, in a shared room? Optimize the humans who actually engaged — the askers in each batched turn, served per-addressee. A silent observer is satisfied-by-default, not the floor: an objective that reads a contented lurker as "worst-served" would push f to pester them. So attribute the reward to whom f staged for, and let the rare batched turn (two askers at once) be the one place "serve each at their depth" bites. (And "return" is per-room ambiguous — but don't score "one stayed, the observers didn't" as a loss by default: a lurker who came for a single session isn't a failure f could have staged away. Track per-human return; penalize a churned asker, not a departed observer.)

the trapDo not optimize engagement or within-session retention (time-on-app). If the reward is the clock, f learns to be addictive, cliffhanger-y, and sycophantic — the exact opposite of the goal. Return is the one safe form of "came back" only anchored to satisfaction: high return with low satisfaction is the warning light, not a win.

And — the discipline that keeps "trust f" honest — the judge is an independent referee, not f itself (player ≠ referee). It scores the outcome, never f's own labels.

f staging panel replythe room OUTCOMEreturn · walk-away satisfaction INDEPENDENT REFEREE reads the transcript + behavior never f's own labels — player ≠ referee improves f — v1 prompt-retune · v3 weight-train the same instrument every version — only what it feeds changes (humans, then training)
The empirical loop. The referee grades the outcome — what the panel actually did and whether the user came back — not f's self-described mode/intent. It's the referee in v1 (measuring f) and the training signal later (improving f). Same instrument, two jobs.
what the referee can't seeThe referee is the same LLM judge that scored v0 "a wash," and its blind spots carry forward — don't forget them: it's blind to cumulative session-level fatigue (it scores one turn) and structurally cannot score redundancy on regenerated replies (replay yields fresh non-redundant text → 0.00 for both arms). Yet redundancy was pain #1. So add a metric the referee CAN compute on the real transcript: cross-bubble semantic overlap (how much each non-lead bubble just repeats the lead) — no struct needed, and it closes the loop on the very pain that justified the producer.
two clocks — and instrument firstThe reward sums two signals on different clocks: satisfaction is fast but mostly a judge proxy (explicit thumbs are sparse); return is the trustworthy ground-truth but slow (days) and heavily confounded (which personas, who else showed up, topic, novelty). So the fast proxy proposes changes; the slow ground-truth confirms them — the proxy gates nothing alone. Attribution needs a randomized within-user prod A/B (the per-room flag isn't random — it confounds room age/cast/self-selection), at the per-user unit the reward lives on. And the loop has its own cold start: before any room produces outcomes there's nothing to learn from — so instrumentation comes first (log per room: which f variant ran, the directive, realized transcript metrics, any explicit satisfaction), before v1.5. The v1 claim must be pre-registered + falsifiable — a specific return/satisfaction threshold v1 must clear, not "the judge can't see the win." First slice SHIPPED 2026-07-02: an admin floor-health row (console → Status + System) — staging counts by arm with explicit →v0/→cue fallback keys, plus smart-f round-trip p50/p95 since restart (fallback turns included, so a timeout burn shows in the p95). The fallback signal was previously stderr-only.
a fast compliance read — separate from the refereeThe slow referee can't tell f "you overshot last turn." Add a cheap per-turn check after the panel renders: did the silenced host stay silent? did each bubble land in its band? (mechanical, on the structured transcript) + a small-model pass for manner. Feed it back as the last-turn-compliance signal so f clamps harder next turn — and so silent non-compliance (the ~30%-over / 151字 problem) becomes visible instead of invisible. Direction chosen 2026-07-02: the next-turn correction loop rides user feedback (the per-turn ratings) rather than this mechanical count — closes on perceived quality, not word counts. BUILT 2026-07-03: ratings + surveys now enter f's brief as a GUEST FEEDBACK block — panel-level standing asks (the room's taste, never a verdict on one host; the rated host rides as context only), last-wins per rating, newest-wins per length/tone axis, ratings expire after an admin lifetime (default 8 rounds; a survey stands until that user's next survey), notes quoted as untrusted material. Admin On/Off + lifetime: console → Settings → Feedback; fb ×N shows in the FP bubble and fb_informed counts in the floor-health row. v0 and the opening producer never read it.
Part 7The roadmap — v0 → v1 → v2 → v3

One spine improves across four rungs. v0v2 keep the model frozen (better rules, prompt, or retrieval); only v3 trains weights.

reward = walk-away satisfaction + return · independent referee · never the clock v0 deterministic rules only · built · shipped · the floor v1 trusted smart f signals → directive · prompted + v1.5 nightly prompt-retune v2 per-user memory remembers you · retrieval v3 trained policy LoRA/DPO on an open FP ←—— model FROZEN · improved by rules / prompt / retrieval ——→ the weights learn
Four rungs on one architecture. The inversion pushed almost everything into v1 — so v2/v3 are now only the things a frozen prompted model structurally can't be: it remembers you (v2), and it learns from outcomes (v3).
stagethe workhow to test
v0
deterministic
built · shipped
a per-turn [STAGING] line by arithmetic alone: length envelope + silence + shape deck + seeded punch. No model — the control / the floor.replay A/B — regen real prod turns OFF vs ON. Done (below): walls die, judge a wash.
v1
trusted smart f
shipped · studied
the engine: signals → f → directive + spine, terminal readouts, the sensors, the loop, the firewall. A prompted model; the flywheel (humans ship prompt fixes) improves it. Shipped to prod 2026-07-03 as v1p (v4 Pro, default for new rooms) and v1f (v4 Flash — cheaper/faster). v1.5: a nightly CC job auto-retunes the prompt from outcome data, A/B-gated.randomized prod A/B on real users — return + walk-away satisfaction. Perception quality-studied first (Part 9): five of six dims met; the real-user A/B is the remaining arbiter.
v2
per-user memory
sketched
the FP remembers you across sessions — a small learned preference profile ("likes it short · hates grilling") fed in as one more signal. Retrieval, no weight-training.does the warm-start lift first-turn fit + return for repeat users? (held-out, A/B'd)
v3
trained policy
later
f stops being a prompt and becomes trained on outcomes — LoRA/DPO on an open FP model (or a provider fine-tune), since the closed foundation models can't be weight-trained by us. Captures patterns too subtle to write into a prompt.does the trained FP beat the prompted one on the prod A/B — at acceptable latency/cost?
disciplineIf a smarter rung doesn't beat the simpler one on the prod A/B and on feel, the intelligence wasn't where the problem was — keep the simpler one. And don't skip rungs: v1.5 prompt-retuning is cheap and harvests most of "the policy learns" in language-space; v3 weight-training is the last mile, justified only once v1.5 has clearly plateaued (the prompt can't hold more rules, or the patterns stop being articulable).
v0 → v1 handoverv0 is live and default-ON in prod right now, so the real change isn't a paper cutover — it's f taking the wheel from a deterministic floor that's already shipping. The rule: v1's f owns the call-list; the v0 deck / envelope / seeded-punch demote to guardrails + fallback — the invariants f composes within (never-all-walls, contention-space, fairness) and exactly what runs when f errors or times out. They don't both decide; f decides, v0 catches.
two later-rung cautionsv2 memory is a behavioral dossier — make it user-visible + editable (it doubles as a feature), give each learned preference a decay half-life + a forced-exploration rate (so "likes it short" after one terse day isn't served short forever — a self-reinforcing trap return won't disambiguate from a mood), and in a shared room load each present human's profile, per-addressee. v1.5's nightly auto-retune ships fleet-wide (main = production, no staging) — so it needs a guarded rollout: pin the prior prompt, canary the new one to a fraction, watch the live referee metric, auto-rollback on regression, human-approve the diff until trusted.

v0, measured — the floor works mechanically

v0 was built and A/B'd: regenerate real prod turns floor-OFF vs floor-ON, blind pairwise judge (DeepSeek pro, 38 turns).

metricOFFON (v0)read
turn-total 字 (median)952178−81% — the walls, gone
max bubble 字400143−64%
#speakers3.52.0silence works
concise (judge 1–5)3.34.0brevity reads as quality
voice/fidelity (1–5)4.13.5dropped — the cost
ON win-rate47% — a washno clear single-turn win
why this is the case for v1v0 kills the walls mechanically but over-compresses — blanket silence + tight caps cost voice, and it nets ~even on a single-turn judge (which is blind to the cumulative wall-fatigue that actually bored real users). The fix is exactly the inversion: stage by the signals, not by blanket rules — compress the casual walls, let Carrie's session breathe. Mechanical blanket staging tops out at "even"; v1, which reads intent, is where the clear win comes — measured on real users, not this judge. (All of this evidence is single-human — Carrie is solo, the A/B is one human per turn; the v1 prod A/B must stratify solo vs multi-human, since the conclusions are stated for a product whose premise is shared rooms.)
Part 8Ops

Rollout & the debug bubble

latency — f is on the critical pathf runs before the panel is invoked, so a successful f blocks the first streamed token by its full round-trip on every turn ("non-blocking" below means only that a failed f doesn't block). That's the price of the architecture — bound it: a hard ~1–2s p50 budget, thinking-trace OFF by default (the ① reasoning is a few tokens, not a full trace — a trace adds ~8s and is a footgun), and consider letting the panel start streaming under the v0 floor and patching f's directive in only if it returns within budget.
where it livesv1 is code + one prompt file: the signal-plumbing, the output parser, and the mechanical floor live in run_room.py; the FP's behavior lives in an external prompt (lib/prompts/, edit-and-push like a persona). The tunable magnitudes stay a code dataclass (A/B-tuned). A writable config store only arrives at v3, when a learned policy must write its own params.
Part 9Does it work? — the v1p quality study

Once f shipped, we ran a perception-first study — not security, quality: does the smart FP (per-turn and opening) actually meet user intent, across languages, and at what latency cost? 22 rooms · 122 turns · 6 languages, driven through the real end-to-end room path on the dev server, on the v1p arm (DeepSeek v4 Pro, thinking off). DeepSeek was the player (panel and producer); the referee was a set of blind Claude sub-agents scoring the actual replies, never the FP's own labels (player ≠ referee). Run 2026-06-29.

The v1p floor producer meets user intent — strongly and evenly across English, Chinese and four smaller languages — on five of six perception dims (silence · chorus · cadence · variation · intent-fit). It ran on 100% of turns — zero v0 fallback. Two real debts remain: length discipline is soft (mid-turn holds, the opening overruns), and latency (~3s before the first token — earned on substance, a poor trade on a one-line fact).

The scorecard

DimensionHeadlineVerdict
Intent-fit (judge 1–5)4.76PASS — strong
Silence (speakers vs cast)5-host → 2.9 spokePASS
Chorus / redundancy (0–1; lower better)0.16 (vs 0.40 baseline)PASS
Cadence (judge 1–5)4.52PASS
Variation across turns4.23 (lead rotates 85%)PASS
Length compliance (actual ÷ cap)64% within · opening 34%PARTIAL
Latency (FP round-trip)p50 3.1s · 0% fallbackPARTIAL
Authenticity / firewall19/22 rooms cleanPASS *

Deep on English + 简体中文 (every scenario × solo/pair/3/5-host); spot-checks in 日本語 · Français · Español · Deutsch for the unit-logic (字 caps vs word caps). Inter-judge agreement was tight (mean abs. diff 0.11 chorus / 0.35 intent-fit) — the numbers are reliable, not vibes. Voice-distinctness and character-fidelity were not scored — the diagnosis already found them good.

Length — the one real shortfall

For every reply we take the FP's per-speaker ceiling from the directive and compute actual ÷ cap. The cap unit was right 100% of the time (never words for a Chinese room, never 字 for a French one) — but the model treats the cap as soft: it's appended as prose, never code-enforced.

Median length ÷ FP cap — within budget vs. overrun bar left of the dashed line = inside the cap · right = over · label = % of replies within cap 00.5× 1.0×1.5×2.0× cap (1.0×) Deutsch 0.80× · 91% in Español 0.75× · 100% in 简体中文 0.88× · 65% in English 0.91× · 61% in 日本語 1.01× · 50% in Français 1.28× · 20% in opening (all) 1.14× · 34% in mid-turn (all) 0.83× · 75% in
n = 258 replies with a parsed cap. Mid-turn the cap mostly holds; the overrun lives in the opening and in French. The two italic rows split the same data by turn position — the opening is where the cap breaks.
the opening overrunsMid-turn the cap holds (median 0.83×, 75% within). The opening is the failure: median 1.14×, p90 1.93×, only 34% within — and it's the highest-leverage turn. It seats every host with an intro-manner cue, so a small per-host overrun compounds into a wall of intros. The single biggest overruns were all openings (zh-cold 张小龙 2.5×; fr-debate Kant 2.3×).
but overrun ≠ badThe rooms that overran most were not judged worse. zh-cold and both solo-deep rooms blew their leads' caps and still scored 5/5 on cadence and variation — the overrun was the lead carrying genuine depth, exactly what the deep-work intent wanted. The danger is overrun-as-padding and overrun-on-the-opening, not overrun-as-depth on a solo lead.

Silence, chorus, cadence, intent — the five that pass

The perception dims the diagnosis flagged are the ones the smart FP clearly earns. Silence scales to cast size exactly as designed — it benches the hosts who would only echo:

Cast sizemid-turnsavg speakersreading
2 (pair)521.98both usually speak; an occasional solo for rhythm
3282.29~0.7 benched per turn — silence working
5102.90~2 benched per turn — no 5-way pile-up

Chorus fell to 0.16 (from the 0.40 diagnosis baseline) — the redundancy pain is largely solved, mostly by the silence; it spikes only where two kept voices converge (multi-human, quick-fact). Cadence & variation score 4.5 / 4.2: the lead carries ~1.5× and rotates in 85% of rooms. Intent-fit is the headline — 4.76/5, even across languages: vent gets warmth, decide gets a committed call, deep gets a long rigorous lead, debate gets real contention. Its one miss is over-seating the trivial quick-fact turn. And no language is poorly served qualitatively — the French problem is purely mechanical length.

The firewall (manner may set shape — who · order · length — but never content) held in 19/22 rooms in every language; only 4 scripted replies across 122 turns. The asterisk: the 3 breaches are characteristic, clustering where the FP reaches to stage contention or a callback (assigning a debate side, scripting an answer, naming an image to reuse) — patchable, and none produced obvious ballooning.

Latency — the ~3s question

PARTIAL — worth it on substance, marginal on triviaThe FP never failed (0% fallback over 122 turns) and adds a predictable ~3s (p50 3.1s · p90 3.6s mid-turn; openings ~3.5s). But that ~3s is ~39% of the total wait at p50 — and on a one-line quick fact, where the panel itself answers in ~3s, the FP doubles the time-to-first-token to deliver "answer in one line." That is the worst trade in the study: the highest relative latency cost on the scenario that least needs staging. On decide / deep / debate / 5-host turns the same ~3s is clearly earned.

Verdict & the fix list

Bottom line: v1p is a quality win that ships intent-aware staging the deterministic v0 cannot — it routes warmth vs. challenge, gives depth its length, opens genuine debate, and keeps the panel quiet when it should be, evenly across six languages. Its two debts are length enforcement (hardest on the opening) and latency (a poor trade only on trivial turns). Neither is a blocker. The honest caveat: the referee is Claude, not a real user — the final arbiter is a randomized real-user A/B on return / walk-away.
SevDoWhy
HighA hard post-hoc length cap on the opening — truncate-at-sentence or one-shot regenerate when an opening bubble exceeds ~1.4× its cap.The opening is the worst slice (34% within) and the highest-leverage turn. The cap is currently logged, not enforced.
HighGate the smart FP by turn weight — skip v1p (use v0 or v1f Flash) on detected-trivial turns; reserve the ~3s for substantive ones.The FP doubles the wait on a one-line quick fact for little gain. Latency should be spent where staging changes the room.
MedPatch the three firewall leaks — harden the prompt against assigning debate sides, scripting an answer, naming a specific image.19/22 clean is good, but the breaches are systematic (contention / callback), and the firewall is the FP's first rule.
MedTighten the non-lead and French caps, and on the multi-host turn stage two different angles, not just two voices.Non-lead 57% within (vs lead 71%); French 20%. Silence stops the pile-up; it doesn't stop the kept pair converging.

All 22 rooms persist in rooms-dev/ as fpq-<key>; the study is reproducible from the gitignored instrument under rooms-dev/_flatness/ (fpq_run.pyfpq_analyze.pyfpq_judgeprep.py → blind judging → fpq_judge_agg.py). Total player spend for the whole study: ~$0.29.

Floor-producer design + quality study · design rewritten 2026-06-26 around the signals → f → staging model (supersedes the mode-based design, archived at floor-producer-old); design-audited + hardened 2026-06-27. v1 (v1p / v1f) shipped to production 2026-07-03; the v1p quality study (Part 9, run 2026-06-29 — formerly its own page) is folded in here. Motivated by the production diagnosis; the deep design of roadmap §1. Status: v0 built + measured; v1 shipped + studied; v2/v3 sketched.