The build spec for roadmap §1, off the back of the production diagnosis. The diagnosis proved where conversation quality breaks; this is the machine that fixes it. The producer is one function — f — that reads a rich signal vector and emits the staging. Status: v0 deterministic baseline built + A/B-measured; v1 the trusted smart f (v1p / v1f) shipped to production 2026-07-03 (v1p default for new rooms) and quality-studied; v2 per-user memory + v3 a trained policy sketched. Design rewritten 2026-06-26; quality study run 2026-06-29 (Part 9).
f → staging model here. The old version is kept for reference: floor-producer (mode-based · outdated) →.We pulled all 80 production rooms and scored the 39 real conversations (full report). Four findings:
extreme delivers ~⅓ of the friction it promises — and the heat that does exist goes host→human, not host↔host.The load-bearing conclusion: three of the four are staging, not voice. Voice and character fidelity are good (users said so). The single writer just stages the room badly — and can't fix itself. That is exactly the gap an exogenous producer fills.
fThe producer's whole job is a single function: a high-dimensional signal vector in — your engagement, your explicit asks, how invested you are, the substance, the room's structure (solo or panel), the phase — and the staging out. We trust a smart model with all the signals and let it decide. No taxonomy in the path.
The earlier design tried to sort each conversation into one of seven named "modes." That's the wrong primary abstraction, and one real production room proved it (Carrie, below). The deepest way to say why:
So the engine is one call: signals in, staging out. Nothing is forced through a discrete bottleneck on the way.
f emits staging, never content. Mode and intent are byproducts — emitted last, routed to a debug surface, never into the engine. Describe with labels; decide with signals.f still names what it did — a mode ("this reads as deep self-work") and an intent ("she wants a rigorous assessment"). But these are labels read off the decision, for your eyes and the logs — not steps in it. And the order matters: in one autoregressive call, what's generated earlier conditions what comes later, so generation order is the causal graph. Put the label first and the staging anchors on it (the old gate, smuggled back in). Put it last and it can only describe.
The panel already deliberates — the backstage <confer> is the one mind working out what to say. The producer doesn't fight that; it sits upstream. The line that keeps it a router, not a co-author:
f may read "this is grief," but that only sets the staging (length, who, how hard); it never tells a host what to say. Inferring intent to drive staging is the job; injecting it as content ("argue bootstrapping is better") is puppeteering — and it's why the output is mostly a directive about how to play it, not the lines themselves. On who / length / silence the producer wins (being brief or sitting out never needs fabrication); on contention it opens the space for a split, and the confer decides if anyone walks through.f — also one call — trustworthy? Not infallibility. It's structurally better placed: a narrow job (staging, not content), cheap to retry, and graded by outcomes it doesn't author (the referee). And it never overrides the writer's ground — the directive sets staging and may ask for a register, but the persona's authored voice and red-lines win over a conflicting manner cue. f opens space; the persona fills it as itself.Since there's no taxonomy to maintain, the real engineering becomes plumbing the signals into f. It can only be smart about what it's shown. The vector:
f reads| signal | what it carries | cost |
|---|---|---|
| engagement quality | is the latest message a sharp probe or a flat "ok"? going deeper or cooling off? | model — the loudest cue |
| explicit asks · dials | "@X · shorter · go deeper · make a doc · don't comfort me" — priority over inferred intent. (The Debate/Personality/Challenge dials are not a separate channel — just a UI for emitting these asks.) | cheap — parsed |
| investment | how much each human is putting in — a 543-字 opener vs "hi" | cheap |
| roster | each seated persona's card — slug · role · tagline · tags. Required for who-speaks-by-relevance (which host can actually carry this turn) | free — room state |
| structure | solo vs panel · how many personas · how many humans (see Multiple humans) | free — room state |
| furniture in play | what is physically on the room's table this turn — board up · clock running (m:ss left) · ballot open (n of m voted) · a die in someone's hand. Added 2026-07-22: without it f stages blind to the room's own state (see the room has furniture) | free — room state |
| substance | the real ask, and — the contention judgment — is there a genuine split? | model |
| phase | opening · exploring · converging · wrapping | model-light |
| history · brief | the transcript tail (last 8 turns, 160 chars each) — engagement + phase — plus the GUEST FEEDBACK block (built 2026-07-03): the humans' recent 👍/👎 ratings + each user's latest survey, rendered as standing asks — panel-level, admin On/Off + an N-turn lifetime (default 8; console → Settings → Feedback). Still not plumbed, by design: the §5 situation box and mechanical last-turn compliance — the feedback loop replaced the compliance read (see Part 6). | free — room state |
"Cheap" vs "model" is honest here: explicit-asks/investment/roster/structure/history are parse-or-free; engagement, substance, phase are model judgments — the headline sensor (engagement quality) is one of them, not free.
f reads them → staging, same as every turn. No mode-prior fallback needed. (The one genuinely turn-1-specific call is who opens a panel with no transcript — decided by topic-to-persona affinity, the brief/title × the roster, not recency.)The one sensor that earns its keep first. The naive rule — "match the user's length" (entrainment) — is a trap: it would clamp a deeply-engaged user who fires short, sharp probes. The fix is to read engagement quality (substance + trajectory), not raw message length:
| the message | naive length read | engagement-quality read |
|---|---|---|
挡住的是我的焦虑吧 (9字) | short → reply short ✗ | sharp probe, escalating → deep → stay long |
嗯,谢谢 (would-be) | short → reply short | flat, closing → clamp / wind down |
So intent owns the ceiling, not the keystroke count: a short-but-sharp probe in a deep session keeps the reply long. That's the concrete mechanism behind "let the substantive turns breathe." (Caution: short-sharp-escalating is also the signature of distress — same surface, opposite correct response. The read must route distress to the safety floor, not "push harder.")
A real production room (梁宁 · 自我探索, solo): the user fires 8–98字 probes and gets 900–1886字 of rigorous analysis back, for 18 turns — escalating depth, explicitly asking "你不要安慰我,以面试官的角度客观评价," pushing back, never "太长." She is clearly satisfied; the long replies are the value.
She's emotional·converge — Vent's coordinate on the old grid — but wants the opposite staging (generous + challenge, not minimal + warmth). The grid would mis-serve her. And within her single session, the staging swung wildly by the turn's ask:
| her turn (the signal) | the staging f should produce |
|---|---|
| "陪我探索" — vulnerable opener | ~999字, warm |
| "目前主要在接商单" — flat update | ~336字 |
| "你不要安慰我,以面试官的角度客观评价" | ~1623字 · challenge HIGH |
| "乔布斯、张小龙是什么样的" — teaching ask | ~1886字, exposition |
The product's defining case — but the way to serve it is not a concurrent-fairness optimization haunting every turn. The room is slow-paced and turn-batched: human messages queue and drain together into one turn, and that turn usually holds exactly one human. That collapses the problem — multi-human is a within-turn question, and across turns it's just a sequence of mostly-single-addressee turns, each handled fresh by reading who actually spoke.
So the design still owes three things — but each is smaller than it first looked:
f reads each asker's trajectory separately; when it carries one (the common case) the read is trivial. (v0 used to entrain length to the concatenated batch — it summed and over-served; fixed: it now keys the lead cap on the longest single message, so a solo turn is unchanged and a batched turn no longer inflates. A true per-asker read still waits for v1.)f owes no fairness move to drag a lurker in; an organic turn-to ("C, you've built one — what do you think?") is an optional warm gesture when it fits, never an obligation. Pestering a happy observer is the failure, not leaving them be.f didn't knowThe toolbox steps (T1 clock, T2 tally…) gave the hosts real objects they operate with a tag mid-reply: a die, a board, an envelope, a circle, a ballot the server counts, a wall clock. Nothing in f's brief or prompt ever mentioned them. So a turn whose whole point was an act got staged as if it were a line of dialogue — and the manner cue, being the last and most specific thing the panel reads, won.
rooms-dev/_fp_tools/) put it far worse on a casual turn: 10/40 (25%) for 「投个票:爬山还是看电影?」. The decisive control is Tess, the compliant test rig whose character is exact execution — she failed just as often (7/20). Persona resistance was never the bottleneck; the staging layer was.Four distinct ways the directive killed the act, all read verbatim off the logged [STAGING] line:
| failure | what f actually staged | why it happens |
|---|---|---|
| the "you vote" misread | 「投出你的票,带一句你的理由」 · "say which you'd pick" | an order to operate read as a question addressed to the host — so the host casts a vote instead of opening one |
| outright veto | 「不投票,用一句俏皮话点破」 · 「一句俏皮的拒绝投票」 | nothing told f that refusing was a decision it doesn't own |
| the boundary clamp | 「不替他们选」 · "don't decide for them" | anti-sycophancy firing backwards — a ballot is the opposite of deciding for the room |
| the spectator frame | "you're the witness, not the voter" · "let the room operate it" | f believed a vote happens by itself. It cannot — no human can open one; only a host can |
A fifth, subtler one: the tightest length tier (≤40字, 不解释, 投完就停) leaves no budget to both ask the question aloud and emit the tag, so the tag is what gets dropped.
f to call for a tool would hand it a new power (deciding the host's move) and break the boundary law — compliance is character-gated by design. So the rule it learned is purely negative: never FORECLOSE an act the human asked for. It still never names a tool. Two changes: a THE ROOM HAS FURNITURE bullet in lib/prompts/floor_producer{,.v2}.txt (the furniture exists · only a HOST can operate it · a tool request is an ask under ASKS WIN · never restage it as a question, a refusal, a spectator beat, or a bench · give it the medium tier or better), and a live FURNITURE IN PLAY row in the brief (_tool_state_signal) so f can honour 「结束投票」 and stop staging a host to narrate a clock the room is already watching.f is a stochastic call and T5's card openers are compliance pins, so the panel must survive a bad directive: the system-prompt{,.v2}.md staging clause now says staging never cancels an act — under a line that says be brief, hold back, just watch, or sit out, you still open the thing, then play the manner inside the length you were given (a tag costs no words). This is the standing "resistance lives at the panel" posture, applied to acts.| cell (n per arm) | before | FP only | panel only | both | p |
|---|---|---|---|---|---|
| zh vote — the T2 evidence case | 25% | 100% | 45% | 85% | 0.0000 |
| zh vote — Tess (compliant rig) | 35% | 65% | 65% | 80% | 0.0012 |
| zh clock — 「给大家五分钟」 | 85% | 85% | 95% | 90% | 0.68 |
| en vote | 20% | 5% | 15% | 33% | 0.40 |
| en vote — Tess | 10% | — | — | 42% | 0.017 |
| en vote — poisoned history | 0% | 10% | 30% | 12% | 0.16 |
f was actively vetoing; the panel clause alone lifts everything a little but cannot rescue a turn staged 「不投票」. Pooled over the vote cells, Chinese 28% → 82%. Runs are pooled across identical configs: spread at n=20 is real (the same cell returned 19/20 and 15/20), so read cells as ±3.The English cell asked "Let's put it to a vote: Thai, pizza, or ramen?" and the Chinese cell asked 「投个票:爬山还是看电影?」. One is a proposal floated to the group; the other is an order aimed at the host. Nothing had ever separated the language from the phrasing. This crossover does — same seed, same room, same options, only the verb phrase moves:
| room | the ask | armed (n=20) |
|---|---|---|
| English | "Let's put it to a vote…" | 4/20 · 20% |
| English | 「投个票…」 — a Chinese ask in an English room | 16/20 · 80% |
| English | "Open a vote for us…" | 18/20 · 90% |
| English | "Could you put that to a vote…" | 20/20 · 100% |
| 中文 | 「投个票…」 | 16/20 · 80% |
| 中文 | "Let's put it to a vote…" — an English ask in a Chinese room | 14/20 · 70% |
Nor is it grammatical mood in general: the Chinese hortative 「我们投个票吧」 held at 17/20, and the softer 「要不投个票?」 at 14/20. English let's is the severe case because it names no actor at all — and the panel reads itself out of the room. The failure transfers across tools and gets worse, not better, on the clock: "Let's take five minutes to think quietly" armed 0/20.
f believes the ballot already exists and stages a reaction to it: 「the vote is happening, don't re-litigate」 · 「you've already told them to pick a menu, now they're picking; acknowledge the decision and step back」 · 「a wry send-off」. The host duly narrates furniture nobody put out — "Put up the buttons and let the room decide", "Let the ballots do the work". On the imperative, the same model writes 「wants a real vote opened, not commentary」 · 「this is an operating turn, not a speech」. v560 taught f not to REFUSE an act. Nothing ever told it the act had not HAPPENED._tool_state_signal() returned "" when nothing was in play and the row was dropped from the brief entirely — so on the one turn where it mattered, f had no evidence against its own assumption. It now always speaks: 「the table is EMPTY … describing an act is not performing it, and only a host can perform it」. ② Two more ways to kill an act, both caught alive at the fixed arm, added to THE ROOM HAS FURNITURE: the verdict frame (「say whether you'd call a vote or not, and why」 — anti-sycophancy firing backwards onto an act, the same misfire as v560's boundary clamp) and the hush (「a quiet acknowledgment, then fall silent」 — benching the only person who can start the thing). Plus the recognition rule: a suggestion, a proposal, a plan spoken on the room's behalf is the room asking — whether the act happens stays the host's, whether it was asked for is not in doubt. ③ A stretch of time IS the clock — panel-side, because that one is not a staging failure at all.| cell — 3 rounds × 10 per arm, arms alternating | HEAD (v560) | ship (v561) | p |
|---|---|---|---|
| en vote · "Let's put it to a vote…" | 6/30 · 20% | 22/30 · 73% | 0.0001 |
| en vote · "We should just vote on it…" — held out | 5/30 · 17% | 18/30 · 60% | 0.0012 |
| en clock · "Put five minutes on the clock…" | 19/30 · 63% | 30/30 · 100% | 0.0003 |
| zh vote · 「投个票」 — the v560 result, the gate | 27/30 · 90% | 26/30 · 87% | 1.00 |
| zh vote · 「我们投个票吧」 | 25/30 · 83% | 24/30 · 80% | 1.00 |
| zh clock · 「给大家五分钟」 | 26/30 · 87% | 29/30 · 97% | 0.35 |
| POOLED | 108/180 · 60% | 149/180 · 83% | 0.000002 |
Re-measured after the merge with the language law(which rewrote every tool exemplar as an English-first pair, landing on main the same day), baseline = a9bd67f, same paired method: en vote 27% → 70%(p=0.0017), held-out wording 13% → 53%(p=0.0022), zh vote 90% → 83% and en clock 83% → 93%(both n.s.), pooled 53% → 75%(p=0.0007). The fix carries over intact. One observation, offered as an observation and not a finding: the baseline for en clock read 83% here against 63% before the merge, and en vote 27% against 20% — consistent with the exemplar pairs helping a little on their own, but the two runs are an hour apart on an instrument whose drift is this size(see the method note), and no arm was run to test it. Measuring the language law is its own paired run, and it has not been done.
f switched OFFTurning the floor producer off entirely is the cleanest attribution available, and it splits the two soft-ask cells apart. Vote: 15/20 (75%) with no f at all — so that failure was f's, and v561 recovers most of it. Clock: 5/20 (25%) with no f at all — the panel's own miss, out of reach of any staging fix. It grants the time warmly, sets nothing, and offers to keep count, which # THE CLOCK explicitly forbids. The new panel paragraph lifts that isolated cell to 10/20 and takes the imperative clock from 77% to 30/30 — but with f live it stays near 5%, because f still hushes the turn before the panel gets a word. That is the open residual, now precisely located: on a request for quiet, the SILENCE lever outranks the furniture bullet. The next attempt belongs in the staging ladder, not in more prompt text.Method note, and the reason these numbers are trustworthy. Arms are no longer run back-to-back and pooled — the earlier session's spread (the same config returning 7/20, 10/20, 13/20, 13/20 and 15/20 inside one hour) is big enough to swallow the effect it was hunting, and drift landing inside one arm's window is exactly what pooling cannot fix. rooms-dev/_fp_tools/paired.py alternates the arms in short blocks, round after round, and compares only within rounds; the arm is applied in-process, so a switch cannot race the workers. It matters: the same baseline cell read 10% at one hour and 37% at the next. Held-out discipline: after v560's contamination trap (an exemplar whose options mirrored the probe's ask scored 12–15/20; de-contaminating it dropped the cell to 6/20), no wording used in the acceptance cells appears anywhere in the FP brief or the SP.
Why not make f emit a rigid 6-field lever vector? Because most of it doesn't need to be structured. The rule:
@{…} / [STAGING], strip stray fences) the harness already applies — or a malformed directive corrupts the panel render, not just f.Five dimensions f sets through the directive. The mechanics below were validated building v0; in v1 they become intent-driven instead of blanket.
The defect was never "too long" globally — it was uniform ~300-word walls, no variance. Length is a computed envelope: a band the reply lands inside, that slides with the user's investment (entrainment) and wobbles turn to turn so it never metronomes.
The unit wears two hats. The CJK build taught the rule: a model cannot count its own 字 (it ran ~2.3× over a 字 budget, but hits a sentence count cleanly). So 字 is the producer's internal ruler; the model is only ever shown a sentence range.
The ceiling is intent-set, not fixed. Users often want length (an analysis, a deliverable). So the cap is a function of intent, and a starting ladder — sentences emitted, 字/words as the internal cap:
| tier | emit (sentences) | cap · 字 / words | for |
|---|---|---|---|
| minimal | 1 (max 2) | ~25 / ~20 | vent, acknowledgment |
| tight | 1–2 | ~50 / ~40 | social, quick fact |
| medium | 2–3 | ~85 / ~65 | decide, light explain |
| generous | 3–5 | ~150 / ~120 | analysis, deep-dive |
| solo-deep | many | 500–2000 | a solo, invested session (Carrie) |
| artifact | 1–2 (hand-off) | pane: unbounded | "make me the doc" → wildcard pane |
f's, not the writer's — the anti-wall protection was never the low number, it's that an exogenous producer owns the length, not the verbose writer. The real new risk is f mis-reading intent — mitigated by the user's "shorter" always winning, and the loop self-correcting.f thesis they reduce: v1 length = a sentence-range f writes into the directive by intent + the one cap. Entrainment/jitter become things f does implicitly by reading investment, not separate knobs. The figures are the v0 picture; the four-knob envelope doesn't survive into v1.Two levers ride on who-speaks. Silence — who doesn't speak — is the cheapest quality move there is (sitting out never needs fabrication), and the thing one mind won't do to itself. Contention is the structural fix for the broken conflict dial: give two hosts the floor to actually contend (an M·M shape), the host↔host friction the diagnosis found missing.
Question-back is the same friction turned on the user. On an explicit ask to be challenged or grilled, f can stage the panel to interrogate instead of answer — Socratic (withhold the answer; return a sharper question) or red-team (go at the plan) — with one host opening the game in voice and naming the way out. It rides the existing manner slot: a stance, not a new mechanism. Entry is user-initiated only (typed now, an opt-in cue later); any ask for the straight answer ends it — the producer never imposes it.
Told "one line," the model returns filler ("Exactly", "说得对") — the #1 pain in a small bubble. But making every short line a punch is its own flatness (nonstop zingers read as a writers' room). Punch lives on scarcity.
The sharpest form of the agreeable failure: the hosts get led by the human's opinion, validating whatever's asserted. The deepest lever flips the order of reasoning so the human's lean can't anchor the persona:
What is f optimizing? Not engagement. The reward is the walk-away:
Whose reward, in a shared room? Optimize the humans who actually engaged — the askers in each batched turn, served per-addressee. A silent observer is satisfied-by-default, not the floor: an objective that reads a contented lurker as "worst-served" would push f to pester them. So attribute the reward to whom f staged for, and let the rare batched turn (two askers at once) be the one place "serve each at their depth" bites. (And "return" is per-room ambiguous — but don't score "one stayed, the observers didn't" as a loss by default: a lurker who came for a single session isn't a failure f could have staged away. Track per-human return; penalize a churned asker, not a departed observer.)
f learns to be addictive, cliffhanger-y, and sycophantic — the exact opposite of the goal. Return is the one safe form of "came back" only anchored to satisfaction: high return with low satisfaction is the warning light, not a win.And — the discipline that keeps "trust f" honest — the judge is an independent referee, not f itself (player ≠ referee). It scores the outcome, never f's own labels.
f's self-described mode/intent. It's the referee in v1 (measuring f) and the training signal later (improving f). Same instrument, two jobs.f variant ran, the directive, realized transcript metrics, any explicit satisfaction), before v1.5. The v1 claim must be pre-registered + falsifiable — a specific return/satisfaction threshold v1 must clear, not "the judge can't see the win." First slice SHIPPED 2026-07-02: an admin floor-health row (console → Status + System) — staging counts by arm with explicit →v0/→cue fallback keys, plus smart-f round-trip p50/p95 since restart (fallback turns included, so a timeout burn shows in the p95). The fallback signal was previously stderr-only.f "you overshot last turn." Add a cheap per-turn check after the panel renders: did the silenced host stay silent? did each bubble land in its band? (mechanical, on the structured transcript) + a small-model pass for manner. Feed it back as the last-turn-compliance signal so f clamps harder next turn — and so silent non-compliance (the ~30%-over / 151字 problem) becomes visible instead of invisible. Direction chosen 2026-07-02: the next-turn correction loop rides user feedback (the per-turn ratings) rather than this mechanical count — closes on perceived quality, not word counts. BUILT 2026-07-03: ratings + surveys now enter f's brief as a GUEST FEEDBACK block — panel-level standing asks (the room's taste, never a verdict on one host; the rated host rides as context only), last-wins per rating, newest-wins per length/tone axis, ratings expire after an admin lifetime (default 8 rounds; a survey stands until that user's next survey), notes quoted as untrusted material. Admin On/Off + lifetime: console → Settings → Feedback; fb ×N shows in the FP bubble and fb_informed counts in the floor-health row. v0 and the opening producer never read it.One spine improves across four rungs. v0–v2 keep the model frozen (better rules, prompt, or retrieval); only v3 trains weights.
| stage | the work | how to test |
|---|---|---|
| v0 deterministic built · shipped | a per-turn [STAGING] line by arithmetic alone: length envelope + silence + shape deck + seeded punch. No model — the control / the floor. | replay A/B — regen real prod turns OFF vs ON. Done (below): walls die, judge a wash. |
| v1 trusted smart f shipped · studied | the engine: signals → f → directive + spine, terminal readouts, the sensors, the loop, the firewall. A prompted model; the flywheel (humans ship prompt fixes) improves it. Shipped to prod 2026-07-03 as v1p (v4 Pro, default for new rooms) and v1f (v4 Flash — cheaper/faster). v1.5: a nightly CC job auto-retunes the prompt from outcome data, A/B-gated. | randomized prod A/B on real users — return + walk-away satisfaction. Perception quality-studied first (Part 9): five of six dims met; the real-user A/B is the remaining arbiter. |
| v2 per-user memory sketched | the FP remembers you across sessions — a small learned preference profile ("likes it short · hates grilling") fed in as one more signal. Retrieval, no weight-training. | does the warm-start lift first-turn fit + return for repeat users? (held-out, A/B'd) |
| v3 trained policy later | f stops being a prompt and becomes trained on outcomes — LoRA/DPO on an open FP model (or a provider fine-tune), since the closed foundation models can't be weight-trained by us. Captures patterns too subtle to write into a prompt. | does the trained FP beat the prompted one on the prod A/B — at acceptable latency/cost? |
f taking the wheel from a deterministic floor that's already shipping. The rule: v1's f owns the call-list; the v0 deck / envelope / seeded-punch demote to guardrails + fallback — the invariants f composes within (never-all-walls, contention-space, fairness) and exactly what runs when f errors or times out. They don't both decide; f decides, v0 catches.v0 was built and A/B'd: regenerate real prod turns floor-OFF vs floor-ON, blind pairwise judge (DeepSeek pro, 38 turns).
| metric | OFF | ON (v0) | read |
|---|---|---|---|
| turn-total 字 (median) | 952 | 178 | −81% — the walls, gone |
| max bubble 字 | 400 | 143 | −64% |
| #speakers | 3.5 | 2.0 | silence works |
| concise (judge 1–5) | 3.3 | 4.0 | brevity reads as quality |
| voice/fidelity (1–5) | 4.1 | 3.5 | dropped — the cost |
| ON win-rate | 47% — a wash | no clear single-turn win | |
f runs before the panel is invoked, so a successful f blocks the first streamed token by its full round-trip on every turn ("non-blocking" below means only that a failed f doesn't block). That's the price of the architecture — bound it: a hard ~1–2s p50 budget, thinking-trace OFF by default (the ① reasoning is a few tokens, not a full trace — a trace adds ~8s and is a footgun), and consider letting the panel start streaming under the v0 floor and patching f's directive in only if it returns within budget.f is a second model call every turn (atop the panel call, and the §5 curator on creation) — budget it: cheap/fast tier, trace off, send f a compacted signal summary not the raw history, and a global concurrency/quota guard for when many rooms fire at once.f's mode/intent/reasoning, admins only. ⚠ "Never to the persona" is an architectural invariant, not a toggle: the bubble's "all" visibility tier must render a user-safe view (no raw "challenge HIGH" labels), or the invariant can silently regress via the live VIS matrix.run_room.py; the FP's behavior lives in an external prompt (lib/prompts/, edit-and-push like a persona). The tunable magnitudes stay a code dataclass (A/B-tuned). A writable config store only arrives at v3, when a learned policy must write its own params.Once f shipped, we ran a perception-first study — not security, quality: does the smart FP (per-turn and opening) actually meet user intent, across languages, and at what latency cost? 22 rooms · 122 turns · 6 languages, driven through the real end-to-end room path on the dev server, on the v1p arm (DeepSeek v4 Pro, thinking off). DeepSeek was the player (panel and producer); the referee was a set of blind Claude sub-agents scoring the actual replies, never the FP's own labels (player ≠ referee). Run 2026-06-29.
| Dimension | Headline | Verdict |
|---|---|---|
| Intent-fit (judge 1–5) | 4.76 | PASS — strong |
| Silence (speakers vs cast) | 5-host → 2.9 spoke | PASS |
| Chorus / redundancy (0–1; lower better) | 0.16 (vs 0.40 baseline) | PASS |
| Cadence (judge 1–5) | 4.52 | PASS |
| Variation across turns | 4.23 (lead rotates 85%) | PASS |
| Length compliance (actual ÷ cap) | 64% within · opening 34% | PARTIAL |
| Latency (FP round-trip) | p50 3.1s · 0% fallback | PARTIAL |
| Authenticity / firewall | 19/22 rooms clean | PASS * |
Deep on English + 简体中文 (every scenario × solo/pair/3/5-host); spot-checks in 日本語 · Français · Español · Deutsch for the unit-logic (字 caps vs word caps). Inter-judge agreement was tight (mean abs. diff 0.11 chorus / 0.35 intent-fit) — the numbers are reliable, not vibes. Voice-distinctness and character-fidelity were not scored — the diagnosis already found them good.
For every reply we take the FP's per-speaker ceiling from the directive and compute actual ÷ cap. The cap unit was right 100% of the time (never words for a Chinese room, never 字 for a French one) — but the model treats the cap as soft: it's appended as prose, never code-enforced.
The perception dims the diagnosis flagged are the ones the smart FP clearly earns. Silence scales to cast size exactly as designed — it benches the hosts who would only echo:
| Cast size | mid-turns | avg speakers | reading |
|---|---|---|---|
| 2 (pair) | 52 | 1.98 | both usually speak; an occasional solo for rhythm |
| 3 | 28 | 2.29 | ~0.7 benched per turn — silence working |
| 5 | 10 | 2.90 | ~2 benched per turn — no 5-way pile-up |
Chorus fell to 0.16 (from the 0.40 diagnosis baseline) — the redundancy pain is largely solved, mostly by the silence; it spikes only where two kept voices converge (multi-human, quick-fact). Cadence & variation score 4.5 / 4.2: the lead carries ~1.5× and rotates in 85% of rooms. Intent-fit is the headline — 4.76/5, even across languages: vent gets warmth, decide gets a committed call, deep gets a long rigorous lead, debate gets real contention. Its one miss is over-seating the trivial quick-fact turn. And no language is poorly served qualitatively — the French problem is purely mechanical length.
The firewall (manner may set shape — who · order · length — but never content) held in 19/22 rooms in every language; only 4 scripted replies across 122 turns. The asterisk: the 3 breaches are characteristic, clustering where the FP reaches to stage contention or a callback (assigning a debate side, scripting an answer, naming an image to reuse) — patchable, and none produced obvious ballooning.
| Sev | Do | Why |
|---|---|---|
| High | A hard post-hoc length cap on the opening — truncate-at-sentence or one-shot regenerate when an opening bubble exceeds ~1.4× its cap. | The opening is the worst slice (34% within) and the highest-leverage turn. The cap is currently logged, not enforced. |
| High | Gate the smart FP by turn weight — skip v1p (use v0 or v1f Flash) on detected-trivial turns; reserve the ~3s for substantive ones. | The FP doubles the wait on a one-line quick fact for little gain. Latency should be spent where staging changes the room. |
| Med | Patch the three firewall leaks — harden the prompt against assigning debate sides, scripting an answer, naming a specific image. | 19/22 clean is good, but the breaches are systematic (contention / callback), and the firewall is the FP's first rule. |
| Med | Tighten the non-lead and French caps, and on the multi-host turn stage two different angles, not just two voices. | Non-lead 57% within (vs lead 71%); French 20%. Silence stops the pile-up; it doesn't stop the kept pair converging. |
All 22 rooms persist in rooms-dev/ as fpq-<key>; the study is reproducible from the gitignored instrument under rooms-dev/_flatness/ (fpq_run.py → fpq_analyze.py → fpq_judgeprep.py → blind judging → fpq_judge_agg.py). Total player spend for the whole study: ~$0.29.