The empirical half of roadmap §1 (the floor producer). §1 asserted that the megaprompt "averages the room." This is the measurement: all 80 production rooms pulled from the live box on 2026-06-23, the 39 real multi-host conversations scored on a rubric built from actual user feedback — what users said is good (voice, fidelity — not re-scored here) and what they said is bad. Every finding is triangulated three ways: a mechanical pass (pure arithmetic), an LLM judge reading every transcript, and the users' own in-room complaints. The diagnosis points directly at what the floor producer must fix.
extreme temperament delivers about a third of the friction it promises. None of this is a voice or fidelity problem — those are good. It is a staging problem.extreme promised 3, delivered 1.0 judgeThe rubric was cut down to only the four things users complained about. Users said voice distinctness and character fidelity are already good, so those are not scored here — every judge token goes at a real pain. The instrument never sees the live model; it reads the recorded event logs. Human turns are anonymized (Human-A/B) before anything reads them; raw transcripts stay off the repo.
| Pain (user's words) | What we score | How |
|---|---|---|
| "One host just paraphrases the last — add an angle or stay silent" | Each non-lead line → new-angle / redundant-paraphrase / filler → redundancy rate | judge |
| "They speak too long — big bubbles. Quality is value, not length" | padded rate + reply-length distribution + short-line ("concur") rate | judge mech |
"The dials don't work — at extreme I didn't feel conflict" | perceived conflict 0–3 vs the dial's promise → delivery gap | judge |
| "Repeated body language — 张小龙 always does his 吸一口烟" | same physical gesture repeated by the same host → max repeats | mech judge |
The single most consistent defect across all 39 rooms. The median host reply is ~314 words (p90 482); per-room medians run from 83 up to 714 words. A comfortable chat bubble is one to three sentences — call it ~40 words. Hosts routinely answer at 7–18× that. And the response gate's "concur in one short line" tier is essentially dead: the short-line rate is ~1%. Nobody nods; everyone gives a speech.
The judge confirms it isn't just long but padded: 36% of all host lines are rated ≥40% cuttable with no loss of substance (preamble, restatement, hedging). 14 of 39 rooms are above 40% padded. And it is not a big-cast artifact — two of the worst-padded rooms are 2-host coaching rooms (padded 0.57 and 0.51), where a small panel simply monologues. Users felt it and said so in the room:
When more than one host speaks in a turn, 40% of the later contributions add no new angle — they paraphrase or re-affirm what was already said. The signature move the judge saw again and again: "X said it well — let me add…" followed by the same point in new words. 23 of 39 rooms sit at or above 40% redundancy.
It scales with cast size and "everyone speaks" turns, but it is not only a chorus effect: high-all-speak rooms average 0.42 redundancy vs 0.37 for the rest — a real but modest gap. The other half is cross-turn self-redundancy: a single host re-running their own signature point every turn (one host re-cited the same "40% survey" rule five times; another re-derived the same pricing formula turn after turn). The good news is that it is clearly solvable — the best rooms scored 0.00–0.19 because each host held a genuinely distinct lens:
This is the user's most specific complaint, and the data is blunt. The extreme temperament promises a sharp, sustained clash (level 3). Across 17 production rooms it delivered a perceived conflict of 1.0 — "polite divergence." 5 of those 17 rooms scored a flat 0 — total mutual agreement, at the maximum setting. The natural dial, by contrast, is well-calibrated (gap 0.33).
extreme, hosts deliver longer, more pointed speeches at the guest — but with each other they "converge and openly agree" (one host literally said "my answer lands in basically the same place as his"). The panel interrogates the human well; it will not make the panelists contest each other. That is exactly why the user "didn't get the feeling" — the heat went the wrong direction.An honest correction: an earlier pass on a handful of dev rooms suggested the dials made rooms more choral. The full production sample (17 extreme rooms vs 3 in dev) overturns that — on real data the dial story is under-delivered conflict, not changed participation. Pulling everything was the right call. (probing and combative each have only n=2 in production — too few to bucket; the natural-vs-extreme contrast carries the finding.)
Pain 4 is real but local. Most rooms are clean; gesture-repetition concentrates in a handful of personas — and where it appears, it is egregious. The worst single room is a textbook case the user named almost exactly:
extreme/theatrical room (3 hosts, 66 turns). 张小龙 lit, drew on, tapped, and stubbed out a cigarette sixty times; 张一鸣 pushed his glasses up 46 times. Mechanically counted from stage directions; judge-confirmed across prose variants.This is different in kind from the first three. It is not a staging problem the room can fix per-turn — it is a persona-authoring problem: a signature gesture licensed in profile.md gets over-deployed at runtime. The fix lives in the persona kit (single-source / cap the gesture, the way the authoring kit already single-sources a signature line), not in the floor producer.
| dial | rooms | redundancy | padded | conflict /3 | gap | silence | all-speak | med words |
|---|---|---|---|---|---|---|---|---|
| natural | 18 | 0.34 | 0.31 | 0.67 | 0.33 | 0.15 | 0.73 | 290 |
| extreme | 17 | 0.42 | 0.41 | 1.00 | 2.00 | 0.29 | 0.50 | 346 |
| probing * | 2 | 0.72 | 0.29 | 0.50 | 1.50 | 0.27 | 0.50 | 227 |
| combative * | 2 | 0.42 | 0.32 | 0.50 | 1.50 | 0.08 | 0.86 | 338 |
* probing / combative n=2 each — shown for completeness, too few to read into. Overall (39 rooms): redundancy 0.40 · padded 0.36 · perceived conflict 0.80/3 · median reply 314 w · p90 482 w · short-line rate ~0.01.
Three of the four pains are the same root problem from different sides: one mind voices the whole panel, and that mind always has something for everyone to say, at length, in agreement. It cannot self-correct, because the writer and the staging-decider are the same context. That is precisely the case §1 makes for an exogenous per-turn floor producer — and this diagnosis tells it exactly which levers to pull.
extreme fails because the writer routes heat at the human, not between hosts. The producer needs to name who disagrees with whom this turn — staging the split, never scripting the position (the dumb-pipe contract holds).extreme room against extreme's own promise), not across confounded groups.extreme (17 rooms) and Chinese-language sessions, light on probing/combative. That is a feature: it reflects real usage, not a lab.rooms-dev/_flatness/ (gitignored; raw transcripts never committed). Leads into the floor-producer design (roadmap §1).