Dialogue · Design notes · Conversation quality
← Design notes

Conversation quality — a production diagnosis

The empirical half of roadmap §1 (the floor producer). §1 asserted that the megaprompt "averages the room." This is the measurement: all 80 production rooms pulled from the live box on 2026-06-23, the 39 real multi-host conversations scored on a rubric built from actual user feedback — what users said is good (voice, fidelity — not re-scored here) and what they said is bad. Every finding is triangulated three ways: a mechanical pass (pure arithmetic), an LLM judge reading every transcript, and the users' own in-room complaints. The diagnosis points directly at what the floor producer must fix.

One line: the panel is too long, too agreeable, and the conflict dial barely works. 40% of non-lead bubbles add no new angle, the median host reply is ~314 words, and the extreme temperament delivers about a third of the friction it promises. None of this is a voice or fidelity problem — those are good. It is a staging problem.

0.40
redundancy — share of non-lead bubbles that add no new angle judge
314 w
median reply — p90 is 482 words mech
2.0 /3
dial gapextreme promised 3, delivered 1.0 judge
60×
one gesture — 张小龙's cigarette, in a single room mech

How it was measured

The rubric was cut down to only the four things users complained about. Users said voice distinctness and character fidelity are already good, so those are not scored here — every judge token goes at a real pain. The instrument never sees the live model; it reads the recorded event logs. Human turns are anonymized (Human-A/B) before anything reads them; raw transcripts stay off the repo.

80 prod rooms pulled from box render + anonymize → 39 multi-host mechanical pass length · silence · gestures 12 LLM judges read every transcript aggregate 39 scorecards this page
The pipeline. Mechanical metrics are free and objective; the judge adds the reading the arithmetic can't do (is this length earned? is this a paraphrase?). The two are cross-checked against what users actually typed in the room.

The rubric — four pains, in users' words

Pain (user's words)What we scoreHow
"One host just paraphrases the last — add an angle or stay silent"Each non-lead line → new-angle / redundant-paraphrase / filler → redundancy ratejudge
"They speak too long — big bubbles. Quality is value, not length"padded rate + reply-length distribution + short-line ("concur") ratejudge mech
"The dials don't work — at extreme I didn't feel conflict"perceived conflict 0–3 vs the dial's promise → delivery gapjudge
"Repeated body language — 张小龙 always does his 吸一口烟"same physical gesture repeated by the same host → max repeatsmech judge
judge biasLLM judges over-reward length, politeness, agreeableness, and confidence — which are the exact failures here. The judge was explicitly instructed to invert that: a polished line that merely restates a prior host is a defect, not a nicety. Confidence comes not from the judge alone but from its agreement with the mechanical numbers and with users' own words — see each finding.

Finding 1 — length is the pervasive tax

The single most consistent defect across all 39 rooms. The median host reply is ~314 words (p90 482); per-room medians run from 83 up to 714 words. A comfortable chat bubble is one to three sentences — call it ~40 words. Hosts routinely answer at 7–18× that. And the response gate's "concur in one short line" tier is essentially dead: the short-line rate is ~1%. Nobody nods; everyone gives a speech.

0100200 300400500 600700 words comfortable bubble ≈ 40 w median 292 p25 171 p75 450 longest room median: 714 w
Per-room median host reply length. The coral box is the middle half of all rooms; even the 25th-percentile room runs ~4× a comfortable bubble. Mechanical — word counts don't lie.

The judge confirms it isn't just long but padded: 36% of all host lines are rated ≥40% cuttable with no loss of substance (preamble, restatement, hedging). 14 of 39 rooms are above 40% padded. And it is not a big-cast artifact — two of the worst-padded rooms are 2-host coaching rooms (padded 0.57 and 0.51), where a small panel simply monologues. Users felt it and said so in the room:

"You're all talking far, far too long — please keep it short and plain."— a guest, mid-conversation (paraphrased / translated)
"Can each of you ask just one, most-essential question?"— a guest, after a wall of multi-paragraph replies

Finding 2 — redundancy by convergence

When more than one host speaks in a turn, 40% of the later contributions add no new angle — they paraphrase or re-affirm what was already said. The signature move the judge saw again and again: "X said it well — let me add…" followed by the same point in new words. 23 of 39 rooms sit at or above 40% redundancy.

non-lead bubbles 60% — new angle 40% — redundant ↑ the 40% the user calls "a waste of time to read"
The value split of non-lead contributions across 39 rooms. Judge-rated; corroborated by the mechanical near-zero short-line rate — redundancy arrives as full speeches, not nods.

It scales with cast size and "everyone speaks" turns, but it is not only a chorus effect: high-all-speak rooms average 0.42 redundancy vs 0.37 for the rest — a real but modest gap. The other half is cross-turn self-redundancy: a single host re-running their own signature point every turn (one host re-cited the same "40% survey" rule five times; another re-derived the same pricing formula turn after turn). The good news is that it is clearly solvable — the best rooms scored 0.00–0.19 because each host held a genuinely distinct lens:

"You're all saying the same thing."— a guest, in a 4-host room (paraphrased)

Finding 3 — the conflict dial barely works the sharpest finding

This is the user's most specific complaint, and the data is blunt. The extreme temperament promises a sharp, sustained clash (level 3). Across 17 production rooms it delivered a perceived conflict of 1.0 — "polite divergence." 5 of those 17 rooms scored a flat 0 — total mutual agreement, at the maximum setting. The natural dial, by contrast, is well-calibrated (gap 0.33).

0123 conflict level → 0 none · 1 polite divergence · 2 real friction · 3 sharp clash promised ~.75 0.67 natural (n=18) · gap 0.33 ✓ promised 3 1.0 gap +2.0 extreme (n=17) · 13 of 17 under-deliver by ≥2
The dial's promise vs what landed. Turning temperament to maximum moves perceived conflict from 0.67 to just 1.0 — a rounding error for a max setting. Judge-rated against each room's own dial.
why it under-deliversThe judges converged on a mechanism the dial text never anticipated: the friction that exists is host→human, not host↔host. At extreme, hosts deliver longer, more pointed speeches at the guest — but with each other they "converge and openly agree" (one host literally said "my answer lands in basically the same place as his"). The panel interrogates the human well; it will not make the panelists contest each other. That is exactly why the user "didn't get the feeling" — the heat went the wrong direction.

An honest correction: an earlier pass on a handful of dev rooms suggested the dials made rooms more choral. The full production sample (17 extreme rooms vs 3 in dev) overturns that — on real data the dial story is under-delivered conflict, not changed participation. Pulling everything was the right call. (probing and combative each have only n=2 in production — too few to bucket; the natural-vs-extreme contrast carries the finding.)

Finding 4 — a few signature gestures, worn to death

Pain 4 is real but local. Most rooms are clean; gesture-repetition concentrates in a handful of personas — and where it appears, it is egregious. The worst single room is a textbook case the user named almost exactly:

张小龙 · cigarette 60× 张一鸣 · push glasses 46× 李安 · lean forward 24× 李安 · light laugh 16× 刘慈欣 · long pause 11× cute once; grating by ~5× →
One extreme/theatrical room (3 hosts, 66 turns). 张小龙 lit, drew on, tapped, and stubbed out a cigarette sixty times; 张一鸣 pushed his glasses up 46 times. Mechanically counted from stage directions; judge-confirmed across prose variants.

This is different in kind from the first three. It is not a staging problem the room can fix per-turn — it is a persona-authoring problem: a signature gesture licensed in profile.md gets over-deployed at runtime. The fix lives in the persona kit (single-source / cap the gesture, the way the authoring kit already single-sources a signature line), not in the floor producer.

The numbers, by dial

dialroomsredundancypaddedconflict /3gapsilenceall-speakmed words
natural180.340.310.670.330.150.73290
extreme170.420.411.002.000.290.50346
probing *20.720.290.501.500.270.50227
combative *20.420.320.501.500.080.86338

* probing / combative n=2 each — shown for completeness, too few to read into. Overall (39 rooms): redundancy 0.40 · padded 0.36 · perceived conflict 0.80/3 · median reply 314 w · p90 482 w · short-line rate ~0.01.

What this means for the floor producer

Three of the four pains are the same root problem from different sides: one mind voices the whole panel, and that mind always has something for everyone to say, at length, in agreement. It cannot self-correct, because the writer and the staging-decider are the same context. That is precisely the case §1 makes for an exogenous per-turn floor producer — and this diagnosis tells it exactly which levers to pull.

1 · redundancy (0.40) 2 · length (~314 w) 3 · dial gap (+2.0) 4 · gesture repeats (60×) FLOOR PRODUCER — per-turn levers · silence — "you add nothing here, sit out" · length budget — "one line each" / "let one run" · contention — assign a real host↔host split staging only — never content, never positions persona authoring — cap the signature gesture 3 of 4 pains are staging the writer can't self-fix · the 4th is a profile fix
The diagnosis as a build spec. Note the new requirement the data surfaced: the producer must be able to assign genuine host↔host contention — the static dial provably can't, and the heat otherwise points only at the human.

What to trust, and what not to

Diagnosis run 2026-06-23 · 80 production rooms → 39 multi-host conversations · mechanical pass + 39 single-room LLM-judge scorecards (12 subagents) · instrument in rooms-dev/_flatness/ (gitignored; raw transcripts never committed). Leads into the floor-producer design (roadmap §1).