g→s→f function stack. It was superseded 2026-06-26 by the signals → f → staging model (the producer is one trusted function; "mode" is a readout, not a gate). Current page → Floor producer — the design. Kept here for reference only.The build spec for roadmap §1, written off the back of the production diagnosis. The diagnosis proved where conversation quality breaks; this is the machine that fixes it. Status: design settled. v0 the deterministic baseline is built, shipped (new chats default-on) and A/B-measured; v1 the smart producer is designed below; v2 is sketched. Reorganized 2026-06-26 — start with the Contents, then read top to bottom.
We pulled all 80 production rooms and scored the 39 real conversations (full report). Four findings drive every design choice below:
extreme delivers ~⅓ of the friction it promises — and the heat that does exist goes host→human, not host↔host.The load-bearing conclusion: three of the four are staging, not voice. Voice and character fidelity are good (users said so). The single writer just stages the room badly — and can't fix itself. That is exactly the gap an exogenous producer fills.
Everything below is one of three things: an input (what the producer reads), the brain (what it infers), or an actuator (the staging it emits) — all bent toward a single objective (what the user walks away with).
The panel already deliberates — the backstage <confer> is the one mind working out what the hosts say. So: does an exogenous producer fight that? It doesn't. They sit at different layers — the producer is upstream input to the confer, never a rival to it.
M·M slot); the confer decides the substance — the door is opened, the confer decides if anyone walks through.Five staging dimensions. Each is presented the same way: the mechanism (how it works), then how v1 turns it from a blanket rule into an intent-driven one. The brain (Part 4) is the conductor; these are the instruments.
The diagnosis defect was never "too long" globally — it was uniform ~300-word walls, no variance, no short lines. The research (below) points the same way: length shouldn't be a target the model fills, it should be a computed envelope — a band the reply lands inside, that slides with the user and wobbles each turn. Four knobs size one bubble:
| knob | what it does | why |
|---|---|---|
| floor | the minimum — a lead must carry a claim and its reason | stops the curt brush-off (the dismissed pain) |
| ceiling | the maximum — the anti-wall cap | the one fence the verbosity-bias work says must stay |
| entrainment | the whole band slides up/down with the user's message length | matching builds rapport; under-matching reads as dismissive |
| jitter | a small seeded wobble so no two turns share a target | breaks the metronome (habituation); seeded ⇒ reproducible & A/B-able |
The CJK build taught the hard rule: a model cannot count its own 字 — it ran ~2.3× over a 字 budget but only ~1.1× over a word budget, and it hits a sentence count cleanly in every language we tested. So the unit wears two hats, and we'd been conflating them:
| finding | source | what it means here |
|---|---|---|
| Repetition is the single biggest drag on conversation quality — cutting it lifts every human score | See et al., NAACL 2019 | the "repeated pattern" worry is the #1 documented failure (they measured content repetition; we extend it to rhythm) |
| Longer ≠ better — optimal length is context-dependent; extra length adds repetitiveness without quality | CHI 2024 | length must follow the turn's job, not a fixed amount |
| People entrain on utterance length; matching builds rapport and lowers latency | entrainment, multi-party | backs the mirror — keep it, on both ends of the band |
| Turn length follows the communicative action, and listeners project the endpoint | Liddicoat 2004 (TCU) | length should track intent — the model is the only part that knows it |
| Predictable repetition habituates faster / lowers arousal — but too much novelty is also bad; moderate is optimal | PLOS One 2020 | jitter yes, but moderate & structured — not pure randomness |
| LLMs & reward models favor longer (GPT-4 picks the longer answer >90% of the time); "length-gaming" | length bias, 2024 | so the length call must be the producer's, not the writer's — but the height is intent-set, not fixed (below) |
In v1 the fence is two-level. A turn-total cap and a per-bubble cap — because a turn-total alone would let the model pour the whole budget into one giant bubble (a wall again). Inside both fences the LLM distributes the length by substance.
A fixed cap is wrong: users often want length — an analysis, a deliverable, or an explicit "go longer." A constant ceiling makes those impossible. So in v1 the ceiling is a function of intent/mode, set by the producer. The anti-wall guarantee survives because the protection was never the low number — it's that an exogenous producer owns the length, not the verbose writer. A high ceiling the producer grants (because the turn earns it) is safe; an unbounded one the writer grants itself is not.
| intent / mode | bubble length |
|---|---|
| vent / emotional | minimal — presence, not paragraphs |
| social / casual | tight (one-liners) |
| quick fact Q&A | tight–medium |
| exploratory / analysis | generous — substance allowed |
| explicit "go longer / shorter" | the user's dial wins — overrides the mode |
Two length regimes, not one. A deliverable — "write me the doc / the code / the 2000-word analysis" — shouldn't be a chat bubble at all; it's an artifact, and the room already routes those to the wildcard pane, unbounded by design. So the bubble stays a tight hand-off ("here's the analysis →") while the product runs as long as it needs to. The producer's output actuator already picks prose-vs-artifact; the length policy just follows it.
max_tokens runaway guard — plumbing, not design.Length sizes one bubble; shape arranges the long and short bubbles across a turn. v0 today always deals the same shape — L·S·S (one host leads, the rest get a line). That single rhythm is exactly what users habituate to — and it's structurally why the conflict dial under-delivers: if one host always leads and the others footnote, two hosts can never both make a case. So the producer carries a small deck of coherent shapes and deals a different one each turn.
M·M splits the weight (two real cases); S·S·S is a rapid-fire beat with no lead. L·L·L — three walls — is the thing we're killing, so it is never in the deck. This is the v0 deck — a fixed, seeded draw. v1 doesn't pick a card from it: it composes the shape per host within these same coherence rules, so it can also produce graded shapes like L·M·S the deck lacks.| shape | feels like | best for |
|---|---|---|
L·S·S | answer + two footnotes | a direct question (today's default) |
S·S·L | build to a payoff | the last voice lands it (authority / contrarian) |
S·L·S | setup · develop · button | a hand-off; the middle is the meat |
M·M | two-way, nobody dominates | a disagreement — both get room |
S·S·S | rapid fire | a light / social beat |
It's a weighted, no-repeat deck — never a shuffle. Three rules keep each shape coherent:
S·S·L the short ones are reactions/setups ("Interesting—" / "I'd push back—"), labelled by role so they scaffold the lead — not orphaned claims.L·S·S; a disagreement → M·M; banter → S·S·S.The deck is keyed by speakers-this-turn, not cast size: the silence lever caps speakers at ~3 even in an 8-host room, so the live keys are 1·2·3 (4 only on a 4-host full house). Participation itself also varies — a 3–4-host room runs 2-up by default, occasionally solo, occasionally a deliberate full house (the deck keeps that to one lead + short lines, never a wall pile-up).
| cast | speakers / turn | shapes it can deal |
|---|---|---|
| 1 | 1 | L — one voice, no shape lever (length envelope only) |
| 2 | 2 | L·S · S·L · M·M |
| 3–4 | 2 default · 1 solo · all on a full house | the 2-up shapes, plus the 3-/4-wide deck (incl. M·M·S) when full |
| 5–8 | 3 (the rest silent, rotating) | L·S·S · S·S·L · S·L·S · M·M·S · S·S·S |
The contention shape scales with the turn: M·M in a 2-speaker turn, M·M·S in a 3-speaker turn — so two hosts can contend at any room size. The one case with no shape lever is a 1-host room: a single voice, so only the length envelope applies.
Two levers ride on the deck. Silence — who doesn't speak — is the cheapest quality move there is (sitting out never needs fabrication), and it is exactly the thing one mind won't do to itself. Contention — the M·M shape — is the structural fix for the broken conflict dial.
M·M is the shape that gives two hosts the floor to actually contend — the host↔host friction the diagnosis found missing (the extreme dial scored conflict 1/3 because friction went host→human, never host↔host). So shape variance isn't only anti-boredom — it's the structural prerequisite for the contention lever. Weight M·M up when a real split is in the air.The firewall holds here especially hard: the producer allocates the space for a split (the M·M slot) but never assigns positions — it can't, it sees only the public cards, not the personas' interiors. It opens the door; the confer decides if anyone walks through. And if the panel genuinely agrees, forcing a fight is manufactured conflict (a failure) — so contention degrades gracefully to distinct reasoning or silence. See the worked example for what "surface but don't assign" looks like in practice.
The short slot was the flattest thing in the room. Told "one line", the model returned filler — "Exactly", "说得对,我补充" — the diagnosis's #1 pain (redundancy 0.40) in a small bubble. But the reflex fix — make every short line a punch — is its own flatness: nonstop zingers read as a writers' room, manufacture wit that isn't there, and flatten distinct voices into one. Punch lives on scarcity.
So the S slot is fixed in two layers: raise the floor for all (every short bubble is now "a tight line — a genuine quick point, never empty agreement" — kills filler without forcing a joke), and reserve punch for the rare earned moment (at most ONE short slot per turn, only ~1-in-3 turns, seeded, is promoted to a craft directive). What that directive should be came from a one-liner study on DeepSeek (EN + ZH):
| the study found | so |
|---|---|
| a positive craft directive (reversal / concrete image / jab) beat "be brief" reliably | tell the model what punch is, not just "one line" |
| few-shot examples HURT (worst of all) — they bred forced, generic mimicry | keep it abstract, never exemplar-based |
| punch is partly model-gated — same prompt, strong model 4.2 vs cheap 2.7 (of 5) | cheap-model rooms have a punch ceiling; for them, a best-of-N on short slots buys what the model won't volunteer |
| ZH wants more room than EN (对仗 / 反转 / imagery) | a language-aware punch directive (short_punch_zh) |
| in the panel, it lands a punch then keeps explaining — a 151-字 "one-liner" shipped to prod (first sentence great, four more after it) | the directive must say STOP after the blade — "no explaining or expanding"; forbidding setup (the front) wasn't enough, the back needs a hard stop too |
| and the stop works best as a structural cut-down, not a 字 count — a panel A/B: "≤20 字" was overshot to ~30 (the model can't count), but "delete the extra down to that one line" held; dropping the number entirely kept the best punch | the shipped directive is structure + craft + a cut-down corrective, no 字/word count — the two-rulers rule again |
punch_every + short_punch_en/zh knobs). v1 the producer beat-gates it — yes on a sharp or playful turn, never on a tender or grieving one (a zinger there is the "inauthentic / lectured" failure). Same lesson as length and conflict: a good thing deployed variably is a lever; slammed on as a constant, it's a new flatness. The general principle this proved — how to phrase any want so DeepSeek obeys — graduates to the brain (what runs it).The sharpest form of the Pandered failure: the hosts get led by the human's opinion — validating whatever's asserted, as if they have no ground of their own. Two pulls stack: the base model is RLHF-trained to be agreeable, and the single-mind megaprompt makes every persona tilt toward the human at once (choral caving).
extreme dial risks). The aim is the green middle.The principle: a reply should track the persona's view, not the human's. Agreement is fine when genuine; sycophancy is caving against one's own ground for no reason but social pressure. The deepest lever flips the order of reasoning so the human's lean can't anchor the persona:
The full fix is layered, root to surface:
| lever | what it does | layer |
|---|---|---|
| Give them ground | documented positions + red lines (what this person won't concede); extrapolate a stance from values, don't default agreeable | persona kit |
| Commit before cave | form my read first, then weigh the human's — the order flip above (anti-anchoring) | structural · strongest |
| Position persistence | track what a host committed to; a reversal needs an explicit reason, not silent drift across turns | structural |
| Detect & allocate | the producer senses caving (agreement-rate, validation density, drift toward the human) → cranks the Challenge dial + opens a "stress-test from your ground" slot — space, never stance | producer |
| Frame | the human is a thinking partner to engage, not an authority to validate; "update only for a real reason" | prompt |
The levers are what the producer can do; the brain is how it chooses. It reads the room into a mode, maps that mode to a staging across all five levers (intent → staging), exposes a sliver of control as dials, optimizes an objective, runs as a small function stack, and is voiced by a specific kind of model.
Step back from the word "mode" for a second. The producer's real job is one function: a high-dimensional signal vector in — your engagement, your explicit asks, how invested you are, the substance, the room's structure (solo or panel), the phase — and a lean lever vector out: length · shape · silence · challenge · punch · closure. That lever vector is the staging. No "mode" sits in that path. The deepest way to say it:
So what is the named-mode vocabulary for? Not for computing the staging — for reading it. A handful of named regions (Vent · Decide · Learn…) give the judge a unit to score, an admin a handle to debug, and a brand-new room a sane cold-start prior before any signal has arrived. But the moment signals exist, they decide; the mode just describes. The 2-axis picture below is exactly that — the dashboard, a coarse legible readout of the manifold and the cold-start default, not the engine. (Earlier diagrams — the system-on-one-line, the function stack — draw mode as an inferred intermediate; read those as this same legibility view, the engine underneath being the signal→lever function.) Seven named regions, spanning the space:
other — it never mints a named mode mid-chat; other clusters are harvested offline into a new anchor + a judge axis. All seven anchors ship in v1, deliberately spanning both halves — the diverge side (Explore · Spar · Banter) as much as converge — so the producer has a validated home for every read, not just convergent ones.| anchor | where it sits | feels like | staging signature |
|---|---|---|---|
| Vent | emotional · converge | "I just need to be heard" | one warm voice, minimal length, no challenge, no pile-on |
| Learn | task · converge | "explain this to me" | the one relevant master solos, unhurried, exposition |
| Decide | task · converge | "A or B?" | M·M — surface the real tradeoff, then hold or synthesize |
| Build | task · converge | "make me the thing" | the artifact is the point → wildcard pane; tight bubble hand-off |
| Explore | task · diverge | "what are the angles?" | breadth — several short distinct takes, no forced landing |
| Spar | task/emotional · diverge | "stress-test me" | high challenge by request, adversarial or scenario |
| Banter | emotional · diverge | "let's riff" | short, snappy, one voice rotating, challenge off, punch welcome |
other, which we harvest into a new anchor later. Closed vocabulary for predictability + measurement; open composition for blend + novelty. The space grows from real rooms — we don't enumerate every conversation up front. The two clocks of novelty: live, the producer only blends known anchors or stages an unrecognized point from the levers and tags it other — it never mints a named mode mid-chat (an invented mode would be an untested policy on a real user, with no judge axis). A new anchor is born offline: other clusters get a deliberate staging signature and a judge axis, then an A/B, before joining the map. So novelty in staging is open and live; novelty in the vocabulary is human-gated.When other recurs it can mean two very different things, with very different costs — so they earn very different bars (a third bucket is just noise):
other. Case 1 (common): a recurring pattern lands at a known spot in the plane → harvest a new anchor. Case 2 (rare): the two axes can't separate it — chats at the same coordinates want different staging — the signal for a possible 3rd axis (e.g. stakes), taken only when no lever absorbs it.other turns out to be… | what it means | the response | bar |
|---|---|---|---|
| a recurring pattern at a known spot | the 2-axis map is right, just under-sampled | add an anchor (the offline flywheel) | low · common |
| won't localize — same coordinates, different staging | the two axes don't span what drives staging | lever first; a true 3rd axis only as a last resort | high · rare |
| a one-off | noise · the long tail | leave it other | a frequency gate |
altitude, closure and challenge out of the mode space as actuators. So when other won't localize, first ask which lever absorbs it. Add a genuine 3rd axis only when the new other can't be reconciled with the 2×2 at all — the factor is orthogonal (present at every task/emot × conv/div spot), re-sorts the staging at the same 2D point, and coordinates many levers at once (mode-like, not lever-like). A 3rd axis turns the plane into a 2×2×2 cube — anchors double, the judge burden doubles, the read gets harder — so the bar stays high. The likeliest real candidate is stakes / gravity (low↔high): a life-decision and a which-laptop decision share a 2D spot but want different care, pace, challenge and closure — too many levers for one knob — and it already leaks in as the acute-distress safety exception. other is the instrument: localizable clusters refine the map; un-localizable ones (high residual) are the alarm that the basis itself is incomplete.This is f, concretely: the four moves by which the producer reads the signals and sets the levers. Each is answered from the substance and the live signals — not by looking a mode up (we name the result a mode afterward, for the record). Same procedure every turn:
M·M; the agree-ers take short slots or go silent. Shape is a consequence of who has substance, not a card drawn blind.Averaged over many real turns, each named region has a typical staging — its centroid. The table below is those centroids — priors, not a function: the real signature is set per turn by the signals, and it varies within a mode (Carrie, below). Read it as "where this region usually lands," not "what this mode always does":
| mode | length | shape | silence | challenge | punch | closure |
|---|---|---|---|---|---|---|
| Vent | minimal | solo | most sit out | off | no | stay with it |
| Learn | generous | solo (the master) | others silent | low | no | synthesize |
| Decide | medium, capped | M·M | third silent | mid | rare | hold the split |
| Build | artifact (pane) | solo + hand-off | others silent | low | no | ship the thing |
| Explore | tight | S·S·S | rotate | low | some | diverge — hold |
| Spar | medium | M·M / grill | — | high | yes | press, then land |
| Banter | tight | S·S·S | rotate | off | welcome | none |
The same input, staged three ways by mode — the VC question as a worked example:
| if the turn reads as… | mode | the producer stages… |
|---|---|---|
| "Should I raise VC money, or bootstrap?" — a real decision with a live tension | Decide | M·M: two hosts each argue a side (control vs speed), third silent; medium length, capped; hold the split — don't smooth; punch off (it's a serious call) |
| "I'm exhausted — I can't tell if I'm about to make a huge mistake" | Vent | one warm voice; minimal length; no challenge, no pile-on; presence over analysis |
| "Draft my seed-round memo" | Build | the doc to the wildcard pane (unbounded); a tight bubble hand-off ("here's a first cut →"); no debate |
And the flip side — the same mode staged many ways, because the signature is per turn, not per mode. One real solo session (梁宁 · 自我探索, prod) — all in a single "Counsel" read:
| her turn (the signal) | the staging f produced |
|---|---|
| "陪我探索" — opening, vulnerable | ~999字, warm |
| "目前主要在接商单" — a flat update | ~336字 |
| "你不要安慰我,以面试官的角度客观评价" | ~1623字 · challenge HIGH |
| "乔布斯、张小龙是什么样的" — a teaching ask | ~1886字, exposition |
The dials are the only part users see. So the names say the effect, not the mechanism. Three clear nouns, each answering "what am I tweaking" — and each a different tension:
| Dial (user sees) | Poles | What it tweaks | Tension | From the diagnosis |
|---|---|---|---|---|
| Debate | Harmonious → Combative | how much the AIs argue with each other | panel ↔ panel | the broken one — needs the producer to surface a real split (never assign sides) |
| Personality | Subtle → Vivid | how strongly each AI plays its character | each AI = itself | unchanged (already good) |
| Challenge | Supportive → Grilling | how hard the AIs push you | panel → you | the easy win — the heat already flows here |
The tension column is the disambiguator: a tiny icon per dial (panel↔panel, the lone figure, panel→you) keeps Debate and Challenge from ever blurring. The diagnosis predicts the order of difficulty: Challenge ships almost for free (the panel already grills humans well), while Debate is the hard one the producer's contention lever exists to fix.
Mode vs dial — who wins. Three layers, and the rule that resolves them: the user's explicit dial always wins. The mode shapes the room and defaults the dials; it never overrides a dial the user has touched — because "stop coddling me, tell me straight" is a real and legitimate request a clamp would wrongly block.
It picks the staging that best serves walk-away satisfaction while dodging the failure modes, with the mode setting the weights:
Closed loop: your newest message is the reward signal on last turn's staging — "太长了" → length↓, you go deeper → depth↑, you @-mention X → rhythm steers to X, you disengage → the mode may be wrong.
"A good chat" is really the absence of a set of failures, and most come in opposing pairs — fix one end and you cause the other. The producer is balancing on several knife-edges at once, and the mode decides where to sit on each.
| ◄ over-correct | under-correct ► | the lever | measured? |
|---|---|---|---|
| Bored — long, redundant | Brushed off — curt | length budget + variance | yes (pains 1–2) |
| Intimidated — over-challenged, pile-on | Underwhelmed — shallow, generic | challenge + depth + #speakers | partial |
| Dismissed — real point ignored | Pandered — empty flattery | focus + genuine friction | no — new |
| Confused — divergence, no through-line | Flattened — one blended view | closure (synthesize vs hold) | no — new |
| Railroaded — panel steers, ignores you | Aimless — meanders | follow-the-user vs give a spine | no — new |
Plus four singletons: lectured (register too high) · creeped out / inauthentic (a host fakes feeling, or breaks the honest-anachronism contract) · whiplash (tone/length lurches; a host contradicts its earlier self) · no payoff (lots of talk, no takeaway — walk-away satisfaction itself).
A way to see the machine — what it reads, infers, and emits. It is a specification, not the implementation: the real producer is one smart-model call that does all of this holistically and fuzzily (weighing subtext, a disclosure buried in a casual line). But the functions earn their keep — they are the prompt's input list + output schema, they draw the code-vs-LLM line (the arithmetic is code; only mode/focus/contention need judgment), and each output is a judge axis.
Public room state only — never the private profiles.
| Input | What the producer sees |
|---|---|
| turn | the turn number |
| message | your latest message (and its length) |
| tail | the recent transcript |
| brief | the §5 "describe your situation" box |
| cast | the AIs and the humans in the room |
| host stats | per host — recency, word-share, @-mentions |
| your dials | the dials you've set explicitly (may be empty) |
focus and contention need the model's judgment; the rest follows by rule once those are set.mode's read), the band arithmetic (entrain · jitter · 字→句) = rules (v0 uses only the rules half — a fixed ceiling, no intent)So the line is clean: llm covers mode, focus, contention, closure, altitude; mechanical covers the shape deck, the dial math, and length's band arithmetic; and the three staging-traffic outputs — rhythm, length, attitude — are hybrid (the model sets the intent, the rules execute it). The mechanical halves are exactly what v0 ships with no model at all.
The producer is a utility model, independent of whoever voices the panel — its only job is to read and stage. Two settled calls: which model, and how it must phrase what it wants.
One discipline governs all three: always measured, v0 first. Each stage states the work (what's built), the expected effect (what it should buy), and how to test (the instrument that proves it). The point of building it this way: if a smarter stage doesn't beat the simpler one on the numbers and the feel, the intelligence wasn't where the problem was — and we keep the simpler one.
| stage | the work | expected effect | how to test |
|---|---|---|---|
| v0 deterministic built · shipped |
a per-turn [STAGING] directive by arithmetic alone: length envelope + silence + shape deck + seeded punch. No model — the control arm. |
kills the walls + breaks the metronome mechanically (turn-total ↓, max bubble ↓, rhythm varies). Stages blind — recency, not relevance. | replay A/B — regen real prod turns OFF vs ON, blind judge. Done: walls die, judge a wash (below). Real test = real users. |
| v1 smart producer designed |
one smart-model call: mode inference + the 4-move intent→staging + relevance who-speaks + the instruction hierarchy. The v0 envelope/deck stay as the floor it composes within. | staging becomes intent-aware: compress casual walls but let substance breathe, un-smooth real splits, stop benching the wanted host. A clear win, not a wash. | randomized prod A/B on real users — return (came back later) + walk-away satisfaction. A judge axis stood up for each new failure (contention, sycophancy) before optimizing it. |
| v2 dials + learning sketched |
dials visible + FP-suggest/auto-dial; the new sensors (engagement, phase, compliance) + actuators (closure, altitude); the learned-policy ladder. | user control + legibility (dials = a live readout); the wider failure set covered; the product improves between ships. | walk-away satisfaction — judge score + an explicit thumbs, never the clock. The learned-policy rung is itself the A/B harness, human-gated. |
The work. A deterministic per-turn [STAGING] directive, appended to the user turn (cache-safe like the dispatch brief), gated by a per-room flag (new chats default-on), non-blocking (errors/timeouts → today's panel). It computes — by arithmetic, no model — a length envelope (entrained ceiling + the shape deck's size profile, emitted as a sentence range), a silence lever (cap speakers, rotate who sits out), the shape deck (no-repeat, seeded), and a seeded punch slot (~1-in-3 turns).
Expected effect. Kill the walls and break the metronome — turn-totals and max-bubble down hard, rhythm varies turn to turn. It is the baseline: it cannot read intent, so it stages blind (recency, not relevance), and that ceiling is the whole reason v1 exists.
How to test. The replay A/B in § v0 measured — and the honest reading of why a single-turn judge under-measures it.
The work. Replace v0's blind decisions (recency · budget math · seeded draw) with one smart-model call that (a) infers the mode as a running belief, (b) runs the 4-move intent→staging — intent sets the length ceiling, substance sets the shape (the host with the meaty point leads; a genuine split → M·M), the beat gates punch, (c) picks who speaks by relevance not recency, and (d) phrases every want the way the model actually obeys (a structure + "fix it if you overshoot," never a count). The mechanical envelope + deck remain as the floor it composes within.
Expected effect. Staging stops being blanket and starts being intent-aware — the direct fix for v0's over-compression (compress the casual walls, let a substantive analysis breathe), un-smoothing the real disagreements the diagnosis found missing, and never benching the host the room most wants.
How to test. A randomized floor on/off prod A/B on real users, measuring two outcomes the single-turn judge can't be: walk-away satisfaction — did the chat deliver what you came for (the judge score + an explicit "that helped") — and return — did you come back for another chat later, value revealed by a free choice after you left (the healthy across-session signal, not within-session time-on-app). Read them together: return alone is gameable by an addictive product, so it's anchored to satisfaction — high return with low satisfaction is the warning light, not a win. Plus a judge axis per new failure before the producer optimizes it. Discipline: if v1 doesn't beat v0 on the judge and on feel, keep v0.
The work. Surface the three dials with FP-suggest + auto-dial; add the new sensors (engagement/affect trend, explicit commands, open threads, phase, last-turn compliance) and actuators (closure, altitude); stand up the learned-policy ladder (flywheel → per-user memory → a trained producer).
Expected effect. User control + legibility (the dials become a live readout of the producer's read), coverage of the wider failure set (confused↔flattened, lectured↔underwhelmed), and a product that improves between ships.
How to test. Walk-away satisfaction + return — judge score + an explicit thumbs + did they come back — never the clock (optimizing engagement breeds an addictive, sycophantic producer, the exact opposite of the goal). The learned-policy rung is itself the A/B harness, human-gated.
The build rule was "always measured." So: a replay A/B — take real prod turns, regenerate the panel reply twice (floor OFF vs floor ON, toggling only the floor on the same real history), judge blind. Model held constant (DeepSeek pro), 38 turns, a blind pairwise judge on the diagnosis pains.
| metric | OFF | ON (v0) | read |
|---|---|---|---|
| turn-total 字 (median) | 952 | 178 | −81% — the walls, gone |
| max bubble 字 (the wall) | 400 | 143 | −64% |
| #speakers | 3.5 | 2.0 | silence works |
| concise (judge 1–5) | 3.3 | 4.0 | brevity reads as quality |
| voice / fidelity (judge 1–5) | 4.1 | 3.5 | dropped — the cost |
| redundancy (0–1) | 0.00 | 0.00 | unmeasurable here (below) |
| ON win-rate (better turn) | 47% — a wash | no clear win | |
v0 trades voice for brevity, and on this judge it nets ~even. The mechanical win is real — the OFF arm reproduces the prod walls (952 字 turns, 400-字 bubbles) and v0 kills them, and the judge agrees the brevity is a quality gain. But compressing and silencing voices flattens their distinctiveness — the thing users said was already good — and the two roughly cancel. A tuning sweep then isolated the cost: loosening the length alone didn't help, but easing the silence (more hosts speaking) recovered the voice almost entirely (3.5 → 4.0, vs OFF's 4.1) while keeping the wall-kill — so the loss was the silence lever, not the compression. Yet no arm clearly beat OFF (best ≈ 50%).
A 3-host room (Buffett · Munger · a founder, Jensen). You ask: "Should I raise VC money, or bootstrap?" Here is that one turn today, then under each version (bubble width ≈ reply length):
History and cache are safe; behavior changes only going forward. Because the producer's directive is append-only (it rides your turn like the dispatch brief), it never rewrites a transcript and never busts a room's cache — old rooms or new.
The test: every failure needs a sensor (an input that detects it) and an actuator (an output that fixes it). Auditing the failure pairs against the current I/O finds the gaps. The additions:
| Add to inputs (sensors) | catches |
|---|---|
| engagement / affect trend — are your messages getting shorter, slower? | the reward signal, made first-class — "cooling off" is the loudest cue to change |
| explicit commands — "@X / shorter / make a doc / go deeper" | you telling it directly — given priority over inferred intent |
| open threads — your asks not yet answered | dismissed |
| phase — opening / exploring / converging / wrapping | no payoff — when to land a takeaway |
| compliance — did last turn obey the budget? | the panel ignoring the producer — escalate if so |
| Add to outputs (actuators) | catches |
|---|---|
| closure — diverge / hold the split / synthesize | confused ↔ flattened |
| altitude — plain ↔ expert register | lectured ↔ underwhelmed (distinct from the Personality dial) |
Honest answer: a frozen model does not learn from new chats on its own. "Self-improving" has to be built as a loop — but we already built the hard half (the judge), so the producer is instrumented for improvement from day one (every output is a judge axis). Three rungs, climbed in order:
| Stage | What learns | Mechanism | Cost |
|---|---|---|---|
| 0 — have it | adapts to this chat | the recurrent running belief (the mode carried forward) | by design |
| 1 — flywheel do first | the product improves | auto-judge prod → trends → ship prompt/mode fixes → A/B | low |
| 2 — memory | per-user prefs + an example bank | learn "likes it short, dislikes grilling" to seed the starting belief; retrieve top turns as few-shot | medium |
| 3 — learned policy | a trained producer | fine-tune / RL on (state → staging → outcome) | heavy · later |