Explored, agreed-direction features for the live room — recorded in idea.md so the reasoning survives until they are built. Captured 2026-06-10 from a design discussion. Since then three have shipped — §2 room dispatches (v47), §5 the smart panel selector (2026-06-18), and §1 the floor producer (deployed to production 2026-07-03): its v0 deterministic baseline landed on main 2026-06-24 as a per-room toggle, and the smart engine (v1p / v1f) — v1p the live default when the producer is On — plus an independent opening producer (ships dark, admin-armed) followed, all wired as admin-console toggles (the Models tab); §3 listen mode and §4 compaction remain planned. Separately, the perceived-speed / local-first infra has shipped to production — see Perceived speed & the slow-network playbook — and the room LANGUAGE system shipped + deployed 2026-06-21: an account default language (4 langs, OS-seeded) and a per-room target the panel speaks in-character, with a faithful translation appended under a hairline when a persona slips script — see Room language. Two further ideas were added 2026-06-21 at discussion-stage — §6 persona mailboxes and §7 send-to-many — recorded while their direction is still forming (earlier in life than the settled §1–§5).
The megaprompt scheme wins on fluency (exp-010: hosts felt real and distinct, conversation flowed) — and its flatness is the same architectural property seen from the other side. One mind holds the whole conversational field, so handoffs are perfectly timed and nobody collides (fluency); and one mind smooths the whole field — it averages the room, gives everyone something to say, resolves tension at a single author's comfortable rate (flatness). Telling the megaprompt "vary your rhythm" drifts back to smooth over a long room: the writer and the staging-decider are the same context, and the writer always has something for everyone to say.
A separate per-turn agent that reads the room and emits a one-line staging directive before each panel call. It decides participation, order, length, and temperature — never content. The single writer stays intact for micro-coherence; unevenness is injected exogenously, as an outside signal the writer must accommodate rather than a disposition it must remember to maintain. The directive (~50 tok) is appended as a small control message before the panel call, fitting §8 append-only caching — the cached prefix stays untouched.
| Producer MAY say | Producer may NOT say |
|---|---|
| who speaks this turn / who stays silent | what anyone should say or argue |
| order ("Plato first, then the room") | topics to raise ("push him on the ethics angle") |
| length budget ("one line each" / "let Plato run") | positions ("disagree with the guest") |
| temperature ("press", "let it breathe", "wrap") | emotions as content ("be angry about X") |
The most valuable directive in the vocabulary is silence ("Confucius sits this one out"). LLM hosts are over-helpful; rhythm comes from suppression more than stimulation. A content-leaking producer manufactures drama on schedule — the exp-004 lesson generalized: an engineered cue upstream produces an engineered beat downstream, invisible in the transcript. In the megaprompt scheme that's "only" a quality bug; if this ever runs over isolated actors, the producer channel would be the intent-leakage vector — so the vocabulary contract is cheap insurance taken early.
dynamics / intensity (creation-time, baked into the cached prefix) define the envelope; the producer modulates within it per turn. The static dials stop being the only rhythm control.main 2026-06-24; per-room toggle, new chats default to it; deploy batched): a deterministic per-turn [STAGING] directive computed from public room state only — silence (recency-rotated lead, @-mention override, occasional solo turn), a length budget (language-aware — 字 for CJK, words otherwise) that mirrors the guest, and variance. No model call; append-only so it never busts the §8 cache. Establishes how much flatness is just participation math before crediting the smart producer with the rest.extreme dial delivers ~⅓ of the conflict it promises — the heat goes host→human, not host↔host. It sizes the levers above and adds one: the producer must be able to assign genuine host↔host contention. → the diagnosis · the full design.Personas are bounded by the base model's knowledge cutoff plus their profile — but "Plato meets today's news" is one of the strongest things the room can do. The fix: search is not a persona tool. A hidden @Web Search that the user @-mentions runs a web correspondent (Haiku 4.5, max 1 search, ~$0.027) that distills a neutral, dated, sourced ~150-word fact-brief. It lands once in the stream (visible to hosts and humans alike) and the personas react under their own extrapolation contract — same fact, divergent voices, the single best live demo of the isolation thesis.
Why search is not a persona tool: (1) voice contamination — raw news prose makes a host paraphrase the journalist instead of reacting in his own voice; (2) it breaks the extrapolation anchor — a current event is a distance-2/3 question with a fact-brief handed across the table, not something the figure "knows"; (3) §8 cache + latency — raw results would bloat every later prefix, and agentic multi-search adds 10–30 s while live humans wait.
The staged live sequence: searching bubble → dispatch card + cost pill + ∞-pane artifact → conferring egg → replies. There is no topic policy — a deliberate user decision (fetch anything, wire-service neutral). Done — see Rendering & speed and the live app.
Read each host's bubbles aloud in a per-character voice. Serves people who'd rather listen than read, mobile use, and the room's nature — it is already a panel show, and audio is its native medium. With the floor producer (§1) pacing it, this composes into a generated-podcast pipeline.
| # | Problem | The shape of the fix |
|---|---|---|
| 1 | Collision within a room | accent×gender is too coarse (Buffett + Jobs + Feynman + Karpathy = four American males). Add a timbre/age/pace axis, plus room-scoped collision-avoiding seating persisted like the colour PALETTE — voice anchors identity only if stable across sessions. |
| 2 | Accent × room language | rooms run in zh/ja/en (now a first-class room-language system); "British English" means nothing when Churchill speaks 中文. Mapping: card voice: → bucket (accent·gender·age) → room seating → per-language provider voice ID. Bilingual voices (e.g. CosyVoice) collapse the per-language column. |
| 3 | The parody line | accent aids identity, never becomes the bit — the audio edition of "amplitude, not parody." The bucket includes a dignified-neutral option; assignment is a judgment call recorded in the persona card by the authoring pipeline, not auto-derived from birthplace. (Jack Ma's English: iconic and his. Confucius: neutral gravitas.) |
| 4 | What gets voiced | on-stage <speak> only — never <confer>, artifacts, or code. Stage directions become SSML pauses (*[long pause]* → an actual pause). Strip markdown before synth. |
Flat reading kills the panel feel — the TTS must take a delivery instruction per utterance. The signal already exists in-band; the style prompt is derived, not authored (no new content channel, no new leak vector):
Instruction-following is a hard filter on provider choice — flat direction-following (gpt-4o-mini-tts, Kokoro) is disqualifying; speech-specialized models are the bar.
China consequence: Gemini / ElevenLabs CDNs are blocked from mainland — direct delivery fails exactly for China-path users, while Alibaba OSS / MiniMax CDNs work there. So: either one China-reachable instructable provider for everyone (DashScope / MiniMax), or per-region routing (CN → DashScope, intl → ElevenLabs/Gemini). This is the evaluation to run when building.
Caching: audio is immutable per (text + style, voice) → lazy synth on first play (readers who never listen cost zero), browser-cached, provider-URL expiry accepted (re-synth is fractions of a cent). First candidate: Qwen3-TTS-Flash on DashScope — instructable, OSS-URL delivery, China-reachable; ~0.4¢ per average bubble. (The vendor-synergy tiebreaker has since lapsed: the dictating mic shipped on iFlytek, not DashScope — TTS is now a standalone vendor choice.)
§8 caching is append-only: every turn re-sends the whole history and the prefix only grows — it never shrinks. Left alone, a long room's model context climbs without bound, and three costs scale with it: dollars (a 500K cached prefix still costs ~5× a 100K one to read, and past the 1h TTL it's a full re-read), latency (TTFT grows with prefix size even when cached), and effective-context quality (lost-in-the-middle and mild drift grow as the working set fills, well before the hard limit).
When a room crosses a high token threshold, fold everything older than the last N turns into a single model-written «story so far» block, and keep the recent turns verbatim. The persona cards (stable prefix head) and the recent tail stay; the long middle becomes a summary.
A summary that drops a load-bearing detail or smooths a persona's position is the megaprompt-flatness failure in a new place. So it preserves, explicitly: the arc (what's been covered), decisions/facts reached, each persona's established positions (so they don't contradict their earlier selves), open threads (so they aren't silently dropped), key human disclosures (the room's whole point), and artifact references (title + id, not full text — artifacts already live in state.json and the wildcard pane). The recent N turns stay verbatim — recency is where micro-coherence lives.
state.json keeps every turn; the user still scrolls the entire chat. It's invisible in the transcript — which makes it safe-ish to get imperfect (the record is intact) and trivial to A/B (replay the event log with/without).§8 append-only → invalidates the prompt cache → a one-time full cache-write at the compaction turn. The threshold must be high enough that the reset amortizes over many subsequent cheap turns. This is exactly why "summarize every turn" is wrong.→ open the interactive mockup — the full new-chat window with the project's real persona library (describe → three panels → editable rows with the reason → hand-pick → folded dials).
§1–4 are all about conversation quality once a room exists. Nothing addresses how a room gets formed — the one cost that grows with every persona added. Today the only way in is "browse all personas and hand-pick," which fails two ways: cold-start (a new user has no idea who 小胖看房 is, or why they'd want Buffett plus a skeptic — the blank picker asks them to choose from strangers) and scale (browsing 8 is fine, 80 is not, and the catalog is designed to keep growing — structural, monotonic cost, the room-formation twin of §4's append-only prefix).
+New pickerNot a new surface — a mode toggle on the v0.3 picker modal. The substrate exists: tags.txt per persona + persona-search.js fuzzy search. The smart door is the default (better cold-start); both are first-class, and the curated set lands in the same grid as a pre-filled, editable selection — never a separate screen, never a locked black box. Augment browsing, don't replace it.
| Deterministic (tag / embedding match) | LLM-curated (chosen) | |
|---|---|---|
| Reads "anxious first-time founder" | weak — keyword overlap only | yes — reads intent |
| Picks complementary angles | no — similarity rank → echo chamber | yes — can be instructed to |
| Explains why each seat | no | yes — the rationale is a feature |
| Cost | ~free | one cheap call per room creation — infrequent, high-leverage |
tags.txt, not full profiles) plus the user's goal; get back 3–5 slugs each with a one-line rationale. DeepSeek tier is plenty, and it fires once per room — not per turn. Full profiles are withheld on purpose: selection needs the card, not the interior, and it keeps the prompt small.kickoff_cue / dispatch_brief), it inherits the floor-producer dumb-pipe contract: frame the topic, never pre-load positions, never leak "the user is anxious" as steering. Same intent-leak vector as §1 — cheap to guard at design time.Captured 2026-06-21 — discussion-stage, nothing built. Earlier in its life than §1–§5: the direction is still forming and the open questions below are real.
→ open the persona-homepage mockup — the page before the DM: who a persona is and why you'd write, with the machinery hidden (switch across four persona types).
§1–5 all live inside the live room — synchronous, group, SSE-streamed, on the critical path. There is exactly one way to reach a persona today: open the app and convene a panel. But a persona is pure data driven by one system prompt; that engine doesn't have to be spoken to only through the room. Email is a second front door — a 1:1, asynchronous, store-and-forward channel that meets the user in their own inbox. "A letter back from Buffett" is a different product feeling than the panel: intimate, unhurried, pen-pal — and it costs the user nothing to keep open (no app, no session). Because it's async, latency stops being a cost (a reply is expected to take minutes), so it's a lighter lift than the room in every way except deliverability ops.
auth.py already issues) closes both, and hands the persona a known correspondent. So this idea carries a real dependency: a user-profile / verified-identity layer that only partly exists today.state.json; an email thread is a durable per-(user, persona) history. Cleanest model: a room with cast = 1 + one human, persisted. This is the first taste of cross-session persona memory — handle it deliberately.functions/ CF footprint) handle inbound parse + the whitelist check and POST a clean payload to the box; use an API sender (Resend / SES / Postmark) for the DKIM-signed reply. MX / SPF / DKIM / DMARC + bounce handling are the real ops tax — that's the cost to weigh.Captured 2026-06-21 — discussion-stage, nothing built. Open question flagged at capture: a feature, or a separate product?
One question — or a whole questionnaire — fired at many personas at once, each answering in a clean, isolated context with no shared transcript. Collect the replies, then aggregate them into something readable (a matrix, a contrast, a summary). Pitched as market research across personalities: how would N different archetypes each react to this pitch / message / decision?
The room is the mixed condition: personas share one append-only context and react to each other (fluency, but convergence pressure). Send-to-many is the isolated condition at its limit — N independent inferences, zero shared context, zero convergence. §2 already named the prize ("one fact → N divergent reactions… a monolithic model given search would converge them"); this is that, with contamination set structurally to zero. It's the legacy AI↔AI instrument's core question — what does shared context do to the answers? — promoted to a user feature. The project already does this internally: the cold-subagent persona audit is send-to-one-cold-context.