The panel's meaning is good. Its register is an essay. This is the measurement of that gap against the humans sitting in the same rooms, the three things that actually move it, and the one that does not. Status: shipped — the meter, the net, and the prompt work. Written 2026-08-16 · second sweep and the survivors' batch 2026-08-20.
Intuition does not survive a frequency problem, so the first thing built was a ruler, not a fix: exam/aitone.py counts fifteen named tells and reports each as a rate per 1,000 characters. The control is not zero — it is the human column of the same corpus. Real people use colons and the odd metaphor, and a turn scrubbed to 0 reads as a telegram, which is its own tell.
The corpus: 40 real production rooms off the SG box, long human-with-panel chats, no game cartridges and no test rooms — 7,563 panel turns against 5,620 human turns (scripts/sweep_prod_chats.py).
Not from the model being stubborn. Running the same meter over the prompts and profiles the panel reads every turn found them saturated with exactly the mark they were producing:
| what the model reads | em dashes / 1k |
|---|---|
| the human turns it is answering (the target) | 0.86 |
the megaprompt (system-prompt.md) | 5.78 |
| a typical persona profile | 5.3 – 8.3 |
| what it then writes | 10.81 |
The model reads roughly 27,000 characters of analyst prose about a person, then writes as that person, and carries the register across. It is not disobeying an instruction; it is obeying the most recent demonstration. A second cause sat in the code: the FORMATTING note told every host its words render as Markdown and to "use it naturally when it helps (a list, a table, emphasis)" — read as a licence to publish into a chat bubble.
Three arms were tried and measured. Only two earned their place.
run_room.de_tell() runs on on-stage speech only, at the single seam where a <speak> becomes a bubble. It spends every em dash for the mark a thumb would have typed. Fenced blocks are left alone (a dash inside a diagram is syntax), and artifacts are never touched — the side pane is where writing is supposed to look written.
Replayed over all 7,563 production turns it removes 64% of the whole excess over human, at zero cost and zero latency, because a punctuation mark carries no meaning and so nothing is lost by spending it.
A How a message looks section in the megaprompt names the tells outright, and a closing chat_register_note() sits after the profiles, telling the model in as many words that the essay prose it has just read is the analyst's and not the host's. Measured by exam/tone_ab.py — 20 real guest messages replayed against the live megaprompt, base arm in a git worktree, both arms scored raw:
The whole「去 AI 味」ecosystem on GitHub — speak-human-tw (38 patterns), shuorenhua, ai-flavor-remover, de-ai-flavor-skill — is one design: a symptom checklist handed to a second LLM pass over finished text. No code, no measurement, and no human control. Their lists nonetheless independently confirm the meter: 破折號過度使用, 否定平行結構「不是A,而是B」, 排比三段式, 粗體過度, 編號切碎, 句長過勻, 說教式深度腔, 刻意換詞循環, and — for Vibes — 幻覺引用「提升 47.3%」. Fifteen detectors built from this room's own data landed on the same list.
Two levers came out of the literature rather than the repos, and both were measured against the shipped baseline (n=20, bench noise ~±0.5):
| arm | index | vs baseline | verdict |
|---|---|---|---|
| human control (the target) | 15.12 | — | — |
| baseline — the shipped prompt work | 56.23 | — | — |
| show the register — quote the guest's own last lines back as a style sample | 53.72 | −4% | off by default |
| positive framing — "punctuate the way a thumb types" rather than "never use an em dash" | 50.81 | −10% | shipped |
Positive framing is the cheap one that worked, and it matches the one solid mechanism in the literature: a negative constraint has to suppress an already-activated path, while a positive one just names the path to take (constraint-compliance study). Note the popular「Pink Elephant」explanation for this is conceptual, not measured — the widely-cited article runs no experiment, so don't lean on it.
Stacked, the prompt arm reaches roughly −15%. The net is still doing four to five times as much work.
The obvious follow-on was to strip the dashes out of the instruction files so the demonstration stops. It was done, measured by eye, and reverted: in instruction prose those dashes are structural, not stylistic. THE FIREWALL — never cross it. became THE FIREWALL, never cross it., and a comma splice in a rule the panel has to parse is a real risk to the thing this whole exercise is meant to protect. The 88 persona profiles were left alone for the same reason, and because the net already zeroes the output.
The axis that decides this already exists and costs nothing: the floor producer's own length ceiling. It has tiered every turn since the flatness work — minimal ≤40字, tight ≤90, medium ≤140, generous ≤280, solo-deep 500–2000. keep_layout() reads the cap straight off the staging line: generous and up keeps its Markdown; below that it is demoted to the same lines without the furniture. No second judge, no extra call, and a missing staging fails open — a turn with no producer read never has a guest's structure stripped from it.
With demotion in play on conversational turns, the net removes up to 76% of the excess; the live figure sits between that and 64%, depending on how many of a room's turns are analysis turns.
A separate surface, the same disease. Vibes posts kept landing invented precision — "47 minutes of paper jam", "it beeped 14 times", "37 times this week". The cause was literal: lib/vibes_write.py asked every post for "one concrete number", and the authoring kit's own worked example was "37 个拒绝". The model was quoting its brief.
how-to-write-vibes-moments.md together, so the rule and its example agree.2026-08-20. The owner, reading the live rooms:「the personas are still using 不是…而是…」. A fresh corpus off the box — this time from the event log, so every turn carries its timestamp and the days before and after the net went live (the box pulled it 08-16/08-17) can be read separately. 17 rooms · 365 panel turns · 260 human turns over ten days; 120 panel turns are post-net.
First, the net is confirmed working in production — and the owner is confirmed right about what it never touched:
| tell, in the post-net production turns | human | panel |
|---|---|---|
| em dash / 1k (the net's job) | 0.32 | 0.00 |
| markdown in a bubble / 1k (also the net's) | 0.00 | 0.00 |
| turns carrying 「不是X,(而)是Y」 | 1.2% | 28.3% |
| …its reversed face,「要的是X,不是Y」— which the meter did not count | 0.4% | 12.5% |
| turns minting ≥2 scare-quoted concepts(「抢风貌」·「真实感」) | 1.2% | 11.7% |
| turns with a colon-to-expound | 10.0% | 50.8% |
| mean turn length, characters | 72 | 196 |
Two more finds, both with a traceable cause:
*[点头]* *[停顿]* *[话锋一转]* on 37 of 37 post-net turns, ~3 a turn, from a fixed little vocabulary. The source was megaprompt rule #6, which said never decoration — while demonstrating the exact two tokens. 点头 and 停顿 are literal translations of the examples it gave. The model kept the vocabulary and dropped the constraint, and *[停顿]* does precisely the job the em dash used to do: a typeset pause. The register escaped the net by changing costume., while every human in the same rooms types 「,」— and the net's own dash→,patches were showing as the only full-width commas in a bubble. A visible seam, and meaning-free.stage_direction, quoted_concept (soft — humans quote too), and ascii_comma_cjk as the regression meter for the seam.ascii_comma_cjk keeps the rate on the books; if a future model stops doing it, the rate reads zero and no code needed changing.The prompt half of the batch, benched the standard way — exam/tone_ab.py, the recent corpus replayed, base arm at the pre-change HEAD in a worktree, both arms scored raw by the upgraded meter (13 effective samples; 7 of 20 hit casts whose app-built personas live only on the box, and failed identically on both arms):
The owner's standing observation — Chinese lines carrying ASCII commas(「五个骰子,每人一个骰盅」)— asked whether the English/Chinese mix in the FP directive is the cause. Census over the dev corpus (10,499 persona CJK lines, 1.24M 字): the base rate is a tiny 0.44 marks per 1000 字, but it is violently concentrated — a handful of rooms run 40–80 per 1000 (100–200× base) while the mass are spotless, and an infected room is infected from its very first line, before any staging ever ran. Only 12 of 541 half-commas sit near any Latin text, so English words in the message are not the trigger; profile language alone isn't either (an all-English profile ran 101k 字 clean; tess is dirty in one room and clean in her others).
The mechanism is a kickoff lottery plus transcript contagion. A/B on one cast (banksy + koons · 中文 · 6 kickoffs per arm): with the production English kickoff cue, 3/6 openings came out comma-broken — and wholesale (~50/1000, exactly the hot-room rate) — while 3/6 were perfectly clean; a Chinese-translated cue dropped it to 1/6 (helps the odds, doesn't cure — n too small to call significant). Whichever way the first turn lands, the transcript teaches every later turn to match. So: the FP directive is acquitted (infection precedes staging); the English-heavy first-turn context is a risk factor; the root is DeepSeek sampling its own punctuation habit once, at the room's first Chinese turn, and the room inheriting the coin flip forever.
zh_punct_first at the _record_turn seam — kickoff or, in a kickoff-less room, the first reply), where a crude regex is provably safe: a greeting cannot be dictating CSV columns, teaching punctuation, or quoting a typo. From a clean opening, contagion and the model's own judgment keep every later turn right — the smart way, in the owner's words: the model, not a regex, decides when a half-width mark is genuinely meant mid-conversation. Guards regardless: tight CJK flanks only (spaced examples untouched), sentence-final ?! included, fences and inline `code` skipped; a resumed mid-life room is past its lottery and never touched. GATE: the lottery A/B re-run — 3/6 infected kickoffs before → 0/6 after. Already-infected old rooms keep their habit by design (their transcript still teaches it); only new openings are inoculated.The A/B builds the base arm in a git worktree, never a checkout in place — parallel sessions share this working copy. It scores raw, before the net, because the question it answers is whether the writing moved, not whether the net can mop up after it. --base-bubbles reuses a paid-for base arm when only the working tree changed.
word_echo, and the turn skeleton the second sweep names above — none of which is safe to rewrite mechanically.exam/aitone.py, the net is run_room.de_tell, the bench is exam/tone_ab.py. Sits beside conversation quality (the flatness diagnosis) and the floor producer, whose length ceiling this page borrows.