← Design notes

Liveliness — what the field knows about feeling alive

A four-sweep web literature review — academic role-play fidelity · shipping companion apps (C.ai, Replika, 星野, 猫箱, Talkie, 筑梦岛) · HCI perception studies · drift-and-convergence engineering — synthesized into a ranked shelf for a consumer app that cares about user-perceived liveliness. Trigger: room c3d8's verdict (Banksy too mild, two artists converging). Written 2026-08-31. Status: research — the shelf is ranked, nothing below is authorized until the owner picks.

The one-sentence synthesis: mildness and convergence are the model's resting state, not a defect — RLHF averaging, in-context style gravity, and same-model output collapse all push every seat toward one polite template — so fidelity is not a thing you author once in a profile; it is a thing the orchestration must actively re-assert every turn, and most of what users read as「alive」is cheap: verbatim voice exemplars, per-seat structural difference, timing and chunking, and being remembered.
why every seat sounds the same by default — the resting state RLHF averaging trained toward the safe middle in-context style gravity output converges to the room's register same-model collapse one brain, shared templates Banksy savage deadpan irony one polite template where every seat lands unopposed Jeff Koons salesman's calm pull pull the counter — re-asserted EVERY turn, because the pull never stops voice: note · example exchanges · per-seat staging · opener bans · a divergence meter
Convergence needs no cause — it is the model's equilibrium. Three standing forces (RLHF sycophancy/averaging · in-context style convergence, stronger than the human baseline · same-model output collapse) pull every seat toward one register, and Chameleon's Limit shows the collapse worsens as history grows. So distinctness is not authored once — it is re-applied per turn by the orchestration, or it decays.

What the four sweeps agree on

findingstrongest evidencewhat it means for -ish
1 · Verbatim exemplars beat descriptions for voice. Retrieved/quoted character speech as few-shot demos is the industry- and lab-validated voice anchor; models absorb register and form from demos even when the content is irrelevant (random demos beat retrieved ones; even corrupted demos helped). Transcripts of the person beat prose about them; affect tags on excerpts help.RoleLLM (style score 41.3 vs 23.2 baseline) · RAGs-to-Riches (+35% reference-token pull under hostile prompts) · ICL study · C.ai Definition = example dialogs, greeting = the strongest single style sampleProfiles carry attested quotes as prose bullets and a corpus/ nobody reads at runtime — but no example-exchange block in dialog form, no affect tags, no group-interjection examples. The highest fidelity-per-token gap we have.
2 · Convergence is the default state. Same-model outputs collapse to shared templates absent a forcing function; distinct personas converge worse over longer histories; per-persona fidelity and population diversity trade off; RLHF preference data actively rewards caving. Instruction drift is significant within ~8 turns — attention to the system prompt decays between turns.Chameleon's Limit · Artificial Hivemind · persona-drift/COLM · Anthropic sycophancy「Two artists converge」is physics, not a profile bug. Fight it structurally, per seat, every turn — contrastive staging, per-seat register, opener bans — and meter it, because it regrows.
3 · The context tail is the leverage slot. Re-injecting persona material before each turn (SPR) is a validated drift baseline; attention is U-shaped and the freshest tokens win; bounded per-turn steering deltas on a fixed base recover most stability (+62%) for little adaptivity cost (−17%).split-softmax/SPR · Lost in the Middle · synchrony–stability frontier · C.ai pinned prefixOur [STAGING] line already rides the tail every turn — a free, cache-safe carrier for a per-speaker voice re-anchor. The voice: card note (shipped) is the payload; appending it verbatim is one deterministic line.
4 · Perception is cheap. What moves perceived humanness most is not model quality: bursty short bubbles beat one block (bigger effect than adding an avatar); human-plausible timing is the tell users hunt (and they still can't beat chance against a well-timed agent); a made-then-corrected typo was the strongest single humanness cue ever tested (N≈3,400); casual, topic-congruent register raises warmth and trust.Chen 2022 (chunking) · HUMA (55.4% ≈ chance) · Imperfectly Human · dynamic delaysThe turn split gives per-host bubbles; the within-host burst (1–3 short bubbles, staggered, WPM-plausible) and per-persona typing cadence are the missing half. UI-layer, model-free.
5 · Relationship beats performance for retention. Memory callbacks are the top「it's alive」trigger in user reviews and block relationship formation when absent; bot self-disclosure reliably deepens engagement (replicated); relevance-gated proactive messages carry the strongest retention evidence of any mechanic — with a compounding annoyance boundary (ask, don't assume; never guilt).Dialoging Resonance · companion-app re-engagement study · ComPeer · 星野 事件簿 · Replika diaryMostly already built: memory row ③, threads, agenda ④⑤ (own-move, own news), ⑦ pings (dry on prod). The gap is visibility — personas rarely cite a specific remembered moment by name — and the ⑦ send switch.

The three mechanics, drawn

1 · Show the voice, don't describe it

prose ABOUT him — what profiles mostly carry
"He is deadpan and self-deprecating; his register is the joke that carries the blade. He deflects with a one-liner rather than a speech."
transcript OF him — what the model actually copies
Guest: "Shredding your own painting just made it worth 18× more." Banksy: "Record price for a Banksy painting set at auction tonight. Shame I didn't still own it." [dry · unbothered]

Models imitate the form of what they are shown, far more than the content of what they are told: quote-demos lift style scores where descriptions plateau (RoleLLM), verbatim transcripts beat wiki-derived prose at equal prompt budget with ~35% more reference-pull under hostile prompts (RAGs-to-Riches), and even random or corrupted demos help — evidence the model absorbs register, not facts (ICL study). Character.ai's whole Definition format is this finding productized (creator book). Affect tags on each excerpt help; the attested quote above is already in the profile — the lever is re-formatting into dialog form, not new material.

2 · The attention sandwich — why re-anchoring rides the tail

one panel call — where attention actually goes persona profiles + room rules primacy — but its share decays each turn conversation history the lost middle [STAGING] + voice re-anchor the freshest tokens — the tail slot low ← attention share → high high, and falling per turn lowest highest drift is significant within ~8 turns — attention to the prefix decays between turns; re-injecting persona material per turn (SPR) is the validated fix, and the tail is its slot
The U-curve is measured (Lost in the Middle); the between-turn decay and the re-injection baseline are measured (instruction stability, COLM 2024); bounded per-turn deltas on a fixed base recover +62% persona stability for −17% adaptivity (synchrony–stability frontier). Our staging line already occupies the tail every turn — shelf lever 2 just gives it a voice payload. Character.ai's production answer is the same shape: an immutable pinned persona prefix, history truncated around it (their prompt-design post).

3 · What actually moves perceived aliveness

measured with real users — none of it is model quality a corrected typo JACR '24 short-bubble bursts Chen '22 human-plausible timing HUMA '25 memory callbacks CSCW '24 in-character casual register HAI '25 relevant proactive pings arXiv '25 emoji flips by context — negative when the topic is serious bar = strength of the measured effect on perceived humanness / social presence (qualitative tiers — units differ across studies); every row is orchestration-level — no model change anywhere on this chart
The typo-correction result is the strongest single cue ever tested, beating name, gender and photo head-to-head across 5 experiments, N≈3,400 (Imperfectly Human); five bubbles beat one block, with a larger effect than adding an avatar (Chen 2022); against a well-timed agent in 4-person chats, humans identified the AI at chance — 55.4% (HUMA); callbacks and disclosure are the replicated relationship levers (Dialoging Resonance); re-engagement pings carry the strongest retention evidence (companion-app study) — and the same paper is why ours must never guilt-trip.

The ranked shelf

Ordered by perceived impact ÷ implementation cost. Status: new not built · part half exists · built exists, needs dialing.

where each lever lands — six layers, one app PERSONA KIT #1 example exchanges — new #3 fellow-panelist stances — new voice: note — shipped 08-31 mined from corpus/, not invented MEGAPROMPT · the panel #4 shared-skeleton bans — new #7 callbacks by name — dial up (the memory store already exists; the gap is the visible reference) FLOOR PRODUCER #2 voice re-anchor at the tail — new neutral English cues — shipped shape variation · mention law — shipped 08-31 HARNESS · UI #6 burst bubbles + cadence — part (turn split exists; the within-host burst + typing pace is the half left) #9 diary · #10 voice notes — tier 3 CONSOLE #8 ⑦ proactive pings — built, running dry on prod the lever is the Initiate tab's send switch — no deploy EXAM/ #5 divergence meter — new #5 boundary-query OOC battery — new the watchdog: convergence regrows, so every lever above needs a ruler coral = new · gold = half exists · green = already shipped — the staging half of this shelf landed 2026-08-31 (floor-producer #fw-register · #fw-voice)
Twelve levers, six layers. Tier 1 (#1–#5) is entirely prompt/server; tier 2 (#6–#8) adds light client work; tier 3 (#9–#12) is new surface area. The green cells are this week's shipped work — the map shows how much of the field's shelf -ish already owns.

Tier 1 — prompt / server only, days

#leverwhat, concretelylands in
1Example exchanges per persona3–6 verbatim in-voice exchanges per profile (mined from corpus/ for real figures), affect-tagged, including at least one group interjection and a distinct entrance line; a dedicated megaprompt slot renders them as dialog, not prose. Consensus finding 1 — the single strongest fidelity lever the field has. ⚠ Our own S-slot study found few-shot HURTS — but that was few-shot inside the FP craft directive; this is few-shot in the persona block, a different layer. Bench it with exam/tone_ab.py against its own before.persona kit + megapromptnew
2Voice re-anchor at the tailCode appends each speaking host's card voice: line to the staging block, every turn — SPR-lite, deterministic, no extra call, cache-append-only. Directly targets the 8-turn drift mechanism.run_room (one line)new
3Fellow-panelist stancesEach persona knows what it makes of the other seated hosts (a card-invisible profile note or a seat-time pairing line): Banksy has an opinion of Koons and it shows unprompted. Fuels genuine contention (the flatness study's missing host↔host friction) and the agent↔agent life only a group room can show — 星野's hidden-depths mechanic, on our home turf.persona kit + megapromptnew
4Kill the shared turn skeletonThe「Daven,/Dan,」opener on 100% of replies, the quote-grade-pivot-verdict shape the AI-tone page already named as the residue — a megaprompt rule that co-seated speakers must not share an opener or shape, plus per-persona structural signatures (Koons name-drops because his profile says so; Banksy must not). Structural convergence is the default (finding 2); this is its cheapest visible symptom.megapromptnew
5Measure it: divergence meter + OOC batteryAn exam/ metric for cross-host structural similarity per turn (Coverage/Uniformity-lite, per Chameleon's Limit) + a per-persona boundary-query battery (ERABAL-style counterfactual probes, OOC rate tracked). Guards every other lever — convergence regrows, so it needs a ruler, not a one-time fix. Judge caveat noted by CharacterEval: LLM judges drift from human perception — same stance as our own studies.exam/new

Tier 2 — server + light client, about a week

#leverwhat, concretelylands in
6Burst bubbles + typing cadenceA host's reply may arrive as 1–3 short bubbles, staggered at human-plausible pace, per-persona typing indicator between them. The turn split already made per-host bubbles; this is the within-host half. Chen 2022: bigger perceived effect than an avatar; every CN companion app treats it as table stakes.harness + room-uipart
7Callbacks made visibleThe memory store exists (row ③) — the megaprompt nudge is for a persona to occasionally reference a specific earlier moment by name. Boundary from the lit: precise recall of something sensitive reads as surveillance, not intimacy — the visibility/incognito rules already fence this.megapromptbuilt
8Flip ⑦ pings liveThe strongest retention evidence of any mechanic in the review — relevance-gated, phrased as asking not assuming, never guilt-tripping (the manipulative variants are documented and to be avoided). Built and running dry on prod; the send switch is the console's Initiate tab.console switchbuilt

Tier 3 — heavier, owner's call on direction

#levernote
9Persona diary — the dream made visibleReplika's single most-cited「why it feels real」feature is the companion-POV diary; our dream/reflection pipeline already writes the content — the build is a reading surface.part
10Occasional voice notesA persona sends a voice note (not calls) — 猫箱 gates voice behind VIP, a revealed-preference signal; tts-providers.js exists.new
11Moments / off-screen life feed筑梦岛/SoulLink: the persona exists when you're not talking to it. New surface + scheduled generation; pairs with the weekly recent.md sweep.new
12Corrected-typo experimentsStrongest single humanness cue tested — but dosage-untested, register-risky in Chinese, and pattern-detectable at frequency. A/B-only curiosity, never a default.new

Cautions the literature is loud about

The experiment the literature cannot answer

our home turfNobody has measured whether visible persona↔persona disagreement and banter reads as aliveness to the humans in the room. The closest work shows multi-agent groups exert real social force and that observers use human descriptors knowing they're AI — but agent↔agent friction as a perceived-liveliness lever is an open gap, and a group-chat app with authored real-figure personas is the instrument to close it. Lever 3 (fellow-panelist stances) builds the mechanism; the return-rate A/B judges it. This is also exactly the flatness study's oldest unpaid debt: extreme never delivered host↔host contention.

Sources

Academic role-play fidelity: RoleLLM · prompt-vs-exemplars ICL · RAGs to Riches · ChatHaruhi · Emotional RAG · CoSER · Ditto · ERABAL · CharacterEval · InCharacter · Chameleon's Limit · RMTBench. Drift & convergence: instruction (in)stability · linguistic convergence · synchrony–stability · Lost in the Middle · sycophancy · Artificial Hivemind · C.ai prompt design. HCI perception: dynamic delays · Imperfectly Human · multi-bubble replies · Dialoging Resonance · ComPeer · Inner Thoughts · HUMA · Multi-Agents are Social Groups · Generative Agents. Consumer practice: C.ai creator book · C.ai memory · 星野评测 · 星野对话设计 · 猫箱试用 · companion re-engagement study.

Liveliness · 2026-08-31 · the trigger room is c3d8; the staging half of its fixes shipped the same day (floor producer #fw-register + #fw-voice). Sits beside AI tone (register), conversation quality (the flatness diagnosis), and many minds (the convergence endgame).