Dialogue.  · The word pool

The word pool — one pool, every persona

The word well generated a bespoke set of secrets for each persona: $0.50 and twenty minutes each. That cannot survive a roster of user-made characters — nobody waits twenty minutes to play with someone they invented thirty seconds ago. So the authorship moves off the persona entirely. One pool is built once for everybody; a persona is served by ranking that pool against its own card. Per persona: one embedding call and a dot product.

The thesis, in one line. A persona's vividness was never in the word — it is in the persona talking about the word. Once you believe that, the secret does not have to be rare, it only has to be theirs enough to be worth their breath; and picking a common word out of a shared pool is a dot product, not a twenty-minute generation job.

1 · The shape

ONCE — shipped in the repo open lexicons → curate → generate pairs → embed 7284 items · 5.6 MB of int8 vectors lib/wordpool/*.json + *.bin ~10 minutes · under $1 · never repeated PER PERSONA — at room open role + tagline + tags → one embedding → dot product vector cached on disk; every later draw is free lib/wordpick.py ~1 s · ~$0.00001 · no build step THE DRAW IS UNCHANGED wordpools.py keeps the CSPRNG and the apply-time draw — the pool is a better-authored bank behind the same door, so a secret still never enters the transcript before it is needed, and still cannot leak from a transcript that never held it. A persona is playable the moment it exists. That is the whole point.
fig 1 · the cost moves from per-persona to once-for-everybody. Nothing about the draw changes.

2 · What the build does, and what it costs

stagesourcewhat comes out
1 · English candidates Brysbaert concreteness ratings — 40k lemmas with concreteness, frequency and part-of-speech in one file nouns, concreteness ≥ 3.9, known by ≥ 90%, spoken often enough to matter
2 · curateone classification pass see §3 — this is the stage that makes the pool playable
3 · Chinese candidates translation of the curated English pool, plus the very top of THUOCL's food and animal lists translate, don't scrape — see §4
4 · curatesame pass, plus breed and brand names named explicitly
5 · pairsgenerated from the curated pool, so both halves inherit the curation no open-source pair list exists worth using — the best found was 101 English pairs. This is the one place generation still earns its cost, and it runs once for everybody rather than once per seat.
6 · judgeseam · symmetry · playability the rubric the word well paid for, reused whole
7 · embed768 dimensions, quantised to int8 int8 at scale 254 reproduces the float ranking's top-30 at 97% for half the bytes. Truncating the vectors instead does not work — 512 dims keeps only 66% and 256 keeps 46%, because the model has already projected once.

3 · The curation pass — the only reason this works

Concreteness ratings cannot make a pool playable on their own, and the reason is sharp: the worst possible secrets are perfectly concrete.

People rank highest of all. Search a pool for words relevant to "Chef & travel host" and the top hits are chef, traveler, journalist, entertainer — because a person is the most similar thing there is to a description of a role. They are also unguessable and unplayable. Every one of them scores 4-plus on concreteness, because a chef is a physical object.

Three classes die in the curation pass, and all three were found by looking at real ranked output rather than by guessing:

The pass keeps about 63% and runs in under a minute for the whole pool, once.

4 · Chinese has no foundation — so translate, don't scrape

English has a genuine free foundation. Chinese does not, and this was the sharpest finding of the whole exercise.

THUOCL has no everyday-objects category at all. Its lists are animals, food, cars, medicine, law, place names, idioms, poems, historical figures. Of the Chinese words the word well generated for Dieter Rams, 4% appear anywhere in its 156,289 entries. 椅子 · 手表 · 书架 · 剃须刀 — a chair, a watch, a bookshelf, a shaver — are simply absent, because no category covers them.

Ranking a designer against a pool of animals, food and cars produced exactly what you would expect:

苹果 欧宝 加去 新灰 欧姆贝 莱姆 德国黄牛 颐达 灰插 嘉好 全保 小灯 力存 黑漂
— Dieter Rams, ranked against THUOCL. Car models and noise.

Translating the concreteness-validated English pool fixes it, and the Chinese side inherits Brysbaert's validation for free:

家具 设备 椅子 机器 工厂 桌子 工具 书桌 沙发 皮革 杯子 橱柜 铅笔 床 木头 衣柜
— the same persona, the same query, after translation.

Attribution settles the recipe. Of the junk left in the top thirty, eight of twelve came from the THUOCL scrape and none from the translated layer. Only the very top of the food and animal lists is kept, for the words translation can never reach — 火锅, 豆腐, 包子 do not fall out of an English pool.

5 · The query is a world, not a name

Never put the persona's name in the query. Including it, 小胖 ranked 小馒头 · 小蜜蜂 · 小灯 · 小黑子 · 小萨摩 — the embedding matching the character — and in English pulled panda · pan · fang · boa off the sound of "Pang". The query is role + tagline + tags.txt, and nothing else. Every one of those fields exists on a persona the instant a user makes one, which is what keeps the per-persona cost at zero.

6 · What it picks

Three personas chosen to be as different as the roster allows, each ranked against the shipped pool. Nothing was generated for them; every word below was already in the pool before any of these characters were considered.

Anthony Bourdain — chef & travel host

poolthe top of the ranking
words.zh熟食 厨房 肉汁 狗肉 食谱 小厨房 游乐场 鸭肉 寿司 汤包 餐厅 巧克力 辣椒粉 肉丸 杂碎 烤肉 食堂 烤肉架 鹿肉 天妇罗 麻辣烫 热狗
pairs.zh麻辣烫/麻辣拌 · 葱段/蒜段 · 汤包/小笼包 · 鹿肉/牛肉 · 热狗/汉堡 · 墨西哥卷/三明治 · 烤肉卷/春卷 · 孜然/胡椒 · 鹅肉/鸭肉 · 油炸锅/蒸锅 · 鱼丸/肉丸 · 红油/酱油 · 面筋/豆腐 · 纸杯蛋糕/松饼 · 扣肉/红烧肉 · 糖霜/糖粉 · 天妇罗/炸鱼 · 水煮鱼/酸菜鱼 · 火锅/烧烤 · 白肉/红肉 · 玉米淀粉/土豆淀粉 · 午餐/晚餐
words.encookbook antipasto restaurant chocolate seasoning leftovers body meatball pastrami giblets sushi tempura hotdog parmesan dessert spices kitchen dumpling skewer tenderloin pork truffle
pairs.enhotdog/corn dog · snack/appetizer · sushi/sashimi · tofu/tempeh · meatball/falafel · spices/herbs · jambalaya/gumbo · spice/herb · beef/pork · salad/slaw · bagel/brownie · omelet/scrambled eggs · griddle/skillet · parsley/cilantro · kebab/satay · turkey/chicken · marinara/bolognese · fondue/raclette · horseradish/wasabi · dumpling/ravioli · toaster/sandwich press · cumin/caraway

Dieter Rams — industrial designer

poolthe top of the ranking
words.zh苹果 高脚椅 削皮器 椅子 咖啡壶 洗碗盆 音箱 扶手椅 木材 床头柜 脚凳 吧凳 苹果酒 折叠床 洗碗机 皮革 书桌 浴缸 洗衣篮 木棚 咖啡机 变速杆
pairs.zh高脚椅/吧台椅 · 把手/旋钮 · 削皮器/切菜器 · 椅子/凳子 · 工作台/办公桌 · 床头柜/边柜 · 烟灰缸/笔筒 · 长椅/板凳 · 肥皂盒/肥皂架 · 凹痕/凸起 · 垃圾桶/回收桶 · 苹果酒/苹果汁 · 折叠床/行军床 · 衣柜/书柜 · 工装裤/背带裤 · 注射器/输液器 · 羽绒床/弹簧床 · 吊灯/壁灯 · 水壶/茶壶 · 座位/坐垫 · 搪瓷/不锈钢 · 柜子/抽屉
words.enapple desk knoll drawer drawers sharpener cologne upholstery knob wastebasket dresser wood screwdriver chair planter bratwurst footstool prune sofa packaging notch plastic
pairs.endesk/table · tables/desks · peeler/grater · rack/shelf · workshop/studio · bin/trash can · bookcase/bookshelf · apple/pear · planter/flowerpot · titanium/aluminum · cot/crib · mug/cup · hanger/hook · corkscrew/bottle opener · clip/clamp · dishpan/washbasin · coffeepot/teapot · drill/screwdriver · folder/binder · shoehorn/shoe tree · cobalt/indigo · broomstick/mop handle

小胖 (Xiao Pang) — Shanghai property advisor

poolthe top of the ranking
words.zh房子 别墅 公寓 顶层公寓 豪宅 家园 地铁 地下室 联排别墅 平房 客厅 阁楼 庄园 屋顶 后楼梯 山顶 地窖 多宝鱼 楼梯间 商场 门挡 踢脚板
pairs.zh阁楼/地下室 · 平房/楼房 · 天窗/落地窗 · 宿舍/公寓 · 楼梯间/电梯间 · 后楼梯/前楼梯 · 窗帘/百叶窗 · 天花板/地板 · 暖气片/空调 · 休息室/客厅 · 蜂巢/蚁穴 · 汤包/小笼包 · 羽绒床/弹簧床 · 浴缸/淋浴房 · 火车/地铁 · 摩天楼/塔楼 · 打底裤/连裤袜 · 毛毛虫/蚯蚓 · 地堡/防空洞 · 门垫/地毯 · 超市/商场 · 羽绒被/蚕丝被
words.enhouse townhouse condo apartments apartment condominium residence cottage bungalow waterfront penthouse basement villa houseboat mansions skyscraper candlestick mansion thermostat summerhouse flooring showroom
pairs.enapartment/condo · townhouse/brownstone · waterfront/shoreline · mansion/villa · perch/bass · skylight/sunroof · baseboard/crown molding · farmhouse/cottage · crawlspace/basement · topsoil/subsoil · balcony/terrace · ceiling/floor · snapper/grouper · drain/shower · cesspool/septic tank · catfish/dogfish · shrew/mole · sinkhole/pothole · dominoes/mahjong · garage/carport · candle/thermostat · drainpipe/gutter
Residual junk runs around 15% — 楔子 · 讲坛 · welt · boa · nutshell among the words, house/home and bookcase/bookshelf (synonyms with no seam) among the pairs. Every one is a curation miss, which means every one is fixable in the shared build rather than per persona. That is the structural advantage: one fix improves every character at once.

7 · The owner's sweep — "not too niche, and not a broad category"

The owner eyeballed the personas starting with A and deleted words on one rationale, stated as two rules that are really one rule seen from both ends:

they have to be common ideas or objects, not too niche — and they can't be a very broad category such as 工具, humans are impossible to guess
A secret must be a thing whose NAME IS THE ANSWER. 工具 is a step toward an answer; 螺丝刀 is one. Ask it as a test: if the table said "is it X?" and heard YES, would the game be over? And a word half the table has never met fails the same test from the other side — nobody can converge on it.

The sweep, and the mistake it took to get right

The first attempt cut 16–20% and was wrong: it dropped 厨房 · 考拉 · 猎鹰 · 瓢虫 · mountain · dragon · paper. It had confused broad with not a small handheld object. The corrected test asks the only question that matters:

Could two players, both correct, be picturing completely different objects? "工具" — one pictures a hammer, another a screwdriver: drop. "考拉" — everyone pictures a koala: keep. It is a question about whether the word SPANS DIFFERENT OBJECTS. Not about size, and not about whether it fits in a hand.

With that, the sweep takes 4–5% of words and 10–14% of pairs — categories (配饰 · 装饰品 · 餐具 · appliance · attire · condiment), genuine niche (铃舌 · 雾号 · plinth · mollusk · plumage), and pairs that are merely synonyms (house/home · 草地/草坪 · port/harbor · 腊肠/香肠), which have no seam and so can never end a round.

Then the same rule again, measured across the whole roster

Per-word judgement cannot see the version of "too broad" that actually hurts. Rank the pool for all 105 personas and count how often each word lands in a top-30:

家园 ×75 · 图片 ×71 · 雪球 ×43 · 绿洲 ×36  |  globe ×59 · top ×59 · soapbox ×54 · nutshell ×40
— out of 105 personas, including a chef, a designer and a mathematician.

A word in the top-30 of a third of every character in the product is about none of them. Some are class words; most are semantic attractors — words near the centre of the embedding space, or ones a biography reaches for metaphorically (soapbox, nutshell, 绿洲, 家园) — and they crowd out what actually belongs.

The fix is NOT to delete them. 房子 is an attractor and the single best word in the pool for a property advisor; deleting it to help a chef robs the one persona it was made for. So the generic component is subtracted at ranking time instead — every word keeps whatever similarity it has above the average persona. The centroid of all 105 queries ships inside the pool, so this costs nothing at run time and nothing per persona.
worst cross-persona repeat, top-30beforeafter
Chinese75 / 105  (71%) 12 / 105  (11%)
English59 / 105  (56%) 16 / 105  (15%)

What that buys, persona by persona — 小胖 keeps 房子 at number one and gains 别墅 · 公寓 · 顶层公寓 · 联排别墅 · 平房 · 阁楼 · 屋顶; Rams goes from 装饰品 · 图片 · 基座 to 高脚椅 · 削皮器 · 咖啡壶 · 洗碗盆 · 扶手椅 · 床头柜 · 脚凳. The penalty is one number (MAD_WORDPICK_BREADTH, default 0.7); above ~0.9 it starts pulling in noise.

8 · The trade, stated plainly

Overlap between what this picks and what the word well generated is 27–36%. Two thirds of the generated words are unreachable from any generic pool, and they are the good ones: 槽钢层 · 半阳台 · 蛏刀 · 透明盖 came out of a persona's own documents and no shared list will ever hold them.

That is the price, and it is worth paying. A pool gives you a chef's words where generation gave you Bourdain's words — but it gives them to every persona ever made, instantly, for a hundred-thousandth of the cost. And the show does not live in the noun. If Bourdain is the one describing the oyster — the boat in France, the mud, the shucking — the vividness is already there and it was never in the word list.

The word well is not deleted. It stays as the deepening path for a persona somebody cares enough about to wait for, and its findings — the stopping-rule trap, the seating rule, the pair rubric — are what made this build possible in an afternoon.

9 · What is not done

  1. Nothing is wired into a game. This ships the pool and the picker; the 20Q and 卧底 overhaul is the next session's job, and it should read §6 and §7 before choosing a target size.
  2. The picker is unmeasured on real play. The samples below are eyeballed, not played. The word well's own lesson applies: the sims cannot find what only happens to real people taking real time.
  3. Relevance is topical, not characterful. Ranking cannot distinguish a word the persona owns from a word merely in their trade. If that turns out to matter at the table, the fix is a cheap re-rank of the top ~200 by one model call — still once per persona, still no build step.
  4. Licences want checking before this ships to production. THUOCL grants commercial use explicitly; the Brysbaert mirror states none; SUBTLEX-CH says research only and is not used here. The translated layer is our own output.
Dialogue · the word pool · built 2026-08-05 · lib/wordpick.py + lib/wordpool/ · siblings: the word well (the per-persona generator this replaces) · the game device · the dice specimen