Owner (09-10), after reading the 09-09 and 09-10 editions:「humans don't speak or write this way」— seven tells named (negation titles · flat「X is Y. X isn't Z」sentences · the house vocabulary · metaphors that don't work · obscure, contrived language · streets that ask and cities that answer · numbers that sound awkward). The plan asked for: a study of our writing against good human writers, quantitative and granular; then a loop that takes the persona out and writes like a human writer until a critic agent rates it 8/10 against the majors' 9; then the persona back, with the quality held. V4.1 Flash makes the loop cheap; the cap is $6 of DeepSeek. This page is the plan; nothing on it is built except the first-cut count in §1.
Revised 09-10 after the owner's comments: the first draft verified a hand list (my regexes for the seven tells). The owner's point — find the structures and the words, don't verify mine — turns §1a into disproportion finders over the whole vocabulary and over sentence shapes, and §1b into open discovery before any catalogue. The numbers below are the second cut: 88 English pieces, every public edition 09-04 → 09-10 (79.5k words), against the 267-piece majors bank (338k words).
The instrument is the log-odds ratio with an informative prior (the「fightin' words」method): every word, every two-word phrase, every sentence shape, and every title word is scored by how far its rate in Ink sits from its rate in the majors, z-scored so a rare word cannot win by luck. Nothing is named in advance. What comes out is not a list of words; it is a register, and it reads the same from four directions.
| direction | over-used by Ink (rate per 10k words, Ink → majors) | under-used by Ink | what it is |
|---|---|---|---|
| words | the 753 → 503 · is 225 → 108 · it 156 → 97 · whether 16 → 3.4 · cannot 12 → 2 · upon 11 → 1.2 · says 11 → 1.6 · remains 8 → 1 · merely 7 → 1.2 · exact 5.4 → 0.7 · machine 5.3 → 0.6 · engine 4.4 → 0.3 · single 7.7 → 1.8 · nobody 5.7 → 0.9 · entirely · completely 5 → 1.1 · plain 3.8 → 0.1 · genuine 3.3 → 0.3 · reader · headline · phrase · sentence · narrative (all 5–16×) | but 32 → 61 · my 20 → 41 · it's 2.3 → 16.6 · about 20 → 40 · was 50 → 71 · had 18 → 30 · people 17 → 32 · like 18 → 30 · some · other · many · a lot (⅓–½) · think 4 → 10 · I'm · don't · you're (⅛–⅕) · kids · women · world · time (⅓–½) · for example 0 → 21 · in fact 0 → 18 · of course ⅙ | Present-tense, definite, certain, abstract.「The X is the Y.」Twice the is, half the but, half the was: Ink states, humans narrate and concede. The intensifiers of certainty (exact · entirely · every single · the whole · the true · the real · genuine · plain · merely · simply) are a performed confidence. Ink talks about text (reader · headline · phrase · sentence) — the brief's own vocabulary leaking. The formal register (whether · cannot · upon · must · shall · remains) is the machine's default English. |
| two-word phrases | that is 22 → 2.6 · is the 33 → 8.5 · is a 29 → 11 · it is 20 → 7 · I have 14 → 4 · do not · does not · will not (5×) · whether the 5.7 → 0.2 · the whole · the exact · the real · the true · every single · here is · the question is · says the 2.8 → 0.06 · the machine 2.6 → 0.09 | I was 3 → 11 · I had 1.7 → 4.9 · if you 4.6 → 11 · he said 0.9 → 3.6 · a lot 1.8 → 5.8 · there are 1.9 → 6 · want to · to do · to get · to have (⅓) · such as 0.5 → 2.9 · the best 0.4 → 4 | The same register, one word wider: the copular claim (that is · is the · is a · it is) against the human's past-tense first person (I was · I had · he said) and the plain verbs of doing (want to · to get). No「the best」in 80k words: Ink does not rank, recommend or enthuse. |
| sentence shapes (opener × length × copula) | 「That is …」≤ 8 words, copular: 152 → 9 per 10k sentences (17×) · 「The … is …」≤ 8 words: 237 → 45 (5×) · 「The …」≤ 8 words: 348 → 65 · 「It …」≤ 8: 136 → 27 · 「A … is …」≤ 8: 47 → 4 · 「Let …」· 「Nobody …」· 「Here is …」 | a sentence opening on a named subject, 21+ words: 311 → 636 (½) · a named-subject sentence with a dash: 12 → 127 (⅒) · a question opening on a name: 10 → 55 · short sentences opening on anything but the / it / that / I: 70 → 239 (⅓) · sentences opening on and · but · so · in · as · for · at (⅓–½) | This is the owner's「X is X. X isn't X」, found rather than assumed: the short copular declarative with a pronoun or「the」as its subject —「That is the whole game.」「The number is the easy thing.」— is the single most over-used shape in Ink, and the human sentence it replaces opens on a named subject, runs long, and is allowed a dash, a question, a conjunction. Paragraph endings are NOT a tell (30% vs 31% end on a short sentence). |
| sentence openers (first two words) | That is 252 → 7 per 10k sentences · I have 135 → 24 · It is 141 → 36 · There is 62 → 15 · The report · Let me · Here is · That's the · I read · You cannot · Do not · The real · It shows | There are 4 → 42 · If you 43 → 83 · I was 14 → 43 · I think 2 → 17 · I don't 6 → 22 · For example 0 → 21 · In fact 0 → 18 · Of course 2 → 12 · And yet 2 → 14 · It's a 0 → 19 · The best 1 → 12 | Ink never says for example, in fact, of course, I think, it's a — the small spoken moves a person uses to concede, illustrate and soften. It opens instead on the verdict:「That is …」. |
| titles | a title that turns on a negation, denial, absence or failure (wide lexicon: not · never · nobody · wrong · lost · refuses · can't · without · quit · fails …): 28% → 10% · title words: the 2×, it 6×, not 6×, now, two | — | The owner's ①, wider than my first count (11 → 25 of 88 titles once wrong · nobody · never · lost count). The lexicon still misses the frame (「Forty Smart Guys Picked the Wrong Team」turns on wrong; 「Nintendo Said Spring 2027 and Nobody Left」on nobody): the honest count is a reader's —「does this title turn on something not happening?」— which §1b's critic answers over every title in both corpora. |
| numbers | ten 6.7 → 1.4 · hundred 5.8 → 1.4 · percent 5 → 0.8 · spelled numbers 13.8 → 8.4 per 1k · digits 7.4 → 14.8 per 1k | digits, figures with a source | Three habits, not one: (a) a figure written out where a person writes digits (「three hundred and one of them children」「three thousand to sixty-five hundred dollars」); (b) the round number as colour (「a hundred times」「ten years」「twenty years」— ten and hundred are 4–5× the human rate); (c) false precision — a specific number the writer could not have counted (「in the twenty seconds between points」「at four annual meetings now」), the owner's point ⑦. On the dev set (c) is checkable by code: a number in the piece that is not in the notebook or the peg is unsourced. |
Reproducible: scratchpad/dispro.py (words · bigrams · shapes · openers · titles · number contexts → dispro.json); to become lib/ink_tells.py. The 09-09/09-10-only first cut stays on the ledger for comparison: sentence length matched (17.5 vs 17.9 words); contractions half; questions a third; exclamations none; dashes a fifth; metaphor nouns 2.5 vs 0.9–1.6 per 1k; disclaimers 0.45 vs 0.15 a piece; paragraphs 109 vs 50–69 words. The owner's box notes on 09-09/09-10 sit beside the numbers: Wolfram「ice says, ice wins」; Kurnit's retirement disclaimer「again and again」; Newton「I will not amuse the world with conjectures」; Lu Xun's zh original「城市问、街问、办公室开口」; Ang Lee「plainer, more approachable」; Rams reviewing a phone he never held (a slate rule, parked in §8).
Counts say how much; they cannot say which metaphor works or why a sentence sounds like a machine, and they only find what they were told to look for. So the reading runs in two passes. Pass one is open: a critic agent (a Claude subagent, not the writer's family) reads 40 matched pairs — each Ink piece beside a major of the same class and beat — with one instruction: list every way the two differ that a reader would feel, in your own words, with the sentence that shows it. No dimensions are given. The notes from 40 pairs are clustered; the clusters that recur become the dimensions. Pass two files the catalogue under those dimensions: the human sentence, the Ink sentence, one line on why one works and the other doesn't. Two worked examples of the kind of entry it files:
Are nine dimensions the major signals? No — nine was my seed list, and the owner's question is the right one. Pass one exists so the list is found. From the code above and from reading the 09-09/09-10 pieces, these are the dimensions I expect pass one to add to the nine (metaphor · sentence shape · vocabulary · numbers · things that speak · title · opening · ending · hedge):
| candidate dimension | what a reader feels | the trace in the code |
|---|---|---|
| the verdict voice | every paragraph lands a judgment; nothing is wondered at, conceded or left open | 「That is …」17× · but ½ · I think · of course · in fact ≈ 0 · the certainty intensifiers |
| tense and time | everything happens in an eternal present; no「I was」, no「he said」, no story told in order | is 2× · was · had · I was · he said ⅓–½ |
| specificity | generic detail dressed as particular (the fake scene:「I read the Guardian piece over coffee」); no kids, no women, no named place | kids · women · people · world ⅓–½ · named-subject sentences ½ |
| the explaining reflex | terms glossed the reader knows; the「which means」after every fact; the piece talking about itself (the reader · this headline · the phrase) | reader · headline · phrase · sentence · narrative 5–16× |
| digression and humour | a person leaves the task — an aside, a joke, a memory that does not serve the argument; Ink never does | exclamations 0 · questions ⅓ · a lot · the best ≈ 0 |
| the argument's arc | the thesis restated each paragraph; the ending that returns to the opening image; paragraphs of one uniform size | paragraphs 109 vs 50–69 words; the landing is not a tell (30 vs 31%) |
| reader address | a generic「you」lectured at, against the human's「if you're …」that imagines one person | if you · if you're · you're ⅕–½ · you cannot · do not 5× |
| false precision | a number that pretends to a count nobody made | ten · hundred 4–5× · unsourced numbers vs the notebook (dev set only) |
Outputs of the study (one session, no DeepSeek spend): docs/ink-writing-gap.html = the record — the disproportion tables in full, pass one's clusters, and the catalogue, for the owner's read; lib/ink/critic.md = the critic's anchors (per dimension, three「this works」and three「this doesn't」); lib/ink/shelf-sentences.json = the writer's example sheet, whole human paragraphs keyed by dimension, so the brief can show instead of instruct. The NYT question: not needed to start — the plain band already has 106 pieces — but it is blogger-heavy (Graham · Urban · Housel · the Conversation); one subscribed week of NYT Opinion guest essays (~40 pieces, harvested in the owner's logged-in Chrome as the New Yorker was) would give the plain band the newsroom columnist it lacks. Worth ~$1; do it in week 1 if convenient.
lib/ink_tells.py, code): the disproportion finders of §1a run every loop over the 12 pieces against the bank — words · phrases · sentence shapes · openers · titles · numbers — so the list of what is over-used is recomputed, never hand-kept (the answer to whack-a-mole: when honest is cured and candid appears, the finder sees it the same night). Each finding carries the majors' rate as its allowance; a piece is「within」when nothing in it sits over the majors' p75 at z ≥ 3. Three tells need a reader to count honestly and use the extraction pattern proven for 隐喻率 — a Flash call that must quote the expression and give the plain replacement, never a score: things that speak · metaphors that don't work · a title that turns on something not happening. The numbers tell is code: a figure not in the notebook or the peg is unsourced; ten · hundred · times as colour are counted apart.docs/ink-writing-loop-ledger.html, generated like the brew log): one row a loop — the change, the tells table, the win rate, the calibrated score, the critic's three notes, dollars to date.Who. A Claude subagent spawned from the session (a different family from the writer; session spend, per the cost rule). DeepSeek Flash shadows the same pairs each loop; its agreement with Claude is reported, and if it holds at 80%+ over five loops Flash takes the bulk and Claude reads a sample. The writer's own family never judges alone.
How. Blind, bylines stripped. Each dev piece is set against a length- and class-matched major from the bank, in both orders; the critic returns the winner, a margin, a 0–10 on the rubric for each, and the losing piece's three worst sentences with the reason. Those sentences are the loop's gradient: the next change is written from them, not from a hunch.
| the rubric | what it asks |
|---|---|
| four reader lines | can an ordinary reader follow it · is a person in it (a stake, a feeling, a voice) · does it make you want the next paragraph · does it say something you didn't have before |
| seven deductions | the owner's tells, each anchored by §1b's specimens: the critic is shown three「works」and three「doesn't」per dimension, so「metaphor that doesn't work」means the same thing in every loop |
| the scale | the majors read 9.0 by construction: each loop the critic scores the 12 majors too, and every score is rescaled so their mean is 9.0. The win rate is the primary number (it cannot drift); the rescaled score is the one the owner asked for. |
| 8/10 means | rescaled score ≥ 8.0 on the 12-piece dev set and a blind win rate ≥ 35% against the majors and every tell within its allowance. Two loops running. |
The exam, before any number counts (re-run at every checkpoint):
| test | pairs | must |
|---|---|---|
| a major vs the same major with five tells injected (a negation title · two house words · one thing that speaks · one spelled number · one disclaimer) | 24 | pick the clean one ≥ 90% |
| a major vs an Ink piece of 09-09 / 09-10 | 24 | pick the major ≥ 80% (the owner's own read of those editions says so) |
| the same pair, both orders | all | agree ≥ 90% — position bias under 10% |
| the owner's ten blind pairs at each checkpoint | 10 | agree ≥ 75%, or the rubric is wrong and the loop stops until it is fixed |
| # | change | aimed at | why first |
|---|---|---|---|
| 1 | paragraphs of 40–70 words; the title written last, from five candidates, one that turns on something happening, a verb, said to a friend | the wall · ① | cheap, structural, the two biggest gaps by ratio |
| 2 | the critic's three worst lines go back to the writer for ONE revision (the notes name the sentence and the reason) | ③ ④ ⑥ | the owner's actual ask: judge, then fix what the judge named — the editor's passes today are generic |
| 3 | the shelf grows from three openings to three whole human paragraphs per dimension the piece is weak on — named subjects, past tense, a concession, a for example | the register of §1a | demonstration beats a rule (the ai-tone law); a word list is whack-a-mole, a register is shown |
| 4 | lift the dash ban to the majors' rate; ask for questions and the odd exclamation where a person would | ⑤ | the house is stricter than any human writer on three counts |
| 5 | a spoken first draft: tell it across the table in 400 words, then write it up | ② ⑤ ③ | the owner's line —「humans don't speak this way」— tested literally |
| 6 | numbers: give the figure or leave it out; never「three times」as colour | ⑦ | small, easy to count |
| 7 | the disclaimer cut: no sentence about what the writer cannot know or did not see | disclaimers | the Kurnit / Newton / Rams note; the gate catches a third today |
| 8 | the loop's own disproportion list (top 40 words and shapes, recomputed) shown to the editor, never to the writer | ③ and the shapes | a listed word echoes when the writer sees it; the editor can see a list safely; the list is this week's, not a fossil |
| 9 | thinking off · temperature 0.7 / 1.0 · Gemini 3.8 Flash on the writer seat (box arm) | all | the cheapest test of all, but each resets the baseline — run once, late |
Persona out. The writer is「a staff writer at Ink」— no biography, no exemplars of its own, first person allowed. The prompt drops from ~47k characters to under 10k (the profile was 72% of it), which is the point: what is left is the brief, the notebook, the shelf, the commission. The loop runs on that until 8/10 holds twice or $4 is spent. The recipe that gets there is the house recipe.
Persona back. The same 12 commissions, the same recipe, the profile in (the writer cut, as today). Two guards, both from the critic: the rescaled score must hold at 8.0 − 0.3 with every tell within allowance, and a point of view must show — shown a pair on the same commission, one by the staff writer and one by the persona, the critic must name the persona's ≥ 70% of the time and say what gave it away (a stance, a memory, a way of putting it). If the score drops, the loop treats the profile the way it treats the brief: what in the profile brings the tells back (the 09-07 finding: the analyst's prose in a profile is read as vocabulary). Stop at 8/10 with both guards green, or at $6 total.
| per loop, 12 pieces | step 2 (no profile) | step 3 (profile in) |
|---|---|---|
| write (V4.1 Flash, thinking on, ~6k / ~13k tokens in, ~1.8k out) | 3¢ | 6¢ |
| one revision from the critic's notes | 3¢ | 4¢ |
| editor passes (up to three) | 6¢ | 6¢ |
| tells extractions (two Flash calls a piece) + the Flash shadow judge | 5¢ | 5¢ |
| DeepSeek a loop, off-peak (peak is 2×) | ~17¢ | ~21¢ |
| the critic (Claude subagent) — session, not the $6 | 1 subagent | 1 subagent |
| wall clock | ~10 min | ~12 min |
$6 buys 25–30 loops off-peak; the plan needs 20–35. Every loop books to its own ledger leaf (app.ink_loop) so the console shows the $6 as it goes and the stop is a number, not a feeling. Heavy runs sit outside DeepSeek's peak (09–12, 14–18 CST).
docs/ink-writing-gap.html and ten blind pairs; the critic's exam runs against your ten. Nothing loops before this.validate_slate, outside this loop.Built and run on 09-10/11 after the owner's go: the study (§1, the record is the writing gap), the harness (exam/ink_loop.py), the frozen dev set (12 commissions from the 09-09/09-10 box brews, each with a fixed major), sixteen loops of the staff writer, and one loop written by a Claude subagent on the same prompts. Every loop was judged blind by two Claude critics (both orders). DeepSeek spend: $1.58 of the $6. The full rows: the ledger.
| loop | recipe | the one change | win | raw ink / majors | gap | spotted | negation titles · flat「X is Y」· words a paragraph |
|---|---|---|---|---|---|---|---|
| L01 | r0 | baseline — the house rules as the brew reads them today, the staff writer instead of a persona | 37.5% | 5.92 / 6.46 | 0.54 | 100.0% | 50.0% · 13.9% · 94.8 |
| L02 | r1 | subtract the three causes in the brief: the everyday comparison, explain-then-judge, say-why-it-matters-to-you | 29.2% | 5.92 / 6.62 | 0.7 | 100.0% | 8.3% · 11.2% · 98.3 |
| L03 | r2 | the paragraph rule: stop on the information, say the point once, no announcing, no closing moral | 41.7% | 6.38 / 6.5 | 0.12 | 100.0% | 16.7% · 9.5% · 79.3 |
| L04 | r3 | examples: three whole human paragraphs instead of three openings | 45.8% | 5.92 / 6.29 | 0.37 | 100.0% | 25.0% · 11.3% · 79.4 |
| L05 | r4 | the title last, from five candidates, turning on something that happened | 29.2% | 6.21 / 6.67 | 0.46 | 100.0% | 0.0% · 12.4% · 81.6 |
| L06 | r5 | one revision from an editor's three marks (the worst sentences, with what a person would say) | 33.3% | 5.92 / 6.54 | 0.62 | 100.0% | 0.0% · 14.1% · 82.2 |
| L07 | r6 | the notebook's NOT ESTABLISHED list is no longer shown to the writer | 41.7% | 5.79 / 6.5 | 0.71 | 100.0% | 8.3% · 10.6% · 87.0 |
| L08 | r7 | Notes and Scorecard written as prose: no numbering, no headings, no watch-list | 45.0% | 6.55 / 6.45 | -0.1 | 100.0% | 10.0% · 12.1% · 86.2 |
| L09 | r8 | no framing device: no owned object or remembered scene opening and closing the piece | 25.0% | 5.67 / 6.96 | 1.29 | 100.0% | 8.3% · 13.9% · 78.5 |
| L10 | r9 | the shape pass: the editor rewrites every short「X is Y」/「That is …」sentence to name its subject and carry its | 25.0% | 5.83 / 6.83 | 1.0 | 100.0% | 16.7% · 7.0% · 76.2 |
| L11 | r10 | a spoken first draft (~300 words, said to a friend) before the piece | 22.7% | 5.55 / 6.86 | 1.31 | 100.0% | 9.1% · 3.1% · 86.2 |
| L12 | r11 | the explain line goes (no more「a term gets its meaning in passing」) + harness fixes: the staff writer never se | 18.2% | 5.09 / 6.86 | 1.77 | 100.0% | 9.1% · 6.7% · 70.4 |
| L13 | r12 | the ending line: not a moral and not an arranged picture — when you have said it, stop | 20.8% | 5.33 / 6.75 | 1.42 | 100.0% | 16.7% · 4.5% · 78.3 |
| L14 | r12 | the ending line: not a moral and not an arranged picture — when you have said it, stop | 50.0% | 6.71 / 6.21 | -0.5 | 100.0% | 8.3% · 7.2% · 91.8 |
| L15 | r13 | the length floors and ceilings halved: a piece stops when it has said it | 22.7% | 5.18 / 6.55 | 1.37 | 100.0% | 9.1% · 3.0% · 62.2 |
| L16 | r15 | keep the good, drop the stack: r2's rules + the title last + the unknowns withheld + Notes in prose; no revisi | 41.7% | 5.75 / 6.29 | 0.54 | 100.0% | 25.0% · 12.7% · 85.3 |
With V4.1 Flash in the writer seat, the brief and the process are worth about half a point of the critic's ten, and no combination tried makes the piece pass as a person's. A Claude subagent on the same prompts scored above its matched majors (6.71 vs 6.21, the only loop to do so clearly) — the seat is worth about a point — and was still named as the machine in every pair, for the same reasons: the aphorism at every paragraph end, the announced move, the thesis restated, the coda. So the habits the owner hears are not one model's accent; they are how a language model writes an essay when asked for one, and neither rules nor passes nor a better model removed them in this round. The tells that are counted went down; the tell that is read did not. The staff-writer step has done its job as a measurement: it says the brief is a half-point lever, the seat a one-point lever, and the paragraph-landing habit the thing still standing.
Short answer: when it comes to spotting the machine, the critic and the owner agree completely — both picked it out in all ten pairs, even from ~250-word openings. When it comes to which piece is worth reading on, they disagree: the owner would rather keep reading the machine piece in 3 of 7 pairs, the critic almost never. So the critic's "spotted 100%" can be trusted; its quality verdicts can't stand in for the owner.
How it was run: ten pairs from the shape test (five from each arm, each on a different assignment; the pair whose human piece is a broken web page left out). Each piece cut to its opening, about 250 words, at a sentence break, titles kept, no bylines, order random. Two questions per pair: which would you rather keep reading, and which one a machine wrote. A blind critic answered the same two questions on the same excerpts. The page, the key and the answers: rooms-dev/_ink/loop/OWNER1/ (local only — the human pieces are copyrighted).
| question | owner | critic | agree |
|---|---|---|---|
| Which one did a machine write? | 10 / 10 (9 by the buttons) | 10 / 10 | 10 / 10 |
| Would rather keep reading the human piece | 4 / 7 | 8 / 10 | 3 / 7 |
Pair 4: the owner's notes name the Hungary piece as the machine, and it was, but both buttons were pressed as if it were labelled A (it was B). The page, the critic's pack and the key were checked and agree, so it is counted by the notes. The owner skipped the first question on pairs 8–10.
Two of the three are shaped pieces built on a memory — exactly the kind the critic marked down for "invented" scenes in the shape test. The owner and the critic value that memory in opposite directions.
| the tell | the example the owner quoted | in the writing gap? |
|---|---|---|
| A title that turns on a negation, or has a dash in it | "The planes were fine — the list broke" | #5 · the dash is new |
| The writer reads the news from afar and comments — "not in the story" | "the piece landed" · "the headline said" · "I have read a lot of these" · "I read the list and laughed" | #3, #20 |
| "That is not X. That is Y." | "That is not statecraft. That is a business decision" | #5 |
| Numbers recited, or precise where nobody cares | the county-by-county dengue count · "It cost eight dollars" · decimals | #17 |
| Filler that sounds wise and says nothing — "someone always talking and nothing meaningful comes out" | "usually nothing, occasionally everything, and invisible in between" | #1 |
| Showy sentence structure, "like a middle school student practising sentence structure in inappropriate places" | "it's a subtraction, and subtraction…" | partly #16 |
| The arranged last paragraph | "and then a hundred miles of land happened, and then a town, and then…" | the closing tableau |
| The word "honest" | "the last honest spec" | #14 (costume words) |
"Very personal. Plain, comfortable language. Stories around himself." · "Good detail depiction — I've never seen our Ink articles illustrate the details to this degree." · "Fun to read, vivid, drawing readers into the pictures… how can we produce articles like those?" (Morgan Housel, Long-Term Money). · "The detailed story and fact writing read true."
scratchpad/tells.py over the 09-09 + 09-10 editions and rooms-dev/_ink/standard/ · earlier: the loop (09-04) · the first read · the register study · what the writer reads