The writing loop — write like a person first, then put the persona back plan 09-10 · checkpoint 1 09-11 · §9

Owner (09-10), after reading the 09-09 and 09-10 editions:「humans don't speak or write this way」— seven tells named (negation titles · flat「X is Y. X isn't Z」sentences · the house vocabulary · metaphors that don't work · obscure, contrived language · streets that ask and cities that answer · numbers that sound awkward). The plan asked for: a study of our writing against good human writers, quantitative and granular; then a loop that takes the persona out and writes like a human writer until a critic agent rates it 8/10 against the majors' 9; then the persona back, with the quality held. V4.1 Flash makes the loop cheap; the cap is $6 of DeepSeek. This page is the plan; nothing on it is built except the first-cut count in §1.

1 · The study Ink vs the majors, by code and by reading the specimen catalogue: what works, what doesn't → the critic's rubric · the writer's example sheet 1 session · $0 DeepSeek 2 · Persona out a nameless staff writer, 12 frozen commissions one change a loop · the critic judges blind stop at 8/10 twice running, or at $4 spent 15–25 loops · 2–3 sessions 3 · Persona back same commissions, the profile in guard: the score holds · a point of view shows stop at 8/10 with the guard green, or at $6 5–10 loops · 1–2 sessions you read the catalogue + 10 blind pairs you read the 12 pieces + 10 blind pairs you read 12 + decide: to the brew?
Three steps, three reads. The owner never grades pieces inside a loop; at each checkpoint the owner reads ten blind pairs, and the critic's calls are checked against that read before the next step starts.
rulings reconciledOn 09-07 the ruling was「a model must not score a piece; no you-write-I-score loop」, and on 09-08「the bank is the target shape, not the target quality」. Today's ask puts a critic in a loop and names the majors as the target quality (9/10). The two fit together like this: the critic never scores in the brew (the production gate stays code, as it is); it judges only inside this loop, blind, against a real human piece of the same length and kind, never a piece alone on a scale; its number counts only after it passes an exam (§3) and only while it agrees with the owner's own blind reads at each checkpoint. The owner's time in the loop is three reads, about fifteen minutes each.

1 · The study — where Ink leaves the human page

Revised 09-10 after the owner's comments: the first draft verified a hand list (my regexes for the seven tells). The owner's point — find the structures and the words, don't verify mine — turns §1a into disproportion finders over the whole vocabulary and over sentence shapes, and §1b into open discovery before any catalogue. The numbers below are the second cut: 88 English pieces, every public edition 09-04 → 09-10 (79.5k words), against the 267-piece majors bank (338k words).

1a · By code — disproportion finders, not hand lists

The instrument is the log-odds ratio with an informative prior (the「fightin' words」method): every word, every two-word phrase, every sentence shape, and every title word is scored by how far its rate in Ink sits from its rate in the majors, z-scored so a rare word cannot win by luck. Nothing is named in advance. What comes out is not a list of words; it is a register, and it reads the same from four directions.

directionover-used by Ink (rate per 10k words, Ink → majors)under-used by Inkwhat it is
wordsthe 753 → 503 · is 225 → 108 · it 156 → 97 · whether 16 → 3.4 · cannot 12 → 2 · upon 11 → 1.2 · says 11 → 1.6 · remains 8 → 1 · merely 7 → 1.2 · exact 5.4 → 0.7 · machine 5.3 → 0.6 · engine 4.4 → 0.3 · single 7.7 → 1.8 · nobody 5.7 → 0.9 · entirely · completely 5 → 1.1 · plain 3.8 → 0.1 · genuine 3.3 → 0.3 · reader · headline · phrase · sentence · narrative (all 5–16×)but 32 → 61 · my 20 → 41 · it's 2.3 → 16.6 · about 20 → 40 · was 50 → 71 · had 18 → 30 · people 17 → 32 · like 18 → 30 · some · other · many · a lot (⅓–½) · think 4 → 10 · I'm · don't · you're (⅛–⅕) · kids · women · world · time (⅓–½) · for example 0 → 21 · in fact 0 → 18 · of coursePresent-tense, definite, certain, abstract.「The X is the Y.」Twice the is, half the but, half the was: Ink states, humans narrate and concede. The intensifiers of certainty (exact · entirely · every single · the whole · the true · the real · genuine · plain · merely · simply) are a performed confidence. Ink talks about text (reader · headline · phrase · sentence) — the brief's own vocabulary leaking. The formal register (whether · cannot · upon · must · shall · remains) is the machine's default English.
two-word phrasesthat is 22 → 2.6 · is the 33 → 8.5 · is a 29 → 11 · it is 20 → 7 · I have 14 → 4 · do not · does not · will not (5×) · whether the 5.7 → 0.2 · the whole · the exact · the real · the true · every single · here is · the question is · says the 2.8 → 0.06 · the machine 2.6 → 0.09I was 3 → 11 · I had 1.7 → 4.9 · if you 4.6 → 11 · he said 0.9 → 3.6 · a lot 1.8 → 5.8 · there are 1.9 → 6 · want to · to do · to get · to have (⅓) · such as 0.5 → 2.9 · the best 0.4 → 4The same register, one word wider: the copular claim (that is · is the · is a · it is) against the human's past-tense first person (I was · I had · he said) and the plain verbs of doing (want to · to get). No「the best」in 80k words: Ink does not rank, recommend or enthuse.
sentence shapes (opener × length × copula)「That is …」≤ 8 words, copular: 152 → 9 per 10k sentences (17×) · 「The … is …」≤ 8 words: 237 → 45 (5×) · 「The …」≤ 8 words: 348 → 65 · 「It …」≤ 8: 136 → 27 · 「A … is …」≤ 8: 47 → 4 · 「Let …」· 「Nobody …」· 「Here is …」a sentence opening on a named subject, 21+ words: 311 → 636 (½) · a named-subject sentence with a dash: 12 → 127 (⅒) · a question opening on a name: 10 → 55 · short sentences opening on anything but the / it / that / I: 70 → 239 (⅓) · sentences opening on and · but · so · in · as · for · at (⅓–½)This is the owner's「X is X. X isn't X」, found rather than assumed: the short copular declarative with a pronoun or「the」as its subject —「That is the whole game.」「The number is the easy thing.」— is the single most over-used shape in Ink, and the human sentence it replaces opens on a named subject, runs long, and is allowed a dash, a question, a conjunction. Paragraph endings are NOT a tell (30% vs 31% end on a short sentence).
sentence openers (first two words)That is 252 → 7 per 10k sentences · I have 135 → 24 · It is 141 → 36 · There is 62 → 15 · The report · Let me · Here is · That's the · I read · You cannot · Do not · The real · It showsThere are 4 → 42 · If you 43 → 83 · I was 14 → 43 · I think 2 → 17 · I don't 6 → 22 · For example 0 → 21 · In fact 0 → 18 · Of course 2 → 12 · And yet 2 → 14 · It's a 0 → 19 · The best 1 → 12Ink never says for example, in fact, of course, I think, it's a — the small spoken moves a person uses to concede, illustrate and soften. It opens instead on the verdict:「That is …」.
titlesa title that turns on a negation, denial, absence or failure (wide lexicon: not · never · nobody · wrong · lost · refuses · can't · without · quit · fails …): 28% → 10% · title words: the 2×, it 6×, not 6×, now, twoThe owner's ①, wider than my first count (11 → 25 of 88 titles once wrong · nobody · never · lost count). The lexicon still misses the frame (「Forty Smart Guys Picked the Wrong Team」turns on wrong; 「Nintendo Said Spring 2027 and Nobody Left」on nobody): the honest count is a reader's —「does this title turn on something not happening?」— which §1b's critic answers over every title in both corpora.
numbersten 6.7 → 1.4 · hundred 5.8 → 1.4 · percent 5 → 0.8 · spelled numbers 13.8 → 8.4 per 1k · digits 7.4 → 14.8 per 1kdigits, figures with a sourceThree habits, not one: (a) a figure written out where a person writes digits (「three hundred and one of them children」「three thousand to sixty-five hundred dollars」); (b) the round number as colour (「a hundred times」「ten years」「twenty years」— ten and hundred are 4–5× the human rate); (c) false precision — a specific number the writer could not have counted (「in the twenty seconds between points」「at four annual meetings now」), the owner's point ⑦. On the dev set (c) is checkable by code: a number in the piece that is not in the notebook or the peg is unsourced.

Reproducible: scratchpad/dispro.py (words · bigrams · shapes · openers · titles · number contexts → dispro.json); to become lib/ink_tells.py. The 09-09/09-10-only first cut stays on the ledger for comparison: sentence length matched (17.5 vs 17.9 words); contractions half; questions a third; exclamations none; dashes a fifth; metaphor nouns 2.5 vs 0.9–1.6 per 1k; disclaimers 0.45 vs 0.15 a piece; paragraphs 109 vs 50–69 words. The owner's box notes on 09-09/09-10 sit beside the numbers: Wolfram「ice says, ice wins」; Kurnit's retirement disclaimer「again and again」; Newton「I will not amuse the world with conjectures」; Lu Xun's zh original「城市问、街问、办公室开口」; Ang Lee「plainer, more approachable」; Rams reviewing a phone he never held (a slate rule, parked in §8).

1b · By reading — discovery first, then the catalogue

Counts say how much; they cannot say which metaphor works or why a sentence sounds like a machine, and they only find what they were told to look for. So the reading runs in two passes. Pass one is open: a critic agent (a Claude subagent, not the writer's family) reads 40 matched pairs — each Ink piece beside a major of the same class and beat — with one instruction: list every way the two differ that a reader would feel, in your own words, with the sentence that shows it. No dimensions are given. The notes from 40 pairs are clustered; the clusters that recur become the dimensions. Pass two files the catalogue under those dimensions: the human sentence, the Ink sentence, one line on why one works and the other doesn't. Two worked examples of the kind of entry it files:

metaphor that doesn't work · Ink, Buffett 09-10「There is no solvent counterparty around to hand you a check if the world ends.」
An insurance term explains a plain thought (nobody can pay you after the end of the world). The complex explains the simple. A person would say the plain thing and keep the insurance joke for one line, if at all.
metaphor that works · Paul Graham,「Life is Short」, the bank「Life is short, as everyone knows. When I was a kid I used to wonder about this. Is life actually short, or are we really complaining about its finiteness?」
A common phrase taken literally and then tested — the figure is already the reader's; the writer just turns it over. Nothing is decoded.
a thing speaks · Ink, Wolfram 09-10「the ice says」·「the ice wins」
The finding was that small icy bodies kept their original colour; the writer makes the ice a speaker to sound lively. The reader now has to translate back.
a person speaks · David Sedaris,「A Long Way Home」, the bank「I don't know how Instagram tagged me as a person who wants to watch this sort of thing, but it was right on.」
The liveliness is in the writer's own reaction, first person, one plain sentence, a question underneath. Nothing personified.

Are nine dimensions the major signals? No — nine was my seed list, and the owner's question is the right one. Pass one exists so the list is found. From the code above and from reading the 09-09/09-10 pieces, these are the dimensions I expect pass one to add to the nine (metaphor · sentence shape · vocabulary · numbers · things that speak · title · opening · ending · hedge):

candidate dimensionwhat a reader feelsthe trace in the code
the verdict voiceevery paragraph lands a judgment; nothing is wondered at, conceded or left open「That is …」17× · but ½ · I think · of course · in fact ≈ 0 · the certainty intensifiers
tense and timeeverything happens in an eternal present; no「I was」, no「he said」, no story told in orderis 2× · was · had · I was · he said ⅓–½
specificitygeneric detail dressed as particular (the fake scene:「I read the Guardian piece over coffee」); no kids, no women, no named placekids · women · people · world ⅓–½ · named-subject sentences ½
the explaining reflexterms glossed the reader knows; the「which means」after every fact; the piece talking about itself (the reader · this headline · the phrase)reader · headline · phrase · sentence · narrative 5–16×
digression and humoura person leaves the task — an aside, a joke, a memory that does not serve the argument; Ink never doesexclamations 0 · questions ⅓ · a lot · the best ≈ 0
the argument's arcthe thesis restated each paragraph; the ending that returns to the opening image; paragraphs of one uniform sizeparagraphs 109 vs 50–69 words; the landing is not a tell (30 vs 31%)
reader addressa generic「you」lectured at, against the human's「if you're …」that imagines one personif you · if you're · you're ⅕–½ · you cannot · do not
false precisiona number that pretends to a count nobody madeten · hundred 4–5× · unsourced numbers vs the notebook (dev set only)

Outputs of the study (one session, no DeepSeek spend): docs/ink-writing-gap.html = the record — the disproportion tables in full, pass one's clusters, and the catalogue, for the owner's read; lib/ink/critic.md = the critic's anchors (per dimension, three「this works」and three「this doesn't」); lib/ink/shelf-sentences.json = the writer's example sheet, whole human paragraphs keyed by dimension, so the brief can show instead of instruct. The NYT question: not needed to start — the plain band already has 106 pieces — but it is blogger-heavy (Graham · Urban · Housel · the Conversation); one subscribed week of NYT Opinion guest essays (~40 pieces, harvested in the owner's logged-in Chrome as the New Yorker was) would give the plain band the newsroom columnist it lacks. Worth ~$1; do it in week 1 if convenient.

2 · The instruments — a frozen dev set and a tells audit

rule · 2026-09-13 · re-baseline on model changeEvery count on this page is taken against a writer version and a critic version. When either changes — a new DeepSeek release under the same wire id, a different critic model, a seat moved to Gemini — the 12 fixed assignments are re-run and re-scored before any before-and-after is read, and the ledger row carries both model ids. Without it a "fix" can be a release note (the industry scan §4: Opus 5 doubled its dashes between releases). The live editions are read the same way, by era, on the over-use profile.

3 · The critic — how it judges, and how we know it judges well

Who. A Claude subagent spawned from the session (a different family from the writer; session spend, per the cost rule). DeepSeek Flash shadows the same pairs each loop; its agreement with Claude is reported, and if it holds at 80%+ over five loops Flash takes the bulk and Claude reads a sample. The writer's own family never judges alone.

How. Blind, bylines stripped. Each dev piece is set against a length- and class-matched major from the bank, in both orders; the critic returns the winner, a margin, a 0–10 on the rubric for each, and the losing piece's three worst sentences with the reason. Those sentences are the loop's gradient: the next change is written from them, not from a hunch.

the rubricwhat it asks
four reader linescan an ordinary reader follow it · is a person in it (a stake, a feeling, a voice) · does it make you want the next paragraph · does it say something you didn't have before
seven deductionsthe owner's tells, each anchored by §1b's specimens: the critic is shown three「works」and three「doesn't」per dimension, so「metaphor that doesn't work」means the same thing in every loop
the scalethe majors read 9.0 by construction: each loop the critic scores the 12 majors too, and every score is rescaled so their mean is 9.0. The win rate is the primary number (it cannot drift); the rescaled score is the one the owner asked for.
8/10 meansrescaled score ≥ 8.0 on the 12-piece dev set and a blind win rate ≥ 35% against the majors and every tell within its allowance. Two loops running.

The exam, before any number counts (re-run at every checkpoint):

testpairsmust
a major vs the same major with five tells injected (a negation title · two house words · one thing that speaks · one spelled number · one disclaimer)24pick the clean one ≥ 90%
a major vs an Ink piece of 09-09 / 09-1024pick the major ≥ 80% (the owner's own read of those editions says so)
the same pair, both ordersallagree ≥ 90% — position bias under 10%
the owner's ten blind pairs at each checkpoint10agree ≥ 75%, or the rubric is wrong and the loop stops until it is fixed

4 · One loop

12 commissions frozen: peg · notebook · form the same every loop write · recipe vN brief + process switches V4.1 Flash · ~5¢ tells audit code + two extractions allowance = the majors' p75 the critic 12 pairs vs majors, blind win rate · score · 3 worst lines one change written from the worst lines the ledger row keep or revert · $ to date vN+1 ~20¢ · ~10 min
A change is kept only if the win rate rose beyond the noise of 12 pairs (a sign test, both orders) and no tell left its allowance; otherwise it is reverted and the next candidate is tried. One change a loop, so the ledger says what worked.

What a loop may change — the parameters

The queue — candidate changes, ranked by the numbers

#changeaimed atwhy first
1paragraphs of 40–70 words; the title written last, from five candidates, one that turns on something happening, a verb, said to a friendthe wall · ①cheap, structural, the two biggest gaps by ratio
2the critic's three worst lines go back to the writer for ONE revision (the notes name the sentence and the reason)③ ④ ⑥the owner's actual ask: judge, then fix what the judge named — the editor's passes today are generic
3the shelf grows from three openings to three whole human paragraphs per dimension the piece is weak on — named subjects, past tense, a concession, a for examplethe register of §1ademonstration beats a rule (the ai-tone law); a word list is whack-a-mole, a register is shown
4lift the dash ban to the majors' rate; ask for questions and the odd exclamation where a person wouldthe house is stricter than any human writer on three counts
5a spoken first draft: tell it across the table in 400 words, then write it up② ⑤ ③the owner's line —「humans don't speak this way」— tested literally
6numbers: give the figure or leave it out; never「three times」as coloursmall, easy to count
7the disclaimer cut: no sentence about what the writer cannot know or did not seedisclaimersthe Kurnit / Newton / Rams note; the gate catches a third today
8the loop's own disproportion list (top 40 words and shapes, recomputed) shown to the editor, never to the writer③ and the shapesa listed word echoes when the writer sees it; the editor can see a list safely; the list is this week's, not a fossil
9thinking off · temperature 0.7 / 1.0 · Gemini 3.8 Flash on the writer seat (box arm)allthe cheapest test of all, but each resets the baseline — run once, late

5 · Step 2 and step 3 — the persona out, then back

Persona out. The writer is「a staff writer at Ink」— no biography, no exemplars of its own, first person allowed. The prompt drops from ~47k characters to under 10k (the profile was 72% of it), which is the point: what is left is the brief, the notebook, the shelf, the commission. The loop runs on that until 8/10 holds twice or $4 is spent. The recipe that gets there is the house recipe.

Persona back. The same 12 commissions, the same recipe, the profile in (the writer cut, as today). Two guards, both from the critic: the rescaled score must hold at 8.0 − 0.3 with every tell within allowance, and a point of view must show — shown a pair on the same commission, one by the staff writer and one by the persona, the critic must name the persona's ≥ 70% of the time and say what gave it away (a stance, a memory, a way of putting it). If the score drops, the loop treats the profile the way it treats the brief: what in the profile brings the tells back (the 09-07 finding: the analyst's prose in a profile is read as vocabulary). Stop at 8/10 with both guards green, or at $6 total.

6 · The bill and the clock

per loop, 12 piecesstep 2 (no profile)step 3 (profile in)
write (V4.1 Flash, thinking on, ~6k / ~13k tokens in, ~1.8k out)
one revision from the critic's notes
editor passes (up to three)
tells extractions (two Flash calls a piece) + the Flash shadow judge
DeepSeek a loop, off-peak (peak is 2×)~17¢~21¢
the critic (Claude subagent) — session, not the $61 subagent1 subagent
wall clock~10 min~12 min

$6 buys 25–30 loops off-peak; the plan needs 20–35. Every loop books to its own ledger leaf (app.ink_loop) so the console shows the $6 as it goes and the stop is a number, not a feeling. Heavy runs sit outside DeepSeek's peak (09–12, 14–18 CST).

7 · Checkpoints, and what I need from you now

  1. After the study: you read docs/ink-writing-gap.html and ten blind pairs; the critic's exam runs against your ten. Nothing loops before this.
  2. When step 2 reaches 8/10: you read the 12 staff-writer pieces and ten blind pairs. If they read like a person to you, the recipe is frozen and step 3 starts.
  3. When step 3 reaches 8/10: you read the 12 persona pieces; you decide whether the recipe goes into the brew for the next morning's edition.
to decideThe 8/10 — is「rescaled score ≥ 8.0 + win rate ≥ 35% + tells within allowance」the bar you mean, or do you want the win rate higher? ② The critic — a Claude subagent on session spend is the design; say if you would rather it were a paid API judge counted in the $6. ③ NYT — DONE 09-11: 74 Opinion pieces (129k words, read in the owner's logged-in Chrome one page every eight seconds after the human-check slider was cleared) sit in the plain band; the bank is 341 pieces and the shipped openings 208; the loop's bank counts re-measure on the next loop. ④ The staff writer — nameless, or a thin card (「40s, two kids, reads on the train」) so a person is in it even without a persona. ⑤ Dashes and questions — lift the one-dash rule to the human rate, yes or no. ⑥ The 09-04 loop page — this plan supersedes it (same idea, a cheaper writer, the critic named, the persona-out step added); mark it so.

8 · Not in this plan

9 · Checkpoint 1 — what sixteen loops say 2026-09-11 · for the owner's read

Built and run on 09-10/11 after the owner's go: the study (§1, the record is the writing gap), the harness (exam/ink_loop.py), the frozen dev set (12 commissions from the 09-09/09-10 box brews, each with a fixed major), sixteen loops of the staff writer, and one loop written by a Claude subagent on the same prompts. Every loop was judged blind by two Claude critics (both orders). DeepSeek spend: $1.58 of the $6. The full rows: the ledger.

looprecipethe one changewinraw ink / majorsgapspottednegation titles · flat「X is Y」· words a paragraph
L01r0baseline — the house rules as the brew reads them today, the staff writer instead of a persona37.5%5.92 / 6.460.54100.0%50.0% · 13.9% · 94.8
L02r1subtract the three causes in the brief: the everyday comparison, explain-then-judge, say-why-it-matters-to-you29.2%5.92 / 6.620.7100.0%8.3% · 11.2% · 98.3
L03r2the paragraph rule: stop on the information, say the point once, no announcing, no closing moral41.7%6.38 / 6.50.12100.0%16.7% · 9.5% · 79.3
L04r3examples: three whole human paragraphs instead of three openings45.8%5.92 / 6.290.37100.0%25.0% · 11.3% · 79.4
L05r4the title last, from five candidates, turning on something that happened29.2%6.21 / 6.670.46100.0%0.0% · 12.4% · 81.6
L06r5one revision from an editor's three marks (the worst sentences, with what a person would say)33.3%5.92 / 6.540.62100.0%0.0% · 14.1% · 82.2
L07r6the notebook's NOT ESTABLISHED list is no longer shown to the writer41.7%5.79 / 6.50.71100.0%8.3% · 10.6% · 87.0
L08r7Notes and Scorecard written as prose: no numbering, no headings, no watch-list45.0%6.55 / 6.45-0.1100.0%10.0% · 12.1% · 86.2
L09r8no framing device: no owned object or remembered scene opening and closing the piece25.0%5.67 / 6.961.29100.0%8.3% · 13.9% · 78.5
L10r9the shape pass: the editor rewrites every short「X is Y」/「That is …」sentence to name its subject and carry its 25.0%5.83 / 6.831.0100.0%16.7% · 7.0% · 76.2
L11r10a spoken first draft (~300 words, said to a friend) before the piece22.7%5.55 / 6.861.31100.0%9.1% · 3.1% · 86.2
L12r11the explain line goes (no more「a term gets its meaning in passing」) + harness fixes: the staff writer never se18.2%5.09 / 6.861.77100.0%9.1% · 6.7% · 70.4
L13r12the ending line: not a moral and not an arranged picture — when you have said it, stop20.8%5.33 / 6.751.42100.0%16.7% · 4.5% · 78.3
L14r12the ending line: not a moral and not an arranged picture — when you have said it, stop50.0%6.71 / 6.21-0.5100.0%8.3% · 7.2% · 91.8
L15r13the length floors and ceilings halved: a piece stops when it has said it22.7%5.18 / 6.551.37100.0%9.1% · 3.0% · 62.2
L16r15keep the good, drop the stack: r2's rules + the title last + the unknowns withheld + Notes in prose; no revisi41.7%5.75 / 6.290.54100.0%25.0% · 12.7% · 85.3

What moved, and what did not

The honest reading

With V4.1 Flash in the writer seat, the brief and the process are worth about half a point of the critic's ten, and no combination tried makes the piece pass as a person's. A Claude subagent on the same prompts scored above its matched majors (6.71 vs 6.21, the only loop to do so clearly) — the seat is worth about a point — and was still named as the machine in every pair, for the same reasons: the aphorism at every paragraph end, the announced move, the thesis restated, the coda. So the habits the owner hears are not one model's accent; they are how a language model writes an essay when asked for one, and neither rules nor passes nor a better model removed them in this round. The tells that are counted went down; the tell that is read did not. The staff-writer step has done its job as a measurement: it says the brief is a half-point lever, the seat a one-point lever, and the paragraph-landing habit the thing still standing.

to decide at checkpoint 1The writer seat. L14 says a stronger seat buys about a point (above the majors' raw score) but not invisibility; the choice is a price question — Gemini 3.8 Flash on the box (already an option on the console's writer card) or a Claude model by API at ~5–8¢ a piece — and worth an A/B on one morning's edition either way. ② Fewer passes on the box. The three copy-editor passes and the antithesis pass are the same machine-on-machine mechanism the loop found harmful; the plan is to A/B them off on one morning's edition (the gate still judges). Decided 09-11 and built without the A/B: the owner ruled no AI editing and no length top-up — every rewrite step is out of the gate and the writer. ③ The recipe to carry into the brew now: r2's paragraph rule and r7's Notes-in-prose are low-risk and measured; the title-last rule fixed the negation titles by count. The staff-writer line 「you are in it」 stays. ④ The habit itself. Since no rule moved the paragraph-landing / restating / announcing habit, the next thing to try is structural rather than verbal: a different form of ask — the piece dictated as reporting notes plus one opinion paragraph, or written in one sitting with no revision at all — measured on the same 12 pairs; and a blind read of ten pairs by the owner, so the critic's 100% is checked against a human's. ⑤ Step 3 (the persona back) waits on ① and ④: a profile in front of a seat that reads as a machine only adds a costume.
decided 09-11 · after the outside researchYes, and wider: no more AI editing — make the first draft right. Every AI rewrite step in the gate goes, measured on one morning's edition (the outside research §7, item 4). ③ Yes — carried as §7 item 9. Also agreed: no more rules in the prompt · generate a few and pick one · all eight changes in §7. ① the seat, ④'s blind read, and the rest of our own open items are proposed there as items 9–17, awaiting the owner's yes.

10 · The owner's blind read — 2026-09-11

Short answer: when it comes to spotting the machine, the critic and the owner agree completely — both picked it out in all ten pairs, even from ~250-word openings. When it comes to which piece is worth reading on, they disagree: the owner would rather keep reading the machine piece in 3 of 7 pairs, the critic almost never. So the critic's "spotted 100%" can be trusted; its quality verdicts can't stand in for the owner.

How it was run: ten pairs from the shape test (five from each arm, each on a different assignment; the pair whose human piece is a broken web page left out). Each piece cut to its opening, about 250 words, at a sentence break, titles kept, no bylines, order random. Two questions per pair: which would you rather keep reading, and which one a machine wrote. A blind critic answered the same two questions on the same excerpts. The page, the key and the answers: rooms-dev/_ink/loop/OWNER1/ (local only — the human pieces are copyrighted).

questionownercriticagree
Which one did a machine write?10 / 10 (9 by the buttons)10 / 1010 / 10
Would rather keep reading the human piece4 / 78 / 103 / 7

Pair 4: the owner's notes name the Hungary piece as the machine, and it was, but both buttons were pressed as if it were labelled A (it was B). The page, the critic's pack and the key were checked and agree, so it is counted by the notes. The owner skipped the first question on pairs 8–10.

The machine pieces the owner would rather keep reading

Two of the three are shaped pieces built on a memory — exactly the kind the critic marked down for "invented" scenes in the shape test. The owner and the critic value that memory in opposite directions.

What gave the machine away, in the owner's words

the tellthe example the owner quotedin the writing gap?
A title that turns on a negation, or has a dash in it"The planes were fine — the list broke"#5 · the dash is new
The writer reads the news from afar and comments — "not in the story""the piece landed" · "the headline said" · "I have read a lot of these" · "I read the list and laughed"#3, #20
"That is not X. That is Y.""That is not statecraft. That is a business decision"#5
Numbers recited, or precise where nobody caresthe county-by-county dengue count · "It cost eight dollars" · decimals#17
Filler that sounds wise and says nothing — "someone always talking and nothing meaningful comes out""usually nothing, occasionally everything, and invisible in between"#1
Showy sentence structure, "like a middle school student practising sentence structure in inappropriate places""it's a subtraction, and subtraction…"partly #16
The arranged last paragraph"and then a hundred miles of land happened, and then a town, and then…"the closing tableau
The word "honest""the last honest spec"#14 (costume words)

What the owner loved in the human pieces

"Very personal. Plain, comfortable language. Stories around himself." · "Good detail depiction — I've never seen our Ink articles illustrate the details to this degree." · "Fun to read, vivid, drawing readers into the pictures… how can we produce articles like those?" (Morgan Housel, Long-Term Money). · "The detailed story and fact writing read true."

what it changes1 · The critic stays, for one job: "does it read as a machine" — it matched the owner 10 for 10. 2 · The critic's scores and win rates stop counting as quality. On which piece is worth reading it agreed with the owner 3 times in 7, below the 3-in-4 bar this plan set in §3. Quality is the owner's read, taken more often and shorter (openings work). 3 · The memory question — DECIDED by the owner, 09-11: "Personas should be able to extrapolate from their own experience and memory to make the story more fun, more personal and convincing. But be discreet, don't overuse. And we need a disclaimer to say that it may not be his real experience." So: an imagined moment is allowed in a persona's piece, built from what the persona really lived and knew, used sparingly, and disclosed to the reader. Written into the extrapolation contract as Track F (Ink pieces; never in chat, the owner's second ruling). Built v1100, without a test (the owner: no need): the writer may imagine one moment and names it; the article reads “Includes an imagined moment” beside its read time; a chat opened from the piece carries the article as it is, with no special instruction (the owner, 09-11). 4 · The tells above go to the critic's examples and the writer's material (the outside research §7, items 8 and 15), in the owner's own words.
Plan · 2026-09-10 · for the owner's review · the numbers in §1a from scratchpad/tells.py over the 09-09 + 09-10 editions and rooms-dev/_ink/standard/ · earlier: the loop (09-04) · the first read · the register study · what the writer reads