The writing loop ledger ledger · 2026-09-13 22:39 · 20 loops

One row a loop, generated by python exam/ink_loop.py ledger from <ink_dir>/loop/ledger.json. The plan and the reading of these numbers: the writing loop. Win = the share of blind pairs (both orders) the critic preferred the Ink piece over a length-and-class-matched major, with its 95% range in brackets — 24 verdicts a loop is a small sample, so two loops whose ranges overlap cannot honestly be told apart (added 09-11, the outside research §7 item 1); score = the critic's raw 0–10 for the Ink pieces and for the majors in the same sitting (the gap is the number; the majors are not all 9s — the bank's plain band carries academic Conversation pieces and a joke DMCA notice); spotted = how often the critic guessed which piece was the machine (50% would mean it cannot tell); agree = how often the two orders of the same pair agreed; over-used = words · sentence shapes still over-used against the bank at z ≥ 3 (the target is 0 · 0); tells = negation titles % · short flat「X is Y」sentences % · contractions /100 · words a paragraph (the majors: 5–10% · 3–3.5% · 1.7–2.5 · 50–69). The models line under each recipe (added 09-13) = the writer model id the loop ran on and the critic's own model id as written into its verdicts — a win rate is read against a writer version and a critic version, never across one; two loops on different models are not a before-and-after (the industry scan §4: a rented model's habits move with each release). Loops before 09-13 show — for both.

looprecipe · writerthe one changewinscore ink / majorsspottedagreeover-usedtells$ loop
L01r0 · staff
— · critic —
baseline — the house rules as the brew reads them today, the staff writer instead of a persona37.5% (21–57)5.92 / 6.46100.0%91.7%40 · 1550.0% · 13.9% · 2.62 · 94.80.0945
L02r1 · staff
— · critic —
subtract the three causes in the brief: the everyday comparison, explain-then-judge, say-why-it-matters-to-you29.2% (15–49)5.92 / 6.62100.0%91.7%40 · 148.3% · 11.2% · 2.61 · 98.30.0918
L03r2 · staff
— · critic —
the paragraph rule: stop on the information, say the point once, no announcing, no closing moral41.7% (24–61)6.38 / 6.5100.0%83.3%40 · 716.7% · 9.5% · 1.92 · 79.30.0886
L04r3 · staff
— · critic —
examples: three whole human paragraphs instead of three openings45.8% (28–65)5.92 / 6.29100.0%75.0%40 · 1225.0% · 11.3% · 2.03 · 79.40.071
L05r4 · staff
— · critic —
the title last, from five candidates, turning on something that happened29.2% (15–49)6.21 / 6.67100.0%75.0%40 · 90.0% · 12.4% · 1.83 · 81.60.1003
L06r5 · staff
— · critic —
one revision from an editor's three marks (the worst sentences, with what a person would say)33.3% (18–53)5.92 / 6.54100.0%83.3%40 · 90.0% · 14.1% · 1.85 · 82.20.1009
L07r6 · staff
— · critic —
the notebook's NOT ESTABLISHED list is no longer shown to the writer41.7% (24–61)5.79 / 6.5100.0%83.3%40 · 78.3% · 10.6% · 1.93 · 87.00.1041
L08r7 · staff
— · critic —
Notes and Scorecard written as prose: no numbering, no headings, no watch-list45.0% (26–66)6.55 / 6.45100.0%90.0%40 · 610.0% · 12.1% · 1.68 · 86.20.1035
L09r8 · staff
— · critic —
no framing device: no owned object or remembered scene opening and closing the piece25.0% (12–45)5.67 / 6.96100.0%100.0%40 · 118.3% · 13.9% · 1.83 · 78.50.1214
L10r9 · staff
— · critic —
the shape pass: the editor rewrites every short「X is Y」/「That is …」sentence to name its subject and carry its reason25.0% (12–45)5.83 / 6.83100.0%100.0%40 · 716.7% · 7.0% · 2.37 · 76.20.129
L11r10 · staff
— · critic —
a spoken first draft (~300 words, said to a friend) before the piece22.7% (10–43)5.55 / 6.86100.0%90.9%40 · 79.1% · 3.1% · 2.33 · 86.20.1288
L12r11 · staff
— · critic —
the explain line goes (no more「a term gets its meaning in passing」) + harness fixes: the staff writer never sees the persona's sentences in the question or the pitch, and the word notebook is gone from the prompt18.2% (7–39)5.09 / 6.86100.0%100.0%40 · 99.1% · 6.7% · 2.16 · 70.40.1272
L13r12 · staff
— · critic —
the ending line: not a moral and not an arranged picture — when you have said it, stop20.8% (9–40)5.33 / 6.75100.0%91.7%40 · 616.7% · 4.5% · 2.03 · 78.30.1322
L14r12 · claude-subagent
— · critic —
the ending line: not a moral and not an arranged picture — when you have said it, stop50.0% (31–69)6.71 / 6.21100.0%66.7%40 · 138.3% · 7.2% · 0.81 · 91.80.0
L15r13 · staff
— · critic —
the length floors and ceilings halved: a piece stops when it has said it22.7% (10–43)5.18 / 6.55100.0%90.9%40 · 49.1% · 3.0% · 2.19 · 62.20.0883
L16r15 · staff
— · critic —
keep the good, drop the stack: r2's rules + the title last + the unknowns withheld + Notes in prose; no revision, no shape pass, no spoken draft41.7% (24–61)5.75 / 6.29100.0%100.0%40 · 925.0% · 12.7% · 2.63 · 85.30.0986
L17r16 · staff
— · critic —
the Gemini seat: r15 unchanged, the writer call on gemini-3.8-flash16.7% (7–36)5.75 / 7.04100.0%100.0%40 · 1016.7% · 3.9% · 1.04 · 95.00.0911
L18r15 · staff
— · critic —
THE SHAPE TEST — control: r15 re-rolled today, three drafts per commission; first drafts judged (the editor's output set aside, owner 09-11)29.2% (15–49)5.25 / 6.29100.0%91.7%40 · 1233.3% · 16.2% · 1.78 · 87.90.0
L19r17 · staff
— · critic —
THE SHAPE TEST — r17: r15 + one shape line with its turn, a different shape per draft; first drafts judged (the editor's output set aside)29.2% (15–49)5.12 / 6.17100.0%91.7%40 · 1125.0% · 10.4% · 2.22 · 82.40.0
L20r18 · staff
— · critic —
five versions with probabilities, the least likely kept (verbalized sampling); the editor off25.0% (12–45)4.0 / 6.46100.0%100.0%40 · 1541.7% · 23.2% · 1.07 · 71.40.0
Generated · the recipes (the weights) are in exam/ink_loop.py · the pieces, pairs and verdicts stay under <ink_dir>/loop/, outside the repo