One row a loop, generated by python exam/ink_loop.py ledger from <ink_dir>/loop/ledger.json. The plan and the reading of these numbers: the writing loop. Win = the share of blind pairs (both orders) the critic preferred the Ink piece over a length-and-class-matched major, with its 95% range in brackets — 24 verdicts a loop is a small sample, so two loops whose ranges overlap cannot honestly be told apart (added 09-11, the outside research §7 item 1); score = the critic's raw 0–10 for the Ink pieces and for the majors in the same sitting (the gap is the number; the majors are not all 9s — the bank's plain band carries academic Conversation pieces and a joke DMCA notice); spotted = how often the critic guessed which piece was the machine (50% would mean it cannot tell); agree = how often the two orders of the same pair agreed; over-used = words · sentence shapes still over-used against the bank at z ≥ 3 (the target is 0 · 0); tells = negation titles % · short flat「X is Y」sentences % · contractions /100 · words a paragraph (the majors: 5–10% · 3–3.5% · 1.7–2.5 · 50–69). The models line under each recipe (added 09-13) = the writer model id the loop ran on and the critic's own model id as written into its verdicts — a win rate is read against a writer version and a critic version, never across one; two loops on different models are not a before-and-after (the industry scan §4: a rented model's habits move with each release). Loops before 09-13 show — for both.
| loop | recipe · writer | the one change | win | score ink / majors | spotted | agree | over-used | tells | $ loop |
|---|---|---|---|---|---|---|---|---|---|
| L01 | r0 · staff — · critic — | baseline — the house rules as the brew reads them today, the staff writer instead of a persona | 37.5% (21–57) | 5.92 / 6.46 | 100.0% | 91.7% | 40 · 15 | 50.0% · 13.9% · 2.62 · 94.8 | 0.0945 |
| L02 | r1 · staff — · critic — | subtract the three causes in the brief: the everyday comparison, explain-then-judge, say-why-it-matters-to-you | 29.2% (15–49) | 5.92 / 6.62 | 100.0% | 91.7% | 40 · 14 | 8.3% · 11.2% · 2.61 · 98.3 | 0.0918 |
| L03 | r2 · staff — · critic — | the paragraph rule: stop on the information, say the point once, no announcing, no closing moral | 41.7% (24–61) | 6.38 / 6.5 | 100.0% | 83.3% | 40 · 7 | 16.7% · 9.5% · 1.92 · 79.3 | 0.0886 |
| L04 | r3 · staff — · critic — | examples: three whole human paragraphs instead of three openings | 45.8% (28–65) | 5.92 / 6.29 | 100.0% | 75.0% | 40 · 12 | 25.0% · 11.3% · 2.03 · 79.4 | 0.071 |
| L05 | r4 · staff — · critic — | the title last, from five candidates, turning on something that happened | 29.2% (15–49) | 6.21 / 6.67 | 100.0% | 75.0% | 40 · 9 | 0.0% · 12.4% · 1.83 · 81.6 | 0.1003 |
| L06 | r5 · staff — · critic — | one revision from an editor's three marks (the worst sentences, with what a person would say) | 33.3% (18–53) | 5.92 / 6.54 | 100.0% | 83.3% | 40 · 9 | 0.0% · 14.1% · 1.85 · 82.2 | 0.1009 |
| L07 | r6 · staff — · critic — | the notebook's NOT ESTABLISHED list is no longer shown to the writer | 41.7% (24–61) | 5.79 / 6.5 | 100.0% | 83.3% | 40 · 7 | 8.3% · 10.6% · 1.93 · 87.0 | 0.1041 |
| L08 | r7 · staff — · critic — | Notes and Scorecard written as prose: no numbering, no headings, no watch-list | 45.0% (26–66) | 6.55 / 6.45 | 100.0% | 90.0% | 40 · 6 | 10.0% · 12.1% · 1.68 · 86.2 | 0.1035 |
| L09 | r8 · staff — · critic — | no framing device: no owned object or remembered scene opening and closing the piece | 25.0% (12–45) | 5.67 / 6.96 | 100.0% | 100.0% | 40 · 11 | 8.3% · 13.9% · 1.83 · 78.5 | 0.1214 |
| L10 | r9 · staff — · critic — | the shape pass: the editor rewrites every short「X is Y」/「That is …」sentence to name its subject and carry its reason | 25.0% (12–45) | 5.83 / 6.83 | 100.0% | 100.0% | 40 · 7 | 16.7% · 7.0% · 2.37 · 76.2 | 0.129 |
| L11 | r10 · staff — · critic — | a spoken first draft (~300 words, said to a friend) before the piece | 22.7% (10–43) | 5.55 / 6.86 | 100.0% | 90.9% | 40 · 7 | 9.1% · 3.1% · 2.33 · 86.2 | 0.1288 |
| L12 | r11 · staff — · critic — | the explain line goes (no more「a term gets its meaning in passing」) + harness fixes: the staff writer never sees the persona's sentences in the question or the pitch, and the word notebook is gone from the prompt | 18.2% (7–39) | 5.09 / 6.86 | 100.0% | 100.0% | 40 · 9 | 9.1% · 6.7% · 2.16 · 70.4 | 0.1272 |
| L13 | r12 · staff — · critic — | the ending line: not a moral and not an arranged picture — when you have said it, stop | 20.8% (9–40) | 5.33 / 6.75 | 100.0% | 91.7% | 40 · 6 | 16.7% · 4.5% · 2.03 · 78.3 | 0.1322 |
| L14 | r12 · claude-subagent — · critic — | the ending line: not a moral and not an arranged picture — when you have said it, stop | 50.0% (31–69) | 6.71 / 6.21 | 100.0% | 66.7% | 40 · 13 | 8.3% · 7.2% · 0.81 · 91.8 | 0.0 |
| L15 | r13 · staff — · critic — | the length floors and ceilings halved: a piece stops when it has said it | 22.7% (10–43) | 5.18 / 6.55 | 100.0% | 90.9% | 40 · 4 | 9.1% · 3.0% · 2.19 · 62.2 | 0.0883 |
| L16 | r15 · staff — · critic — | keep the good, drop the stack: r2's rules + the title last + the unknowns withheld + Notes in prose; no revision, no shape pass, no spoken draft | 41.7% (24–61) | 5.75 / 6.29 | 100.0% | 100.0% | 40 · 9 | 25.0% · 12.7% · 2.63 · 85.3 | 0.0986 |
| L17 | r16 · staff — · critic — | the Gemini seat: r15 unchanged, the writer call on gemini-3.8-flash | 16.7% (7–36) | 5.75 / 7.04 | 100.0% | 100.0% | 40 · 10 | 16.7% · 3.9% · 1.04 · 95.0 | 0.0911 |
| L18 | r15 · staff — · critic — | THE SHAPE TEST — control: r15 re-rolled today, three drafts per commission; first drafts judged (the editor's output set aside, owner 09-11) | 29.2% (15–49) | 5.25 / 6.29 | 100.0% | 91.7% | 40 · 12 | 33.3% · 16.2% · 1.78 · 87.9 | 0.0 |
| L19 | r17 · staff — · critic — | THE SHAPE TEST — r17: r15 + one shape line with its turn, a different shape per draft; first drafts judged (the editor's output set aside) | 29.2% (15–49) | 5.12 / 6.17 | 100.0% | 91.7% | 40 · 11 | 25.0% · 10.4% · 2.22 · 82.4 | 0.0 |
| L20 | r18 · staff — · critic — | five versions with probabilities, the least likely kept (verbalized sampling); the editor off | 25.0% (12–45) | 4.0 / 6.46 | 100.0% | 100.0% | 40 · 15 | 41.7% · 23.2% · 1.07 · 71.4 | 0.0 |
exam/ink_loop.py · the pieces, pairs and verdicts stay under <ink_dir>/loop/, outside the repo