The process as it stands (settled 09-13): four drafts, the ruler sorts them, the lowest wins, the gate marks the winner. You named the flaw: a ruler that counts defects cannot be beaten by more whole drafts, because a long piece always carries many. The only way down is to select at the grain where the defects live, the paragraph. This page draws the two ways to do that, one you proposed and one I proposed, with the arithmetic, the cost, and what each one risks. Nothing is built yet.
Short answer: picking the cleanest of four whole drafts removes about a third of the defects. Picking the cleanest of four candidates at every paragraph removes nearly all of them.
| p | N | one draft | pick of N drafts (today) | A · repair | B · sequential | paragraph calls: A / B |
|---|---|---|---|---|---|---|
| 0.3 | 4 | 4.5 | 2.7 | 0.02 | 0.12 | ~12 / 60 |
| 0.4 | 4 | 6.0 | 4.1 | 0.10 | 0.38 | ~16 / 60 |
| 0.4 | 8 | 6.0 | 3.4 | 0.00 | 0.01 | ~24 / 120 |
| 0.5 | 4 | 7.5 | 5.5 | 0.34 | 0.93 | ~24 / 60 |
| 0.5 | 8 | 7.5 | 4.8 | 0.02 | 0.06 | ~40 / 120 |
Expected counted defects per piece. The middle column is your law of large numbers: doubling the drafts from four to eight buys less than one defect. The two right columns are the paragraph grain. A comes out ahead of B on this table only because its spine has already been selected once, so it has fewer paragraphs to re-roll; the two converge on the same floor.
Short answer: keep today's pick as the spine, run the ruler paragraph by paragraph, and re-roll the failing paragraphs where they stand, several candidates each, the cleanest kept.
What it risks. A seam: the new ¶3 must lead into the old ¶4, which was written after a different ¶3. Giving the candidate ¶4 to connect to is the mitigation; your blind read is the check. And a repaired paragraph can be clean and empty, the wire-round-up failure, which no counter sees.
Short answer: write the first paragraph several times, keep the best; give the task and the kept paragraph back, write the second several times, keep the best; continue to the end. Your proposal.
What it risks. Three things, all structural. The arc: nothing in fifteen local choices builds toward a turn, and a defect-only ruler prefers the safe paragraph at every step. The closer: a model asked for "the next paragraph" writes each one as an ending; the gavel key catches the sentence, but the habit is manufactured by the loop. The tie: once candidates are clean, the ruler is silent and the choice is random unless a positive key exists.
B has no arc without one. A has the spine's, but a repaired paragraph should know where the piece is going too. One call before the writing: the piece's beats, five or six lines, what happens in what order and where it turns. They enter the writer's prompt as material, beside the interview answers, never as an instruction on how to write. This is the shape menu's idea done by the writer itself rather than picked from our list; the menu failed as a rule, and this is the test of whether it works as material.
Both algorithms end in ties: several clean candidates and a ruler with nothing to say. Today the tie falls to the lint score, which is zero for all of them. The candidate to prefer is the one that uses the most of what it was given: names, numbers, places and quotes from the notebook and the interview that appear in the paragraph, counted by the same anchor matching the brew already does for the dealt hand. It is the first key in the pipeline that rewards what a paragraph has rather than what it lacks, and it is the only defence against the clean-and-empty paragraph. A count, not a judgment; no model scores anything.
| A · repair | B · sequential | |
|---|---|---|
| selection acts on | the failing paragraphs of a chosen draft | every paragraph |
| expected defects, p = 0.4, N = 4 | 0.10 | 0.38 |
| paragraph calls a piece | ~16 | ~60 |
| added cost a piece, Flash off-peak | ~0.3¢ | ~1¢ |
| time a piece | one round, ~15 s | 15 rounds, 2–4 min |
| coherence | by conditioning; one seam per repair | structural; no seams |
| the arc | the spine's, kept | none unless the beats are given |
| the ending | the spine's | the writer must flag it; the loop invites closers |
| what it changes in the brew | a pass after the pick | replaces the drafts and the pick |
| where it fails | a clean, empty repair; a bad seam | a clean, shapeless piece; a closer every paragraph |
| how it can grow | N up; re-roll borderline paragraphs too | keep two prefixes alive instead of one |
Short answer: build both as loop recipes, run them on the twelve fixed assignments against today's pick, count, and then read blind.
| run | what | cost |
|---|---|---|
| baseline | the four-draft pick as it runs today | ~50¢ |
| A | baseline + the repair pass, N = 4 | ~55¢ |
| B | sequential, N = 4, with the beats | ~65¢ |
By count: defects per piece by the ruler, over-use excess, gavel share, contraction rate, words, cost, minutes. Then your blind read of the three versions of the same assignment, unlabelled, for the two things no counter sees: does the piece go somewhere, and is anything under the clean sentences. The critic's pairwise read against the majors can follow if the blind read is promising; it is not the decider here.
Short answer: you chose A with the beats as material (09-14). Built as loop recipes the same day and run on the twelve fixed assignments, persona writers, DeepSeek Flash, off-peak: r19 = three drafts, the code picks (today's process, without the dealt hand); r20 = the same, then the beats in the writer's own call and the repair pass, four candidates per failing paragraph.
| per piece, mean of 12 | r19 · the pick | r20 · Algorithm A |
|---|---|---|
| counted defects (the lint, paragraph by paragraph) | 7.9 | 0.5 |
| over-used phrase excess | 12.9 | 7.4 |
| paragraphs closing on the gavel | 0.83 | 0.25 |
| gavel share of paragraphs | 7.9% | 3.1% |
| whole-piece lint: pass | 17% | 67% |
| whole-piece lint score | 10.4 | 3.4 |
| words | 1,017 | 933 |
| contractions per 100 words | 2.44 | 2.22 |
| paragraphs re-rolled · replaced | — | 6.9 · 4.8 |
| candidate calls | — | 28 |
| the loop: seconds · dollars | 158 · $0.20 | 201 · $0.29 |
Three readings. The counted defects fell by an order of magnitude, as the simulation said they would, and the whole-piece verdict flipped from one in six passing to two in three. The over-use excess fell by less than half: the ruler's order puts the lint's defects first, so a candidate that trades one antithesis for two over-used phrases wins; a weighted key would trade differently. And the cache held: four of five prompt tokens in the repair calls were served from cache, which is why 28 candidate calls a piece cost three-quarters of a cent, though more paragraphs failed than the page assumed, so the cost is about double the estimate.
exam/ink_loop.py) on ink_brew.repair_pass; four candidates a paragraph; the positive tie-break key (facts from the reporting) in as the last key only. B is not built. Not yet in the brew: that waits on the read.