Selection by the paragraph — two algorithms for a ruler that only counts defects A chosen · built · first run 2026-09-14

The process as it stands (settled 09-13): four drafts, the ruler sorts them, the lowest wins, the gate marks the winner. You named the flaw: a ruler that counts defects cannot be beaten by more whole drafts, because a long piece always carries many. The only way down is to select at the grain where the defects live, the paragraph. This page draws the two ways to do that, one you proposed and one I proposed, with the arithmetic, the cost, and what each one risks. Nothing is built yet.

the answer in one paragraphBoth algorithms cut the counted defects by an order of magnitude on paper, and both keep coherence, because every candidate paragraph is written into its place with the accepted text in front of it — this is not stitching drafts together. Algorithm A, repair, keeps today's four-draft pick as the spine and re-rolls only the paragraphs the ruler fails, in place: fewer calls, the arc preserved, one seam per repair to watch. Algorithm B, sequential, builds the piece one paragraph at a time, several candidates a step, the cleanest kept: selection on every paragraph, no seams, but the writer has no arc unless it is given one, and a loop that asks for "the next paragraph" manufactures the very closer we are trying to lose. On the simulation, A reaches fewer defects with a third of the calls. B is the stronger idea if the tie-break stops being defect-only. I'd run both on the twelve fixed assignments and read the results blind.

1 · The problem, in numbers

Short answer: picking the cleanest of four whole drafts removes about a third of the defects. Picking the cleanest of four candidates at every paragraph removes nearly all of them.

p is the share of paragraphs that carry a counted defect on any one throw. On yesterday's editions it is around 0.4: ten marked pieces of sixteen, and the over-use profile's rates. K = 15 paragraphs a piece, the median. N = candidates bought per unit. The table is a simulation of 20,000 pieces per row, with every throw independent. That last assumption is the optimistic one: four throws of the same model in the same context cluster (the research's mode collapse), so the real N is smaller than the bought N. The ordering of the three columns holds; the absolute numbers are a ceiling.
pNone draftpick of N drafts (today)A · repairB · sequentialparagraph calls: A / B
0.344.52.70.020.12~12 / 60
0.446.04.10.100.38~16 / 60
0.486.03.40.000.01~24 / 120
0.547.55.50.340.93~24 / 60
0.587.54.80.020.06~40 / 120

Expected counted defects per piece. The middle column is your law of large numbers: doubling the drafts from four to eight buys less than one defect. The two right columns are the paragraph grain. A comes out ahead of B on this table only because its spine has already been selected once, so it has fewer paragraphs to re-roll; the two converge on the same floor.

2 · Algorithm A — the spine, repaired in place

Short answer: keep today's pick as the spine, run the ruler paragraph by paragraph, and re-roll the failing paragraphs where they stand, several candidates each, the cleanest kept.

the commission peg · notebook · interview draft 1 · DS draft 2 · DS draft 3 · DS draft 4 · Gemini the pick six keys, lowest wins = today's process, unchanged the spine, paragraph by paragraph ¶1 ✓ ¶2 ✓ ¶3 ✗ ¶4 ✓ ¶5 ✗ ¶6 ✓ …¶15 the ruler on each paragraph alone: antithesis · constructions · over-use · gavel · metaphor nouns re-roll ¶3 in its place same prompt · ¶1–¶2 accepted before it · ¶4 as the paragraph to connect to the fault is never named — resampled, not edited cand. a cand. b cand. c cand. d the ruler picks the cleanest, or keeps ¶3 if none beats it the prefix is identical across candidates → served from cache back then the whole piece once: metronome · contractions · the mark unchanged from today A · one round, the failing paragraphs re-rolled in parallel · ~16 paragraph calls at p = 0.4 · about 15 seconds a piece
Algorithm A. The four-draft pick is untouched; what it chooses becomes the spine. Only the paragraphs the ruler fails are re-rolled, each in its own place with the text before and after it, and the model is never told what was wrong. The arc comes from the spine, so it survives.
  1. Spine. Today's pick, unchanged.
  2. Ruler, per paragraph. The counters that can judge a paragraph alone. Zero allowance for the strong over-use phrases, because a paragraph is short.
  3. Re-roll in place. For each failing paragraph: the same writer prompt, the accepted paragraphs before it, the next paragraph after it as material to connect to, and the request for this one paragraph. Four candidates. No note on the fault.
  4. Pick. The cleanest candidate by the same per-paragraph keys; the original stays if it ties or wins.
  5. Whole piece. The meters that need the whole run once at the end, as now.

What it risks. A seam: the new ¶3 must lead into the old ¶4, which was written after a different ¶3. Giving the candidate ¶4 to connect to is the mitigation; your blind read is the check. And a repaired paragraph can be clean and empty, the wire-round-up failure, which no counter sees.

3 · Algorithm B — one paragraph at a time, the best of several, forward

Short answer: write the first paragraph several times, keep the best; give the task and the kept paragraph back, write the second several times, keep the best; continue to the end. Your proposal.

the commission peg · notebook · interview + the beats (see §4) step 1 · ¶1 a b c d the ruler keeps b step 2 · ¶2, given ¶1 a b c d the ruler keeps d step 3 · ¶3, given ¶1–¶2 a b c d the ruler keeps a … step 15 each call answers "is this the last?" the piece grows only forward — no seams, every paragraph selected ¶1 b ¶2 d ¶3 a ¶4 c ¶15 what the ruler cannot see at a step: whether ¶4 repeats ¶2's point · whether the piece is going anywhere whether this paragraph is a closer because the loop asked for "the next paragraph" the tie: four clean candidates and a ruler with nothing left to say today's order ends on the lint score, zero for all four → one positive key is needed: how many notebook facts the paragraph uses (§4) B · fifteen sequential rounds · ~60 paragraph calls · about 2–4 minutes a piece · the prefix cached at every step
Algorithm B. Nothing is written twice into the same slot from different pieces; the piece only grows forward, so coherence is structural. Every paragraph passes through selection. The price is that the writer has no arc of its own and the loop's question, "the next paragraph, please", invites a closer every time.
  1. Step 1. The commission, plus the beats (§4). Four candidates for the first paragraph; the ruler keeps one.
  2. Step k. The commission, the beats, the accepted ¶1…¶k−1, and the request for the next paragraph with a yes-or-no "is this the last?". Four candidates; the ruler keeps one.
  3. Stop when the kept candidate says it is the last and the piece is inside its word band; or at a hard cap.
  4. Whole piece. As in A, the whole-piece meters once at the end.

What it risks. Three things, all structural. The arc: nothing in fifteen local choices builds toward a turn, and a defect-only ruler prefers the safe paragraph at every step. The closer: a model asked for "the next paragraph" writes each one as an ending; the gavel key catches the sentence, but the habit is manufactured by the loop. The tie: once candidates are clean, the ruler is silent and the choice is random unless a positive key exists.

4 · Two parts both need

The beats — an arc as material, not a rule

B has no arc without one. A has the spine's, but a repaired paragraph should know where the piece is going too. One call before the writing: the piece's beats, five or six lines, what happens in what order and where it turns. They enter the writer's prompt as material, beside the interview answers, never as an instruction on how to write. This is the shape menu's idea done by the writer itself rather than picked from our list; the menu failed as a rule, and this is the test of whether it works as material.

One positive key — the first in the pipeline

Both algorithms end in ties: several clean candidates and a ruler with nothing to say. Today the tie falls to the lint score, which is zero for all of them. The candidate to prefer is the one that uses the most of what it was given: names, numbers, places and quotes from the notebook and the interview that appear in the paragraph, counted by the same anchor matching the brew already does for the dealt hand. It is the first key in the pipeline that rewards what a paragraph has rather than what it lacks, and it is the only defence against the clean-and-empty paragraph. A count, not a judgment; no model scores anything.

5 · Side by side

A · repairB · sequential
selection acts onthe failing paragraphs of a chosen draftevery paragraph
expected defects, p = 0.4, N = 40.100.38
paragraph calls a piece~16~60
added cost a piece, Flash off-peak~0.3¢~1¢
time a pieceone round, ~15 s15 rounds, 2–4 min
coherenceby conditioning; one seam per repairstructural; no seams
the arcthe spine's, keptnone unless the beats are given
the endingthe spine'sthe writer must flag it; the loop invites closers
what it changes in the brewa pass after the pickreplaces the drafts and the pick
where it failsa clean, empty repair; a bad seama clean, shapeless piece; a closer every paragraph
how it can growN up; re-roll borderline paragraphs tookeep two prefixes alive instead of one
Cost basis. A candidate call reads the writer's prefix, about 13,000 tokens, from cache at $0.003 a million, and writes one paragraph, about 150 tokens at $0.60 a million: about a hundredth of a cent. The candidates of one step share a prefix byte for byte, so only the first of them pays the uncached read. Peak hours double it. Both algorithms run on Flash; Gemini's caching works differently and its seat stays as the fourth draft in A.

6 · How to decide between them

Short answer: build both as loop recipes, run them on the twelve fixed assignments against today's pick, count, and then read blind.

runwhatcost
baselinethe four-draft pick as it runs today~50¢
Abaseline + the repair pass, N = 4~55¢
Bsequential, N = 4, with the beats~65¢

By count: defects per piece by the ruler, over-use excess, gavel share, contraction rate, words, cost, minutes. Then your blind read of the three versions of the same assignment, unlabelled, for the two things no counter sees: does the piece go somewhere, and is anything under the clean sentences. The critic's pairwise read against the majors can follow if the blind read is promising; it is not the decider here.

7 · The first run — Algorithm A, built and measured

Short answer: you chose A with the beats as material (09-14). Built as loop recipes the same day and run on the twelve fixed assignments, persona writers, DeepSeek Flash, off-peak: r19 = three drafts, the code picks (today's process, without the dealt hand); r20 = the same, then the beats in the writer's own call and the repair pass, four candidates per failing paragraph.

per piece, mean of 12r19 · the pickr20 · Algorithm A
counted defects (the lint, paragraph by paragraph)7.90.5
over-used phrase excess12.97.4
paragraphs closing on the gavel0.830.25
gavel share of paragraphs7.9%3.1%
whole-piece lint: pass17%67%
whole-piece lint score10.43.4
words1,017933
contractions per 100 words2.442.22
paragraphs re-rolled · replaced6.9 · 4.8
candidate calls28
the loop: seconds · dollars158 · $0.20201 · $0.29

Three readings. The counted defects fell by an order of magnitude, as the simulation said they would, and the whole-piece verdict flipped from one in six passing to two in three. The over-use excess fell by less than half: the ruler's order puts the lint's defects first, so a candidate that trades one antithesis for two over-used phrases wins; a weighted key would trade differently. And the cache held: four of five prompt tokens in the repair calls were served from cache, which is why 28 candidate calls a piece cost three-quarters of a cent, though more paragraphs failed than the page assumed, so the cost is about double the estimate.

your read · owedThe counts cannot see the two things that decide it: whether a repaired piece still goes somewhere, and whether the seams show. The read page holds the twelve assignments as unlabelled pairs, shuffled; the key sits beside the loop, not on the page. For each pair: which side you would keep reading, and which reads as the machine.
decisions taken · 2026-09-14Algorithm A, with the beats as material, built as recipes r19 and r20 (exam/ink_loop.py) on ink_brew.repair_pass; four candidates a paragraph; the positive tie-break key (facts from the reporting) in as the last key only. B is not built. Not yet in the brew: that waits on the read.
decisions asked · 2026-09-14 · superseded above 1. Build both as loop recipes and run the three-way test, or only one — which? 2. The beats call: in for both, in for B only, or out? 3. The positive tie-break key, notebook facts used: in or out? It is the one piece of this page that changes the ruler's nature. 4. N = 4 or 8 candidates a paragraph; the cost difference is small, the clustering caveat applies to both.
Reviewed 2026-09-14 · A built as loop recipes r19/r20 the same day · related: the over-use profile · the industry and the habits §2 (the anti-slop sampler is this idea at the token level) · the outside research §5 ③ (pick, never merge) · the writing loop · the brew