← Design notes

The quality score — Ink's target functionlive · calibrating

Ink's pieces are written once and read by everyone, so quality is the first priority (owner, 2026-09-02). This page is the method: how an editorial is scored, how the persona × topic crossover is measured so a piece reads as written by the persona and not by an imposter, what the first readings said, and how the readers' verdicts calibrate the judges. The brew selects by this number; every knob in the brew is tuned against it.

1 · The number

IQS = 100 × (wE·E + wV·V + wS·S + wL·landed)   over pieces that pass the lint

The weights follow the format's promise (owner, 09-02: does a fable need a different target function? — yes, in the weights, not the axes). Tonight's reading showed why: the fable and the kicker scored specificity 0.0 by construction and fell below plainly worse pieces. Same scale for every format — a 75 fable and a 75 essay are both good of their kind.

Format classEVSlandedThe light formats' E rubric
reported — Essay · Notes · Review · Scorecard0.400.350.25for the light formats the「position」line reads does it make ONE point that lands — a premise pushed to its end, a turn the reader didn't see coming — rather than a thesis dressed as a joke?; a kicker with a thesis is a bad kicker
voiced — Letter · Advice · Dispatch · Dialogue0.400.450.15
light — Fable · Shouts · Overheard · Kicker · Riddle0.300.450.000.25

landed is the gate's cold read of a light piece — a second model, as a stranger, answers one question: is it actually funny or delightful, or an essay wearing a joke hat, a pun list, a laboured premise? It was a pass/fail at the gate (it dropped Yu Hua's kicker as「a somber reflection, not a joke」); for the light formats it now counts in the score too, where specificity cannot.

AxisWhat it measuresHowWho judges
L · the lintthe machine's tells — the negation pivot (not X, but Y), em-dash cadence, the rule of three, the metronome, the AI vocabulary; in English and in Chineseregexes, no call (lib/ink_lint.py); a hard gate: a piece with a tell is edited or dropped, never scoredcode
E · editorialthe Pulitzer test for editorial writing, adapted:「clearness of style, moral purpose, sound reasoning, power to influence」→ clarity · reasoning · specificity · a position a reader could disagree with · carry (one thing the reader can repeat tomorrow)1–5 per line, each with one quoted line of evidence; (mean − 1) / 4the rival family (Gemini Flash; the writer is DeepSeek)
V · voicethe crossover: does the piece read as the persona's own writing?three independent measures, none of which sees the profile: trait match (½) — five checkable writing traits extracted once from the persona's REAL prose, scored yes/partly/no on the piece; the imposter test (¼) — the judge gets the persona's real paragraphs and three unlabelled pieces (the persona's, the same piece rewritten in another persona's voice, another persona's piece) and must pick from the writing alone, both orderings; the stylometric rank (¼) — function words, punctuation, sentence shape against the persona's own prose, ranked against the cast (code)the rival family + code
S · specificityreported opinion beats asserted opinionnumbers, dates, names, attributed quotes per 100 words, squashed to 0–1 (3 per 100 words is a reported piece; 0.5 is a sermon)code
R · the readersthe ground truth👍/👎 at the end of every piece, the written Feedback (the ⋯ menu), the door taps (rooms opened from the piece)people

2 · Where the method comes from

FindingSourceWhat Ink took from it
The test of excellence for editorial writing is「clearness of style, moral purpose, sound reasoning, and power to influence public opinion」; since 2026 the prize is Opinion Writing —「well-reasoned and compelling arguments… originally researched and reported, or informed by personal experience」pulitzer.orgthe five E lines, and「reported」as an axis of its own (S)
LLM judges agree with human reviewers about as often as humans agree with each other — IF the rubric separates its criteria, pairwise judgments swap positions, length is controlled, the judge is never the writer's own family, and the judge is re-calibrated against human labels monthlythe 2025–26 judge-bias literatureevery judge is the rival family; the imposter test runs both orderings; the readers' verdicts are the labels the judges are calibrated to
The authorship gap: personalized LLM text sits below the human cross-author floor (LUAR 0.48–0.51 vs 0.63) — closer to the model's own fingerprint than to any human. Single metrics lie (cross-metric |r| < 0.07). A judge that reads the same profile the writer read measures instruction-following, not voicePersonalBench; Theory-grounded evaluation of LLM personalization (both 2026)three independent voice measures, none of which sees the profile; the honest expectation: climb toward the human floor, do not expect to cross it
Voice comes from real prose samples, not a description; stripping tells without restoring voice leaves sterile textSudowrite's voice matching; the two-classifier findingthe exemplar bank — the persona's own quoted paragraphs — rides the writer's brief, the editor's voice pass, and the judges

3 · What the first readings said

Seventeen proof pieces, three editions, scored in one evening (about 0.9¢ a piece). IQS ran 46–89.

PieceIQSEVStrait match
Buffett · letter88.91.000.800.840.6
Li Ka-shing · letter84.00.900.850.700.7
Jobs · essay77.90.851.000.201.0
Banksy · Shouts71.80.850.860.181.0
Einstein · notes65.70.750.600.590.2
Asimov · fable64.40.850.710.110.7
Mouratoglou · Shouts63.90.400.950.490.9
Brownlee · review62.00.850.650.100.8
Dan Wang · notes61.00.750.540.460.4
Roosevelt · essay60.50.600.590.640.4
Churchill · essay58.00.800.440.410.2
Federer · kicker55.60.650.590.300.4
Darwin · dispatch50.90.750.410.220.2
Chestnut · fable48.00.600.310.580.0
Twain · kicker46.20.500.490.320.3
Dylan · Shouts46.10.550.400.410.3

4 · The two levers

The reporter's notebook Stage 2½ of the brew. For every commission: the peg article fetched in full where the site allows (the link-unfurl module; walled sites fall through to the snippet) · three searches — the subject, the editor's angle, the persona × the subject — with the two best pages fetched · one Flash distillation into dated facts, numbers, people, attributed quotes, what is disputed, what is not established, and what the persona has said on this before, every line with its source number. The writer's rule changes from「facts only from the peg card」to「facts only from the notebook」; the reader's peg card grows to the article's opening paragraphs when the fetch succeeded. About 0.6¢ a piece.
The voice harvest For a persona whose exemplar bank is thin — the historic greats above all (Ink's staples: fewer rights questions, and the owner's ask: make them read great) — search the open web for PRIMARY text (letters, notebooks, speeches, essays), fetch the best pages, and have Flash copy out verbatim paragraphs written by the persona, never about. The bank grows to 14; the persona's traits are re-extracted from the fuller bank; the writer's brief picks the three paragraphs nearest the piece's subject. python lib/ink_brew.py voices runs it over every fixed-horizon persona; about 1¢ each.

5 · Calibration — the readers are the label

Every piece ends in a 👍/👎 pair and, in its ⋯ menu, a Feedback… sub-page where a reader writes what they thought. During calibration the score is shown to every reader on the piece, so the number is public and argued with. Monthly (or whenever a hundred verdicts have landed): the judges' E and V are compared with the readers' verdicts and the owner's own reads; where they disagree, the weights move toward the people, and a judge that keeps disagreeing is swapped — treated as an instrument change, not a config change. The first known disagreement is on record above.

The optimisation loop Every knob in the brew — the writer's brief, the exemplar count, the editor's passes, the notebook's depth, the register mix, the model — is changed one at a time and read on IQS over a week of editions, with the readers' verdicts as the check that the judges still agree with people. The persona × beat map (IQS by persona and by beat) is the first thing to read each week: it says which crossovers work and which personas still write as imposters.
Method live from 2026-09-02 · code: lib/ink_score.py · lib/ink_lint.py · the brew's stages in lib/ink_brew.py · siblings: the brew · the essays · the look