Ink's pieces are written once and read by everyone, so quality is the first priority (owner, 2026-09-02). This page is the method: how an editorial is scored, how the persona × topic crossover is measured so a piece reads as written by the persona and not by an imposter, what the first readings said, and how the readers' verdicts calibrate the judges. The brew selects by this number; every knob in the brew is tuned against it.
The weights follow the format's promise (owner, 09-02: does a fable need a different target function? — yes, in the weights, not the axes). Tonight's reading showed why: the fable and the kicker scored specificity 0.0 by construction and fell below plainly worse pieces. Same scale for every format — a 75 fable and a 75 essay are both good of their kind.
| Format class | E | V | S | landed | The light formats' E rubric |
|---|---|---|---|---|---|
| reported — Essay · Notes · Review · Scorecard | 0.40 | 0.35 | 0.25 | — | for the light formats the「position」line reads does it make ONE point that lands — a premise pushed to its end, a turn the reader didn't see coming — rather than a thesis dressed as a joke?; a kicker with a thesis is a bad kicker |
| voiced — Letter · Advice · Dispatch · Dialogue | 0.40 | 0.45 | 0.15 | — | |
| light — Fable · Shouts · Overheard · Kicker · Riddle | 0.30 | 0.45 | 0.00 | 0.25 |
landed is the gate's cold read of a light piece — a second model, as a stranger, answers one question: is it actually funny or delightful, or an essay wearing a joke hat, a pun list, a laboured premise? It was a pass/fail at the gate (it dropped Yu Hua's kicker as「a somber reflection, not a joke」); for the light formats it now counts in the score too, where specificity cannot.
| Axis | What it measures | How | Who judges |
|---|---|---|---|
| L · the lint | the machine's tells — the negation pivot (not X, but Y), em-dash cadence, the rule of three, the metronome, the AI vocabulary; in English and in Chinese | regexes, no call (lib/ink_lint.py); a hard gate: a piece with a tell is edited or dropped, never scored | code |
| E · editorial | the Pulitzer test for editorial writing, adapted:「clearness of style, moral purpose, sound reasoning, power to influence」→ clarity · reasoning · specificity · a position a reader could disagree with · carry (one thing the reader can repeat tomorrow) | 1–5 per line, each with one quoted line of evidence; (mean − 1) / 4 | the rival family (Gemini Flash; the writer is DeepSeek) |
| V · voice | the crossover: does the piece read as the persona's own writing? | three independent measures, none of which sees the profile: trait match (½) — five checkable writing traits extracted once from the persona's REAL prose, scored yes/partly/no on the piece; the imposter test (¼) — the judge gets the persona's real paragraphs and three unlabelled pieces (the persona's, the same piece rewritten in another persona's voice, another persona's piece) and must pick from the writing alone, both orderings; the stylometric rank (¼) — function words, punctuation, sentence shape against the persona's own prose, ranked against the cast (code) | the rival family + code |
| S · specificity | reported opinion beats asserted opinion | numbers, dates, names, attributed quotes per 100 words, squashed to 0–1 (3 per 100 words is a reported piece; 0.5 is a sermon) | code |
| R · the readers | the ground truth | 👍/👎 at the end of every piece, the written Feedback (the ⋯ menu), the door taps (rooms opened from the piece) | people |
| Finding | Source | What Ink took from it |
|---|---|---|
| The test of excellence for editorial writing is「clearness of style, moral purpose, sound reasoning, and power to influence public opinion」; since 2026 the prize is Opinion Writing —「well-reasoned and compelling arguments… originally researched and reported, or informed by personal experience」 | pulitzer.org | the five E lines, and「reported」as an axis of its own (S) |
| LLM judges agree with human reviewers about as often as humans agree with each other — IF the rubric separates its criteria, pairwise judgments swap positions, length is controlled, the judge is never the writer's own family, and the judge is re-calibrated against human labels monthly | the 2025–26 judge-bias literature | every judge is the rival family; the imposter test runs both orderings; the readers' verdicts are the labels the judges are calibrated to |
| The authorship gap: personalized LLM text sits below the human cross-author floor (LUAR 0.48–0.51 vs 0.63) — closer to the model's own fingerprint than to any human. Single metrics lie (cross-metric |r| < 0.07). A judge that reads the same profile the writer read measures instruction-following, not voice | PersonalBench; Theory-grounded evaluation of LLM personalization (both 2026) | three independent voice measures, none of which sees the profile; the honest expectation: climb toward the human floor, do not expect to cross it |
| Voice comes from real prose samples, not a description; stripping tells without restoring voice leaves sterile text | Sudowrite's voice matching; the two-classifier finding | the exemplar bank — the persona's own quoted paragraphs — rides the writer's brief, the editor's voice pass, and the judges |
Seventeen proof pieces, three editions, scored in one evening (about 0.9¢ a piece). IQS ran 46–89.
| Piece | IQS | E | V | S | trait match |
|---|---|---|---|---|---|
| Buffett · letter | 88.9 | 1.00 | 0.80 | 0.84 | 0.6 |
| Li Ka-shing · letter | 84.0 | 0.90 | 0.85 | 0.70 | 0.7 |
| Jobs · essay | 77.9 | 0.85 | 1.00 | 0.20 | 1.0 |
| Banksy · Shouts | 71.8 | 0.85 | 0.86 | 0.18 | 1.0 |
| Einstein · notes | 65.7 | 0.75 | 0.60 | 0.59 | 0.2 |
| Asimov · fable | 64.4 | 0.85 | 0.71 | 0.11 | 0.7 |
| Mouratoglou · Shouts | 63.9 | 0.40 | 0.95 | 0.49 | 0.9 |
| Brownlee · review | 62.0 | 0.85 | 0.65 | 0.10 | 0.8 |
| Dan Wang · notes | 61.0 | 0.75 | 0.54 | 0.46 | 0.4 |
| Roosevelt · essay | 60.5 | 0.60 | 0.59 | 0.64 | 0.4 |
| Churchill · essay | 58.0 | 0.80 | 0.44 | 0.41 | 0.2 |
| Federer · kicker | 55.6 | 0.65 | 0.59 | 0.30 | 0.4 |
| Darwin · dispatch | 50.9 | 0.75 | 0.41 | 0.22 | 0.2 |
| Chestnut · fable | 48.0 | 0.60 | 0.31 | 0.58 | 0.0 |
| Twain · kicker | 46.2 | 0.50 | 0.49 | 0.32 | 0.3 |
| Dylan · Shouts | 46.1 | 0.55 | 0.40 | 0.41 | 0.3 |
python lib/ink_brew.py voices runs it over every fixed-horizon persona; about 1¢ each.Every piece ends in a 👍/👎 pair and, in its ⋯ menu, a Feedback… sub-page where a reader writes what they thought. During calibration the score is shown to every reader on the piece, so the number is public and argued with. Monthly (or whenever a hundred verdicts have landed): the judges' E and V are compared with the readers' verdicts and the owner's own reads; where they disagree, the weights move toward the people, and a judge that keeps disagreeing is swapped — treated as an instrument change, not a config change. The first known disagreement is on record above.
lib/ink_score.py · lib/ink_lint.py · the brew's stages in lib/ink_brew.py · siblings: the brew · the essays · the look