Yesterday's page read the research. This one reads the businesses. Four sweeps: the writing apps people pay for (Sudowrite, NovelAI, Lex, Jasper, Grammarly and a dozen more) · the model makers and the open-source people who train against "slop" · the detector-and-humanizer trade · the Chinese market, from 阅文 to the 降AI shops on 淘宝. About 120 sources, most read at the source. Each answer is set against our own 20 habits and yesterday's seven untried fixes.
Short answer: four kinds of company, and they don't fight the same thing.
| kind | who | what they sell | which machine habit they fight | how |
|---|---|---|---|---|
| Writing apps | Sudowrite · NovelAI · Novelcrafter (fiction) · Lex · Type (essays) · Jasper · Copy.ai · Writer · Grammarly · Notion (business copy) | a place to write with a model inside | clichés, purple words, "sounds like a robot", not sounding like you | own models (fiction apps); the writer's own samples in the prompt (everyone else) |
| Model makers | OpenAI · Anthropic · Moonshot (Kimi) · DeepSeek · xAI · Google | the model | sycophancy, length, bullet-heavy answers; recently, their own tics | training data and reward rules; a "personality" dial; advice pages |
| Open-source trainers | Sam Paech (Antislop, EQ-Bench) · the fine-tune community (TheDrummer, Sao10k) | methods and tuned models | over-used words and phrases, sameness | count what the model over-uses against human text, then train or block it |
| Detectors and humanizers | Pangram · GPTZero · Originality · Turnitin | Undetectable.ai · StealthGPT · 笔灵 · 火龙果 · 淘宝 shops | a score; a rewrite that lowers the score | whatever a classifier sees | classifiers; paraphrase, synonym swaps, sentence splitting, even injected typos |
Short answer: the deeper the layer, the better the evidence — and the harder it is to reach on a model you rent.
This is what most people try first, and what most of the trade sells. It doesn't hold. Three findings, from three different directions:
The manner survived; the grammar didn't. Our own loop found the words layer the same way: never-lines lowered the counts and the critic still spotted the machine every time.
The standard answer of every essay and business tool. Sudowrite asks for at least five paragraphs of the writer's real prose and injects them into every request; Writer.com trained one model to extract a "voice profile" from samples and another to write to it; Grammarly builds a passive profile from what you type; Noren, a ghostwriting startup, takes 15–20 samples and claims to recover 90% of a hand-written voice guide; Type lets you attach 100,000 words. Lex and Jasper add rule-lists on top.
What it gets you is known from yesterday's page: samples in the prompt reach about 82% of a writer's style, they carry sentence length and rhythm, and they do not carry the voice. Two products admit the limit in their own words: Writer's voice "degrades over extended content" (a reviewer), and Lex's founder says uploading samples to a general chatbot "typically produces disappointing results" — the reason Lex built its own layer. Our shelf of real openings was this layer, and it was one of our better rounds. It's a floor, not a fix.
The only layer with numbers, and where every credible "no AI-isms" claim lives.
| who | what they did | what they claim or measured |
|---|---|---|
| Sudowrite · Muse (Mar 2025) | fine-tuned on licensed fiction with the authors' consent; base model undisclosed; runs a multi-step pipeline (analyse → plan → refine → revise) rather than one call; "we specifically measure AI clichés during training and have systematically removed them" | "no more tapestries and delving"; internal claim of 40% fewer revision passes for voice, no method; a user: "thoroughly cut back AI-isms… not totally gone, but way less" |
| NovelAI · Erato, Xialong (2024–26) | starts from a base model, never an assistant one, continues pretraining on its own fiction corpus for hundreds of billions of tokens, then a storytelling fine-tune; reinforcement learning used only against repetition and loops | no numbers; the design bet is that there is no "Certainly!" to un-train because it was never trained in |
| Novelcrafter (Feb 2024) | per-author fine-tune: 50–75 pairs of (scene beat → the author's own prose) on a hosted model; ~$2.40 per 100k tokens of training | "remove AI-isms at the source"; one example, no measurement |
| Antislop / FTPO (Paech et al., ICLR 2026) | count which words and phrases a model over-uses against human text (some over 1,000× more often), then adjust only the model weights that start those phrases | about 90% less slop with reasoning tests held and judged writing up; the ordinary preference-training method (DPO) on the same data removed less and damaged writing quality and word variety |
| Moonshot · Kimi K2 (Jul 2025) | reward rules in training that ban "opening with compliments" and "sentences explaining why the response is good or how it fulfils the request" | top of two creative-writing leaderboards at launch; still reuses the same story devices across pieces |
| Apple/CMU; PASTA (2026) | keep meaning-annotations through training (6× less collapse into sameness); or find the "assistant direction" inside an open model and subtract it at generation time | research on open models only |
Short answer: the fiction apps fight the habits; the business apps fight only "doesn't sound like our brand".
| product | sells | habits it names (our #) | mechanism | evidence |
|---|---|---|---|---|
| Sudowrite | novel drafting, $19+/mo | clichés, "flowery language" (#9), "way too telly", exposition (#18), em-dashes (#19), flattery | own model + raw samples + pipeline + a 0–10 "creativity" dial | internal only |
| NovelAI | story continuation, $10–25/mo | repetition, name sameness, assistant voice | base-model training; user-set phrase bias and banned tokens (the docs warn the model "may attempt workarounds using alternate spellings") | none published |
| Novelcrafter | bring-your-own-key workspace | "AI-isms": palpable, tangible, crystalline, ethereal; adverbs; the names Elara, Marcus, Nakamura | per-author fine-tune; a manual highlighter that marks the tell-words red in the manuscript | one example |
| Lex | essay editor, Claude inside | sycophancy, "bullet-point-y, robot-generated" register | feedback, not generation: an anti-flattery critic ("being nice does the writer a disservice"), an Awesome / Boring / Confusing / Didn't-believe rubric, critic personas; voice from an imported Substack | churn fell 20–30% after moving to Claude; no prose measure |
| Type | long-form business writing | bland defaults | up to 100k words of the writer's material in context; ask for rare qualifier words | none |
| Writer.com | enterprise, own Palmyra models | off-brand voice | a voice-extraction model + a voice-generation model; rule checker | reviewer: profile "static", voice "degrades over extended content" |
| Jasper · Copy.ai | marketing copy | "generic, machine-made" | brand voice from ≥300 words; a post-generation rule checker that flags violations | none |
| Grammarly | the editor everyone has | "overly formal tone, repetitive phrasing" | passive voice profile; an AI Humanizer (Sep 2025) that rephrases; an Authorship record of typed vs pasted vs generated | "95% of users report confidence" — a testimonial |
| Noren · Oiti · Postiv | ghostwriting for founders | not sounding like the founder | 15–20 samples → an "identity layer" (recurring words, analogy domains) | internal 90% |
| Notion · Copilot · Wordtune · Rytr | copy inside other tools | — | custom-instruction fields only | Rytr's output is "pattern-heavy enough that detection tools flag it" |
Two shapes recur. The fiction apps split the work — outline, beats, then prose; Sudowrite's model is a pipeline of steps, not a call. And the essay app with the best reputation, Lex, doesn't write: it criticises what a person wrote, with a prompt that tells the critic not to be kind. Neither addresses a paragraph that ends on a verdict.
Short answer: the labs fight sycophancy and length, not prose habits; and their own tics are getting stronger with each release.
| lab | what they did about the machine sound | what it tells us |
|---|---|---|
| OpenAI | rolled back a sycophantic release (Apr 2025) by adding training examples that used to draw over-agreement; GPT-5 "less like talking to AI"; then GPT-5.1 "warmer by default" with eight personality presets; later notes name "teaser-style phrasing" and "bullet-heavy responses" as things reduced; the Model Spec's whole style section is one-liners ("be clear and direct", "don't be sycophantic") | register is a dial they turn both ways; no lab-level work on clichés, antithesis or closers |
| Anthropic | Opus 5's tics were measured by a public leaderboard in Jul–Aug 2026: 510 words a reply vs 158 for Opus 4.5, sentences 58% longer, em-dashes 2.3×, "load-bearing" 2×, "honestly / frankly" up 50%; a 1,778-point forum thread and a bug report titled "increasingly default to repetitive rhetorical tics". Anthropic's answer (Sep 2026) was a prompting-guide section on mannered prose — "substitutes metaphor and flourish for direct statement… makes the reader work harder so the writer can perform" — and a paste-in paragraph | the community's list of "Claudisms" — load-bearing, worth stating plainly, full stop, isn't just X — it's Y — is our list; the fix offered is a prompt, i.e. layer 1 |
| Moonshot · Kimi K2 | the one disclosed anti-slop reward design (§2) | a hosted model trained against habit #7 and against flattery |
| DeepSeek | nothing published on writing; the community explains "DeepSeek味" — forced lyricism, stacked images, 金句 — by an estimated 40% literary share in its training data against 10–20% for rivals; Chinese lists of the flavour: 工整的对仗句和排比句, 模版化, 括号量化 | our writer's metaphor habit (#9) and the wise-sounding line (#8) are the model's diet, not our prompt |
| xAI · Google | Grok 4.1 uses reasoning models as reward models for style; Gemini forum threads report a long-form regression 2.5 → 3.1 | nothing usable |
Short answer: the detectors that work read the whole text, not words; heavy chatbot users spot machine prose almost perfectly; and the tell lists agree with ours on the words but miss our top two.
| fact | number | who |
|---|---|---|
| The best detector's false-positive rate on human text, in a lab test | 0.1% | Pangram, tested by Chicago Booth / NBER, Sep 2025 |
| …and on humanized text | 95–100% caught | same; VU Brussel, Jun 2026 |
| Word-list detection on Claude's writing | near random | KU Leuven, ACL workshop 2025 |
| Non-native English essays flagged as AI by seven detectors | 61% average | Liang et al., Patterns 2023 |
| Five heavy ChatGPT users, majority vote, on 300 articles | 1 wrong | Russell et al., ACL 2025 — casual readers were at chance |
| How much an AI judge marks down text merely labelled "AI" | −34 points | Haverals & Martin, 2025; humans −14 |
| Undetectable.ai traffic | ~2.3M visits/mo | Similarweb, Aug 2026; mostly students |
Two consequences. Substack switched on a Pangram-powered scan for readers in July 2026, so a column on a public surface will be scored whether or not it is disclosed. And the readers who matter to a column — people who use these models daily — are the 299-of-300 group. Our critic's 100% spot rate isn't a harsh judge; it is the audience.
Sources: Wikipedia's Signs of AI writing (13 sections), GPTZero's vocabulary list, Originality's corpus, Kobak's 15-million-abstract study, the editors' guides, and the Chinese RUC 新闻坊 study (§6).
| tell | named by | ours | note |
|---|---|---|---|
| "Not X, but Y" — negative parallelism | Wikipedia · Atlantic · Barron's · editors | #1 | Barron's counted it in Fortune-500 filings: 50 → 200+ from 2023 to 2025; the Atlantic (Jul 2026) calls it "the most mysterious" — nobody knows what in training drives it, and suppressing it may just move it |
| The rule of three | Wikipedia · GPTZero · RUC (排比) | #15 | humans triple sometimes; models every few lines |
| Em-dash overuse | Wikipedia · WaPo · NPR · a medRxiv study | #19 | a population-level signal, not a per-piece one: 4% → 12% of discussion sections; and in Chinese the dash is a myth — AI 0.58% vs human 0.85% |
| Significance inflation — testament, pivotal, underscores, evolving landscape | Wikipedia (its most consistent observation) · GPTZero | #8, #12 | "reads like promo copy" |
| Era-tagged vocabulary — delve (2023), showcase/foster (2024), emphasizing/highlighting (2025) | Wikipedia · GPTZero · Kobak | — | rotates each model generation; a list is stale in a year |
| Analysis tacked on with -ing — "highlighting…", "ensuring…" | Wikipedia · editors | #4 (cousin) | a clause that adds opinion, not information |
| Avoiding "is" — serves as, stands as, represents | Wikipedia | — | |
| Vague attribution — experts argue, observers note | Wikipedia · editors | #17 (cousin) | |
| Outline-shaped endings — "Despite these challenges…" | Wikipedia | #10 | |
| Throat-clearing — "it's important to note" | editors | #7 | the nearest anyone comes to announcing its moves |
| Specificity sanded off — an unusual fact replaced by a generic-positive one | LitHub summarising Wikipedia · the 2025 editing study | #8, #18 | the conceptual root of most of the list |
| Every paragraph ends on its own verdict | nobody | #3, #4 | only the Humanizer skill's "one-line closers"; no study, no product, no detector counts position |
| Numbers nobody needs · the credential paragraph | nobody | #14, #17 | ours alone, as yesterday |
Short answer: the platforms police machine writing by declaration and rank, the editors by eye, and nobody by model. The 降AI trade wrecks prose the same way its Western twin does.
| who | what they do | what they say the machine flavour is |
|---|---|---|
| 阅文 · 妙笔 / 作家助手 | own web-fiction model (Jul 2023, corpus size undisclosed) → DeepSeek-R1 inside the author tool (Feb 2025); "AI 是创作的金手指,主角永远是作家"; says it detects and punishes AI水文 | at the launch, authors asked whether the tool would make 同质化 worse; the VP's answer was that the market isn't saturated |
| 起点 | bans AI as the core of a book; removal from 月票榜 and a monthly public list of violators (Apr 2026), veterans named; editors reject even AI 润色 | detection method undisclosed — the editors read |
| 番茄小说 | the richest toolbox (改写 · 扩写 · 续写 · 卡文锦囊); the 2024 training-clause revolt; a mandatory "是否使用AI" checkbox since 23 Sep 2025; first-show new books 5,606 a month, peak 3,549 a day after DeepSeek | readers: the same opening recombined — "熙熙攘攘的街道,阳光如何如何"; authors: 人物关系前后矛盾 · 喜欢修饰语句,不推进情节 · 文笔漂亮但没有"时间"的概念; the platform's own detector scored one book's chapters 0% / 44% / 87% |
| 晋江 | the anti-AI pole (Feb 2025): allowed = 校对级 · 元素级 · 粗纲级; banned = AI 润色 and any AI plot; reports need a detector score over 60% | "hurts 人作为创作主体的原创性" |
| Editors (起点 · 番茄 · 盐言) | spot it "两分钟内", by eye, not tools; manuscripts up 50% a day after DeepSeek, "几乎全是AI稿" | five tells: 华美的空洞 (piled adjectives, nothing under them) · logic breaks across long spans · AI adds rather than cuts · description that doesn't move the plot · no feeling. Also leftover assistant lines: "以下是为您修改、润色和优化后的内容" |
| 唐家三少 (光明日报, 9 Sep 2026) | coins AI泔水: "通顺但空洞,正确但平庸,读起来像人话,细品没灵魂"; AI 洗稿 makes 5 million characters in 48 hours; asks for text labelling to be enforced | the Chinese word for slop, three days old |
| RUC 新闻坊 (人大新闻学院, Sep 2025) | 142 posts → 215 traits, 35 interviews, 7 models against a school-essay archive | AI uses 对偶 4× per text vs students 0.67; colons and semicolons up; rare words up to 110× human rate; 三段式 "首先…其次…最后"; grand nouns 智慧 · 时代 · 力量. And the dash is not a tell in Chinese. |
| The 降AI trade | 知网 / 维普 / 万方 detectors with university caps (C9 ≤15%); 淘宝 人工降写 ¥300 per 13,000 字, one shop over 4,000 orders; tools 笔灵 · 火龙果 · 千笔 · 灵笔 | the detectors scored 朱自清's 《荷塘月色》 at 63% AI and 《滕王阁序》 at 100%; the tools' output, by the vendors' own tests: 破碎感强 · 因果倒置 · 缺主语; students "被迫删掉精彩段落" |
| 宝玉 (Feb 2026) | the one Chinese source arguing the prompt approach is structurally wrong | AI味 = "用所有训练数据的平均风格写作"; everyone using the same 去AI味 prompt creates a new sameness; users only say what not to do; his fix is a living style file, revised by diffing your own edits |
One gap worth stating plainly: no Chinese study or platform names 不是…而是… as a tell. It appears only in prompt-craft lists on 知乎. The measured Chinese canon is 对偶 · 排比 · 比喻 · 金句 · 三段式 · 升华结尾 — the manner, again, not a word.
Short answer: nothing to buy; four levers within reach on a rented model; one ceiling closer than we thought.
| lever | who proved it | reachable on a rented model? | what it would be at Ink |
|---|---|---|---|
| Over-use profile against a human baseline, as a measure | Antislop (the profiling half) | yes — it reads outputs, not weights | count every word and 2–4-word phrase in the brew's pieces against the 269-piece human bank; anything 10× over is a tell, found rather than guessed; the same tool measures sameness across an edition (yesterday's item 6) |
| A drafter trained against announcing and flattery | Moonshot, Kimi K2 | hosted, ~2.2¢ a draft (6× Flash, ≈ Gemini) | Declined by the owner, 09-13: the rubric is from the K2 report of Jul 2025 and K2.6, the model on sale, has no published writing test. The pool stays three Flash drafts plus one Gemini, picked by code — built 09-12 as item 8. |
| Re-baseline on every model change | the Opus 5 case | yes — free | re-run the 12 fixed assignments when a writer or critic version changes, before reading any before-and-after |
| Raw prose samples, not descriptions | Sudowrite, Noren, Type | already done (the shelf) | keep; expect rhythm, not voice |
| Decode-time bans (phrase bias, backtracking sampler) | NovelAI, Antislop | no — needs the model's internals | nothing; OpenAI's crude version needed 106 tokens to stop one dash |
| Weight-level fixes (FTPO, PASTA, base-model training) | Paech; Anand; NovelAI | no | nothing today |
| Per-persona fine-tune on a hosted model | Novelcrafter (50–75 pairs); Chakrabarty (30 authors) | partly — where a provider offers fine-tuning | a pilot: one real-figure persona with a corpus, beats → the author's own paragraphs, judged blind against the same persona on the ordinary writer |
| Humanizer pass · detector gate · more never-lines | the trade | — | don't |