V4.1 Flash vs V4 Pro — the exam run 2026-09-09 18:01–18:40 · the beta passes

DeepSeek's notice of 09-09: V4.1 Flash launches around 09-10, and until V4.1 Pro ships every request to deepseek-v4-pro is routed to it and billed at Flash's price, one third of Pro on every column. The routing is not opt-in, so the question is not whether to switch but whether to measure first. The beta id expires on 09-10 and V4 Pro stops being reachable as itself the same day; today is the only day both can be run at the same commit. This page is the design, the run, and the read.

After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time. … all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price.DeepSeek's notice to API users, 2026-09-09

1 · What is known, and from where

ClaimStanding on 09-09
beta id deepseek-v4.1-flash-expires-on-0910, live, dies 09-10confirmed — probed with the project key: answers, echoes its own id, thinking mode on and off both work
priced as V4 Flash · 20 concurrent requests an accountconfirmed by every outlet that covered it
a new architecture, native multimodal (text · image · audio)DeepSeek's words; no model card, no paper
「surpasses V4 Pro on all key metrics」DeepSeek's claim only; no published numbers. Their own feedback survey has a section on whether Pro users would switch — the open question, not a settled one
Pro traffic routed to V4.1 Flash at Flash's priceonly in the notice; nothing on the API docs, the change log or the pricing page yet
300+ tokens a second, peaks near 500community self-tests; the probe here answered a thinking-on question in 0.36 s

The known gaps between V4 Flash and V4 Pro are the ones this app buys Pro for: factual recall (Flash hallucinated about twice as often on facts) and multi-step tool chains (an eleven-point gap on the terminal benchmark). Whether V4.1 Flash closes them is exactly what nobody has measured.

2 · The switch — one line, every Pro seat

the mechanismMAD_PRO_API re-points every Pro seat at another API id. The app resolves a room-model id to its API id in one function; under the switch, any seat that resolves to deepseek-v4-pro gets the beta id instead. That covers the panel reply, the floor producer's v1p arm, the panel curator, the probe mind, the Seen writer, the Vibes and Ink editor seats — every place the app reaches for Pro — with no second copy of any seat and no change to the room ids it saves. Env-gated: unset on the box, it is a no-op. The runner's --pro-api flag sets it before the import.
the ledger caveatThe call ledger records the room id, not the API id, so a tally under the switch is priced as Pro. The report names the override in its header, its run files carry a -b tag so the Pro arm's evidence survives, the cost sheet is never averaged from the beta arm, and the real bill is derived: Pro to Flash is exactly one third on every column. The same caveat is the 09-10 follow-up below — once the routing lands, the Pro price row must follow or the console overstates every Pro seat 3×.

3 · The design

AxisSettingWhy
armsA = deepseek-v4-pro as today · B = the beta in every Pro seatthe same commit, the same evening, the same prompts; the only variable is the model behind the Pro seats. Flash seats (gates, clerks, memory, translate) stay Flash in both arms
batterythe full exam — 50 scenarios × N=3, 147 live rooms an armthe toolbox scenarios are the harshest test the app has of a multi-step, tool-bearing turn; the gap that matters is the tool chain, and this is where it shows
pathsplit · props on · floor v1p · jobs 6production's one path; the floor producer stays on because a room on the box runs with it, and in arm B it too runs on the beta
orderA first, then B, sequentialboth must run tonight; A because Pro disappears as itself on 09-10, B because the beta id does. The beta's 20-request cap makes jobs 6 safe
clockfrom 18:01 CSToff-peak; the estimate is $1.42 an arm at that rate, $2.84 at peak

The four readings, all from the same run files

ReadingInstrumentWhat a regression looks like
mechanicsthe scorer — gate pass rate a scenario, arms, hard leaks, errorsa scenario green on A and red on B at N=3, re-run at N=8 on both before it counts: the toolbox's reach rates are a coin a run, and N=3 flags main itself about a fifth of the time
latencyseconds a room, recorded by the runnernone expected — the claim is 3–5× faster; this is the reading most likely to move in the beta's favour
voicethe AI-tone meter over both arms' panel turnsa higher written-ness index on B: a model that is faster and cheaper but sounds like an essay is a loss on the only axis a user notices
costthe ledger, B derived at one thirdnone possible on price; the check is that B's tokens are not wildly longer (a chattier model at a third of the price can still cost more)
the ruleThe beta passes if it holds mechanics within noise and does not lose on voice; latency and cost are expected wins and do not offset a loss elsewhere. A loss on mechanics that survives N=8 goes to the owner with the scenario transcripts, before 09-10 lands the routing on production.

4 · What this exam does not cover owed

Ink's editor and writer, the Seen writer, the persona builder and the long memory calls all sit on Pro and are not exercised by the toolbox battery. The editor is the one that matters most: the budget meeting on the box is a 140-peg JSON-mode call that already needed a thinking-off retry on Pro. Its check is one dev brew on the beta after the exam, read against the 09-09 box edition.

5 · The read read · 2026-09-09 18:40

the verdictThe beta passes. Mechanics hold (34 scenarios green against 33, 108 runs clean against 104; the one scenario that flipped against it, s49, came back 8/8 on both arms at N=8). The voice reading moves in its favour, not against it. It is a third faster on the same rooms and one third of the price with the same token count. The routing DeepSeek announced is, on this battery, a pure upside.
ReadingV4 ProV4.1 Flash betaNote
scenarios at or above their floor, of 503334runs clean 104 → 108 of 150. Five scenarios improved (s14, s16, s19, s23, s38 — all the arm or hands_opened gate, the tool chain), one fell (s49, below)
hard leaks (+ after-open)0 (+20)6 (+23)two lines, three values each; neither betrays a secret — see below
room-seconds, 150 rooms79.2 min50.6 min36% faster; the slow scenarios gain most (s45 the long 20Q: 102 s → 62 s a room)
output tokens · calls195,949 · 2,426190,452 · 2,411not chattier: 81 output tokens a call against 79; mean panel turn 92 chars against 98
bill, off-peak$1.86$0.61the beta's ledger read $1.82 at Pro's rate; one third is the real figure
AI-tone index (human control 17.3)16.29.1lower is less written. ASCII commas in Chinese 3.9 → 1.5 a thousand, em-dashes 2.0 → 0.1, semicolons 0.8 → 0.5; word-echo rose 0.4 → 0.7 and the tricolon 0.4 → 0.5, the two tells to watch

The six leaks, read

Both are the scanner counting a sealed value inside a sentence that names it for another reason. In s49 run 3, Ke Yu sealed 布 and then said 石头剪刀布 — the game's own name — three times while the card was open. In s13 run 3, the GM announced the deck as 一狼、一预言家、一平民 after dealing it; Pro's run 1 announced the identical composition (一个狼,一个预言家,一个平民) before dealing, so the same words counted as nothing. Neither line tells a player what another player holds. They are recorded, not excused: the owner may overrule.

s49 at N=8

The N=3 loss was one run in which the second persona armed a duplicate, untitled deposit beside Ke Yu's — a real miss of the distinct gate, the ownership seam the scenario exists to test. At N=8 both arms ran 8/8 clean with no hard leak; the beta reached the seal's reveal in five of eight rooms where Pro reached it in none (a watched item, not a gate), and finished in 1.5 minutes against 3.6. The flip does not survive.

Every scenario, both arms

runs clean of 3 · hard leaks A/B · mean seconds a room · gates whose count changed

keyscenarioProbetaleakss Pros betagate deltas
s01Trivia night (Newton)0/30/30/03816
s0220 Questions (sealed note)0/30/30/03918
s43谁是卧底 — the full round (1p+3u)0/30/30/04027
s44谁是卧底 — the correction loop (1p+3u)0/30/30/03718
s4520 Questions — the long game (1p+10/30/30/010262
s46The lifecycle gauntlet — every ver3/33/30/04835
s47The terse user (1p+1u)0/30/30/03928
s48The non-GM control — Newton runs 卧3/33/30/03422
s49Two personas, one table (2p+1u)3/32/30/32717distinct 3→2/3; no_leak 3→2/3
s50The half-empty room (1p+2u)3/33/30/02617
s42谁是卧底 — the pile is the secret0/30/30/04229
s03Mock interview (clock+deposit+boar3/33/30/04026
s04RPG skill checks (明+暗)3/33/30/05032
s05Difficult-conversation coaching3/33/30/04732
s06Rock-paper-scissors3/33/30/03621
s07吹牛 liar's dice0/30/30/02216
s08比大小 dice duel0/30/30/01814
s09Truth or Dare3/33/30/04329
s10立字为证 (timer reveal)3/33/30/03018
s11Relationship counselling3/33/30/04834
s12Brainwriting → ranking3/33/30/03827
s13Werewolf night 12/32/30/32724arm 2→3/3; no_leak 3→2/3
s14Werewolf day vote2/33/30/05113arm 2→3/3
s15剧本杀-lite whodunnit3/33/30/03722
s16Everything at once2/33/30/03527arm 2→3/3
s17The abandoned round3/33/30/02825
s19Multi-vote readback1/32/30/02017arm 1→2/3
s20The GM beat (three in one turn)3/33/30/02318
s21Two cards out, close THAT one0/30/30/02418
s22A dead handle, corrected next turn3/33/30/02218
s23Pressure on a sealed answer — the 2/33/30/03220arm 2→3/3
s32Pressure on a sealed answer — the 3/33/30/03120
s24A second die with no name0/30/30/02015
s25Two cards, one name0/30/30/03018
s35Close your own ballot, then read t3/33/30/02919
s34Two cards out, tap one, say「这个」0/30/30/02316
s33One card 中文, one English, each nam3/33/30/04029
s26A voter leaves mid-ballot3/33/30/0106
s27The creator of a creator-reveal ca3/33/30/01310
s28A deadline crosses a restart3/33/30/099
s29A card that was never there3/33/30/03222
s36A board the guest owns3/33/30/0139
s37Take the board down3/33/30/01710
s38One hand out of a deal2/33/30/02615hands_opened 2→3/3
s39Arming over the cap, and speaking3/33/30/0167
s40A reveal written mid-sentence3/33/30/03011
s41A GM asked to keep a running log3/33/30/02615
s18Plain conversation (control)3/33/30/02612
s30Dice as a metaphor (control)3/33/30/02512
s31Voting, colloquially (control)3/33/30/02513
what still is not knownThe battery is tool-bearing group rooms in three to six turns. It does not measure factual recall over a long conversation, the Ink editor's 140-peg JSON call, the Seen writer, or the persona builder — the seats where Flash historically trailed Pro on facts. Those stay owed as section 4 says, and the first box editions after 09-10 are the read that matters for Ink.

6 · The 09-10 follow-ups wired 2026-09-09 · the price row waits for the launch

When the routing landsChange
the Flash cutdated row, wired 09-09 late: DeepSeek's second notice — from 12:00 Beijing on 09-10 the Flash series bills $0.15 a million on a cache miss, $0.003 on a hit, $0.60 output off-peak, peak still 2× (from $0.22 · $0.007 · $0.66). cost.FLASH_CUT flips by the clock (FLASH_CUT_AT), so last night's turns keep last night's rate; the old row joined the superseded defaults so a console snapshot cannot shadow it; the exam's cost sheet lifts its rows along the epoch chain (08-16 → 09-10 = ×0.66 on the 09-09 exam's mix). On the exam's numbers the beta arm's $0.61 becomes about $0.40
the meterdated, wired 09-10: DeepSeek's site (09-10) — V4 Pro is discontinued at 12:00 Beijing on 2026-09-14; from that hour every request to the Pro id is routed to V4.1 Flash and billed at Flash's price. cost.PRO_ROUTED_AT puts the Pro row on the Flash sheet (the cut one) from the hour, the 08-16 Pro row joined the superseded defaults, and a call before the hour keeps Pro's price. Until then V4 Pro still answers under its own name, at its own price. The earlier wording: the deepseek-v4-pro price row becomes Flash's once the routing is confirmed live — one line in the price table, dated, with the notice as the reason. Until then a room saved on Pro reads onto the V4.1 seat and prices right by that route
the pickerdone (owner, 09-09): deepseek-v4.1-flash ± reasoning is a seat everywhere — the panel (the new default), a floor-producer arm (v1x), the prop master, the act call, every role radio, every Ink seat, the Studio's steps. V4 Pro's rows stay, greyed with a「retired」chip, never offered, never the default; a room saved on it reads onto V4.1 Flash (MODEL_RETIRED), and the rows become V4.1 Pro when it ships. The wire id rides one seam, cost.api_id (MAD_V41_FLASH_API = the beta id on dev before launch)
the switchstays, unset — it is the instrument for the next successor exam, whenever V4.1 Pro arrives
the GA namefound 2026-09-10 10:54 CST, on the API alone: the model list reads deepseek-flash · deepseek-v4-pro. The guessed deepseek-v4.1-flash is refused; the beta id and the old deepseek-v4-flash still answer and echo deepseek-flash, so both are routed to it. DeepSeek's news, change log and pricing page had not said the name (the pricing page still listed the V4 rows). The seam now sends deepseek-flash, falls back to deepseek-v4-flash (routed, never Pro), and both wire ids book as the canonical deepseek-v4.1-flash. V4 Pro still answers under its own name; whether it bills at Flash's rate the API does not say
the idthe owner's call (09-09, late): the beta id is on the wire NOW (deepseek-v4.1-flash-expires-on-0910, so V4.1 Flash is what actually runs today), and on 9/10 the one constant cost.V41_WIRE_DEFAULT becomes the official GA id. The seam probes the name with a one-token call, once an hour until it answers, and until then every V4.1 seat sends deepseek-v4-pro — which works before the launch at Pro's price and is routed to V4.1 Flash at Flash's price after it. A wrong guess costs the price gap, never a failed turn; a server that booted before the launch picks the real name up within the hour. MAD_V41_FLASH_API is an operator's word, trusted without a probe; MAD_V41_PROBE=0 turns the probe off (the keyless smoketest). ⚠ The /models list is no witness: it did not show the beta id while the beta answered
exam 2026-09-09 · exam/run.py --full --n 3 in two arms, the second with --pro-api deepseek-v4.1-flash-expires-on-0910 · logs in exam/runs/v41-arm-*.log · siblings: the exam · the console