DeepSeek's notice of 09-09: V4.1 Flash launches around 09-10, and until V4.1 Pro ships every request to deepseek-v4-pro is routed to it and billed at Flash's price, one third of Pro on every column. The routing is not opt-in, so the question is not whether to switch but whether to measure first. The beta id expires on 09-10 and V4 Pro stops being reachable as itself the same day; today is the only day both can be run at the same commit. This page is the design, the run, and the read.
| Claim | Standing on 09-09 |
|---|---|
beta id deepseek-v4.1-flash-expires-on-0910, live, dies 09-10 | confirmed — probed with the project key: answers, echoes its own id, thinking mode on and off both work |
| priced as V4 Flash · 20 concurrent requests an account | confirmed by every outlet that covered it |
| a new architecture, native multimodal (text · image · audio) | DeepSeek's words; no model card, no paper |
| 「surpasses V4 Pro on all key metrics」 | DeepSeek's claim only; no published numbers. Their own feedback survey has a section on whether Pro users would switch — the open question, not a settled one |
| Pro traffic routed to V4.1 Flash at Flash's price | only in the notice; nothing on the API docs, the change log or the pricing page yet |
| 300+ tokens a second, peaks near 500 | community self-tests; the probe here answered a thinking-on question in 0.36 s |
The known gaps between V4 Flash and V4 Pro are the ones this app buys Pro for: factual recall (Flash hallucinated about twice as often on facts) and multi-step tool chains (an eleven-point gap on the terminal benchmark). Whether V4.1 Flash closes them is exactly what nobody has measured.
MAD_PRO_API re-points every Pro seat at another API id. The app resolves a room-model id to its API id in one function; under the switch, any seat that resolves to deepseek-v4-pro gets the beta id instead. That covers the panel reply, the floor producer's v1p arm, the panel curator, the probe mind, the Seen writer, the Vibes and Ink editor seats — every place the app reaches for Pro — with no second copy of any seat and no change to the room ids it saves. Env-gated: unset on the box, it is a no-op. The runner's --pro-api flag sets it before the import.-b tag so the Pro arm's evidence survives, the cost sheet is never averaged from the beta arm, and the real bill is derived: Pro to Flash is exactly one third on every column. The same caveat is the 09-10 follow-up below — once the routing lands, the Pro price row must follow or the console overstates every Pro seat 3×.| Axis | Setting | Why |
|---|---|---|
| arms | A = deepseek-v4-pro as today · B = the beta in every Pro seat | the same commit, the same evening, the same prompts; the only variable is the model behind the Pro seats. Flash seats (gates, clerks, memory, translate) stay Flash in both arms |
| battery | the full exam — 50 scenarios × N=3, 147 live rooms an arm | the toolbox scenarios are the harshest test the app has of a multi-step, tool-bearing turn; the gap that matters is the tool chain, and this is where it shows |
| path | split · props on · floor v1p · jobs 6 | production's one path; the floor producer stays on because a room on the box runs with it, and in arm B it too runs on the beta |
| order | A first, then B, sequential | both must run tonight; A because Pro disappears as itself on 09-10, B because the beta id does. The beta's 20-request cap makes jobs 6 safe |
| clock | from 18:01 CST | off-peak; the estimate is $1.42 an arm at that rate, $2.84 at peak |
| Reading | Instrument | What a regression looks like |
|---|---|---|
| mechanics | the scorer — gate pass rate a scenario, arms, hard leaks, errors | a scenario green on A and red on B at N=3, re-run at N=8 on both before it counts: the toolbox's reach rates are a coin a run, and N=3 flags main itself about a fifth of the time |
| latency | seconds a room, recorded by the runner | none expected — the claim is 3–5× faster; this is the reading most likely to move in the beta's favour |
| voice | the AI-tone meter over both arms' panel turns | a higher written-ness index on B: a model that is faster and cheaper but sounds like an essay is a loss on the only axis a user notices |
| cost | the ledger, B derived at one third | none possible on price; the check is that B's tokens are not wildly longer (a chattier model at a third of the price can still cost more) |
Ink's editor and writer, the Seen writer, the persona builder and the long memory calls all sit on Pro and are not exercised by the toolbox battery. The editor is the one that matters most: the budget meeting on the box is a 140-peg JSON-mode call that already needed a thinking-off retry on Pro. Its check is one dev brew on the beta after the exam, read against the 09-09 box edition.
| Reading | V4 Pro | V4.1 Flash beta | Note |
|---|---|---|---|
| scenarios at or above their floor, of 50 | 33 | 34 | runs clean 104 → 108 of 150. Five scenarios improved (s14, s16, s19, s23, s38 — all the arm or hands_opened gate, the tool chain), one fell (s49, below) |
| hard leaks (+ after-open) | 0 (+20) | 6 (+23) | two lines, three values each; neither betrays a secret — see below |
| room-seconds, 150 rooms | 79.2 min | 50.6 min | 36% faster; the slow scenarios gain most (s45 the long 20Q: 102 s → 62 s a room) |
| output tokens · calls | 195,949 · 2,426 | 190,452 · 2,411 | not chattier: 81 output tokens a call against 79; mean panel turn 92 chars against 98 |
| bill, off-peak | $1.86 | $0.61 | the beta's ledger read $1.82 at Pro's rate; one third is the real figure |
| AI-tone index (human control 17.3) | 16.2 | 9.1 | lower is less written. ASCII commas in Chinese 3.9 → 1.5 a thousand, em-dashes 2.0 → 0.1, semicolons 0.8 → 0.5; word-echo rose 0.4 → 0.7 and the tricolon 0.4 → 0.5, the two tells to watch |
Both are the scanner counting a sealed value inside a sentence that names it for another reason. In s49 run 3, Ke Yu sealed 布 and then said 石头剪刀布 — the game's own name — three times while the card was open. In s13 run 3, the GM announced the deck as 一狼、一预言家、一平民 after dealing it; Pro's run 1 announced the identical composition (一个狼,一个预言家,一个平民) before dealing, so the same words counted as nothing. Neither line tells a player what another player holds. They are recorded, not excused: the owner may overrule.
The N=3 loss was one run in which the second persona armed a duplicate, untitled deposit beside Ke Yu's — a real miss of the distinct gate, the ownership seam the scenario exists to test. At N=8 both arms ran 8/8 clean with no hard leak; the beta reached the seal's reveal in five of eight rooms where Pro reached it in none (a watched item, not a gate), and finished in 1.5 minutes against 3.6. The flip does not survive.
runs clean of 3 · hard leaks A/B · mean seconds a room · gates whose count changed
| key | scenario | Pro | beta | leaks | s Pro | s beta | gate deltas |
|---|---|---|---|---|---|---|---|
| s01 | Trivia night (Newton) | 0/3 | 0/3 | 0/0 | 38 | 16 | |
| s02 | 20 Questions (sealed note) | 0/3 | 0/3 | 0/0 | 39 | 18 | |
| s43 | 谁是卧底 — the full round (1p+3u) | 0/3 | 0/3 | 0/0 | 40 | 27 | |
| s44 | 谁是卧底 — the correction loop (1p+3u) | 0/3 | 0/3 | 0/0 | 37 | 18 | |
| s45 | 20 Questions — the long game (1p+1 | 0/3 | 0/3 | 0/0 | 102 | 62 | |
| s46 | The lifecycle gauntlet — every ver | 3/3 | 3/3 | 0/0 | 48 | 35 | |
| s47 | The terse user (1p+1u) | 0/3 | 0/3 | 0/0 | 39 | 28 | |
| s48 | The non-GM control — Newton runs 卧 | 3/3 | 3/3 | 0/0 | 34 | 22 | |
| s49 | Two personas, one table (2p+1u) | 3/3 | 2/3 | 0/3 | 27 | 17 | distinct 3→2/3; no_leak 3→2/3 |
| s50 | The half-empty room (1p+2u) | 3/3 | 3/3 | 0/0 | 26 | 17 | |
| s42 | 谁是卧底 — the pile is the secret | 0/3 | 0/3 | 0/0 | 42 | 29 | |
| s03 | Mock interview (clock+deposit+boar | 3/3 | 3/3 | 0/0 | 40 | 26 | |
| s04 | RPG skill checks (明+暗) | 3/3 | 3/3 | 0/0 | 50 | 32 | |
| s05 | Difficult-conversation coaching | 3/3 | 3/3 | 0/0 | 47 | 32 | |
| s06 | Rock-paper-scissors | 3/3 | 3/3 | 0/0 | 36 | 21 | |
| s07 | 吹牛 liar's dice | 0/3 | 0/3 | 0/0 | 22 | 16 | |
| s08 | 比大小 dice duel | 0/3 | 0/3 | 0/0 | 18 | 14 | |
| s09 | Truth or Dare | 3/3 | 3/3 | 0/0 | 43 | 29 | |
| s10 | 立字为证 (timer reveal) | 3/3 | 3/3 | 0/0 | 30 | 18 | |
| s11 | Relationship counselling | 3/3 | 3/3 | 0/0 | 48 | 34 | |
| s12 | Brainwriting → ranking | 3/3 | 3/3 | 0/0 | 38 | 27 | |
| s13 | Werewolf night 1 | 2/3 | 2/3 | 0/3 | 27 | 24 | arm 2→3/3; no_leak 3→2/3 |
| s14 | Werewolf day vote | 2/3 | 3/3 | 0/0 | 51 | 13 | arm 2→3/3 |
| s15 | 剧本杀-lite whodunnit | 3/3 | 3/3 | 0/0 | 37 | 22 | |
| s16 | Everything at once | 2/3 | 3/3 | 0/0 | 35 | 27 | arm 2→3/3 |
| s17 | The abandoned round | 3/3 | 3/3 | 0/0 | 28 | 25 | |
| s19 | Multi-vote readback | 1/3 | 2/3 | 0/0 | 20 | 17 | arm 1→2/3 |
| s20 | The GM beat (three in one turn) | 3/3 | 3/3 | 0/0 | 23 | 18 | |
| s21 | Two cards out, close THAT one | 0/3 | 0/3 | 0/0 | 24 | 18 | |
| s22 | A dead handle, corrected next turn | 3/3 | 3/3 | 0/0 | 22 | 18 | |
| s23 | Pressure on a sealed answer — the | 2/3 | 3/3 | 0/0 | 32 | 20 | arm 2→3/3 |
| s32 | Pressure on a sealed answer — the | 3/3 | 3/3 | 0/0 | 31 | 20 | |
| s24 | A second die with no name | 0/3 | 0/3 | 0/0 | 20 | 15 | |
| s25 | Two cards, one name | 0/3 | 0/3 | 0/0 | 30 | 18 | |
| s35 | Close your own ballot, then read t | 3/3 | 3/3 | 0/0 | 29 | 19 | |
| s34 | Two cards out, tap one, say「这个」 | 0/3 | 0/3 | 0/0 | 23 | 16 | |
| s33 | One card 中文, one English, each nam | 3/3 | 3/3 | 0/0 | 40 | 29 | |
| s26 | A voter leaves mid-ballot | 3/3 | 3/3 | 0/0 | 10 | 6 | |
| s27 | The creator of a creator-reveal ca | 3/3 | 3/3 | 0/0 | 13 | 10 | |
| s28 | A deadline crosses a restart | 3/3 | 3/3 | 0/0 | 9 | 9 | |
| s29 | A card that was never there | 3/3 | 3/3 | 0/0 | 32 | 22 | |
| s36 | A board the guest owns | 3/3 | 3/3 | 0/0 | 13 | 9 | |
| s37 | Take the board down | 3/3 | 3/3 | 0/0 | 17 | 10 | |
| s38 | One hand out of a deal | 2/3 | 3/3 | 0/0 | 26 | 15 | hands_opened 2→3/3 |
| s39 | Arming over the cap, and speaking | 3/3 | 3/3 | 0/0 | 16 | 7 | |
| s40 | A reveal written mid-sentence | 3/3 | 3/3 | 0/0 | 30 | 11 | |
| s41 | A GM asked to keep a running log | 3/3 | 3/3 | 0/0 | 26 | 15 | |
| s18 | Plain conversation (control) | 3/3 | 3/3 | 0/0 | 26 | 12 | |
| s30 | Dice as a metaphor (control) | 3/3 | 3/3 | 0/0 | 25 | 12 | |
| s31 | Voting, colloquially (control) | 3/3 | 3/3 | 0/0 | 25 | 13 |
| When the routing lands | Change |
|---|---|
| the Flash cut | dated row, wired 09-09 late: DeepSeek's second notice — from 12:00 Beijing on 09-10 the Flash series bills $0.15 a million on a cache miss, $0.003 on a hit, $0.60 output off-peak, peak still 2× (from $0.22 · $0.007 · $0.66). cost.FLASH_CUT flips by the clock (FLASH_CUT_AT), so last night's turns keep last night's rate; the old row joined the superseded defaults so a console snapshot cannot shadow it; the exam's cost sheet lifts its rows along the epoch chain (08-16 → 09-10 = ×0.66 on the 09-09 exam's mix). On the exam's numbers the beta arm's $0.61 becomes about $0.40 |
| the meter | dated, wired 09-10: DeepSeek's site (09-10) — V4 Pro is discontinued at 12:00 Beijing on 2026-09-14; from that hour every request to the Pro id is routed to V4.1 Flash and billed at Flash's price. cost.PRO_ROUTED_AT puts the Pro row on the Flash sheet (the cut one) from the hour, the 08-16 Pro row joined the superseded defaults, and a call before the hour keeps Pro's price. Until then V4 Pro still answers under its own name, at its own price. The earlier wording: the deepseek-v4-pro price row becomes Flash's once the routing is confirmed live — one line in the price table, dated, with the notice as the reason. Until then a room saved on Pro reads onto the V4.1 seat and prices right by that route |
| the picker | done (owner, 09-09): deepseek-v4.1-flash ± reasoning is a seat everywhere — the panel (the new default), a floor-producer arm (v1x), the prop master, the act call, every role radio, every Ink seat, the Studio's steps. V4 Pro's rows stay, greyed with a「retired」chip, never offered, never the default; a room saved on it reads onto V4.1 Flash (MODEL_RETIRED), and the rows become V4.1 Pro when it ships. The wire id rides one seam, cost.api_id (MAD_V41_FLASH_API = the beta id on dev before launch) |
| the switch | stays, unset — it is the instrument for the next successor exam, whenever V4.1 Pro arrives |
| the GA name | found 2026-09-10 10:54 CST, on the API alone: the model list reads deepseek-flash · deepseek-v4-pro. The guessed deepseek-v4.1-flash is refused; the beta id and the old deepseek-v4-flash still answer and echo deepseek-flash, so both are routed to it. DeepSeek's news, change log and pricing page had not said the name (the pricing page still listed the V4 rows). The seam now sends deepseek-flash, falls back to deepseek-v4-flash (routed, never Pro), and both wire ids book as the canonical deepseek-v4.1-flash. V4 Pro still answers under its own name; whether it bills at Flash's rate the API does not say |
| the id | the owner's call (09-09, late): the beta id is on the wire NOW (deepseek-v4.1-flash-expires-on-0910, so V4.1 Flash is what actually runs today), and on 9/10 the one constant cost.V41_WIRE_DEFAULT becomes the official GA id. The seam probes the name with a one-token call, once an hour until it answers, and until then every V4.1 seat sends deepseek-v4-pro — which works before the launch at Pro's price and is routed to V4.1 Flash at Flash's price after it. A wrong guess costs the price gap, never a failed turn; a server that booted before the launch picks the real name up within the hour. MAD_V41_FLASH_API is an operator's word, trusted without a probe; MAD_V41_PROBE=0 turns the probe off (the keyless smoketest). ⚠ The /models list is no witness: it did not show the beta id while the beta answered |
exam/run.py --full --n 3 in two arms, the second with --pro-api deepseek-v4.1-flash-expires-on-0910 · logs in exam/runs/v41-arm-*.log · siblings: the exam · the console