Jev, the decision model — can it seat the company? bench · 2026-09-22

Short answer: yes, for the shortlist; not for the whole job. Jev is real — TypeSafe AI's「decision model」, out 2026-09-15 — and it does one thing very well here: shown the whole roster of 105 personas and one brief, it ranks who the story belongs to, in under a second, for a hundredth of a cent. It cannot write the licence (it returns numbers, never words), it cannot find「the outsider」(asked for a far lens it names the watch repairer and Roger Federer on every story), and its yes/no on「does this touch the reader's own money, home, health or safety」is the best proxy gate we have tried. So: Jev makes the shortlist and gates the proxy seat, Flash makes the pick and cites the passage. Measured on the twelve 09-22 dev briefs, ≈0.25¢ a brief, ≈3¢ a morning.

the terms A decision model (TypeSafe's phrase is「System One」) is a model that reads text and answers a fixed question with a number, not with prose — which of these options, how much on this scale, yes or no — all questions in one pass, each answer with a probability. A brief is one item of the morning's edition, the block JSON of the Background page (gist · story · questions · numbers · people). The company is who speaks on a brief — the seats of reading with the masters §1b: the subject, the expert, the outsider, the reader's proxy. The licence is the rule that a seat must be earned by a passage of the persona's own profile, quoted word for word and checked in code — no passage, no seat. Flash is DeepSeek V4.1 Flash, the brew's cheap model.

1 · What Jev is

FactWhat we found
WhoTypeSafe AI, San Francisco; founder Diogo Almeida (ex-OpenAI). Two years in stealth, a $40 m seed led by DCVC. Jev is their first model; early access opened 2026-09-15, reported open signup with $5 free credit by the 21st.
What kindNot a chat model, not an embedding, not a reranker in the usual sense. A non-autoregressive model: you send a state (any text or JSON) and a set of typed questions; it answers them all at once, as numbers. Three question types: Choice (pick one of up to 255 options, with a probability on every option and a confidence), Score (a place on 2–10 described levels), Noul (a yes/no as a 0–1 probability). Trained, they say, for calibrated probabilities rather than for pleasing text.
The APIPOST https://api.typesafe.ai/v1/systemone, a bearer key. 64k tokens a request, 32k for the state plus the longest question. Text only. Returns answers keyed by your question names, plus usage. Python and JS SDKs. Served through OpenRouter (/api/v1/systemone, model typesafe/jev-1.13), Vercel's AI Gateway and Cloudflare too.
Price$0.042 per million input tokens, output free — the same through OpenRouter. Our whole-roster pick call is ~15.6k tokens → $0.00066. Rate limit 250k tokens/s, 1,200 requests/min.
LatencyTypeSafe claims 70–500 ms. Measured from the dev PC through OpenRouter: 0.6–1.2 s a call, median ~0.8 s (the proxy and the TUN in the path). An independent bench (liteLLM) saw a 127 ms median direct.
LanguagesTypeSafe's models page: English is primary; other languages「including CJK scripts work but with reduced accuracy — test on your content」. The one Chinese-titled brief (无锡) ranked fine (Xiao Pang 0.59, Li Ka-shing 0.35). A Chinese-heavy morning is untested.
Known rough edges (their own page)Reads literally (「answers the question you wrote, not the one you meant」); accuracy falls as the state fills with text unrelated to the decision; no counting, no date arithmetic, no double negatives; not hostile to injected text by default; a black box — no reason comes back with the number.
From China / from the boxDev PC: api.typesafe.ai resolves to the TUN's fake IP and answers in ~0.9 s (a 403「supply an API key」— reachable); OpenRouter the same. The box is in Singapore, so no wall in the way; both services are US-hosted, expect ~0.3–0.5 s. Not tested from the box (research only). TypeSafe has no China presence; a mainland box would need the same proxy as everything else.
Access todayNo TypeSafe key in .env; the bench ran on the owner's existing OpenRouter key through OpenRouter's Decisions endpoint (marked alpha). No account was created. Pin the model version (jev-1.13.0) if thresholds are tuned.

2 · The task, and the three ways tried

One brief in, two to four seats out: the expert (1–2, whose field this is), the outsider (0–1, another field with a lens on this one), the reader's proxy (0–1, the person on the street, only when the story touches the reader's own money, home, health or safety). Every seat must be licensed by a passage of the profile. The roster: the 92 repo personas plus 13 app-built ones (ink_brew.roster(), both roots), Tess excluded.

A · the baseline code shortlist: tags.txt × the brief's beat, gist, story, questions, the notebook's people and terms → 8 near + 3 far (drawn from the zero-score rest) → ONE Flash call (the gist, the questions, each candidate's card line + three code-picked profile passages) → picks with a quoted licence → the licence checked in code B · Jev alone ONE Jev call, the whole roster as the options of two Choices (expert · outsider) + one Noul (reader touched) + one Choice over the everyday personas (proxy) → the top of each list; then ONE more Jev call on the top expert: a Noul per profile paragraph —「does this passage license the seat?」— the passage-finder C · the hybrid Jev's top-6 expert + top-2 outsider = the shortlist, Jev's outsider top-3 = the far three → the same Flash call as A → the same code check

Judged three ways: (a) the seats the owner named in §1b (the crater: Hawking, Einstein; the mortgages: Xiao Pang, Max Rockbank, an outsider Dan Wang or Feddo, the proxy Xiao Pang) and, for the other ten, the seat a reader would expect (my read, marked as such); (b) whether the pick's licence is a passage that truly speaks to the subject, not merely a string found in the profile; (c) cost and seconds per brief. Everything below is one run, cached; the total spend was about six cents.

3 · The bench — twelve briefs, 2026-09-22

BriefExpected expertA · tags → FlashB · Jev, top-3 (p)C · Jev → Flash
New York startup funding rises, share fallsMoritz, TakuMoritz · Sean Ellis ~Moritz .89 · Taku .11 Moritz · Taku ✓✓
Anthropic weighs new model ahead of IPOKarpathy (at Anthropic), HintonKarpathy · Peak Ji ~Hinton .45 · Karpathy .22 · Musk .09 Hinton · Karpathy ✓✓; outsider Buffett (the IPO)
New Moon crater, bigger than the Colosseumowner: Hawking, Einstein; outsider KeynesTerence Tao · Darwin (Hawking was on the shortlist and passed over)Hawking .49 · Wiwiwik .20 · Musk .16 Einstein absentFeynman · Cameron (Hawking first on the list, passed over again); outsider Liu Cixin
Bay Area buyers weigh 7% mortgagesowner: Xiao Pang, Max Rockbank; Dan Wang / Feddo; proxy Xiao PangXiao Pang ; outsider Dan Wang ; proxy Max Rockbank (an agent is not the street)Xiao Pang .34 · Keynes .22 · Max Rockbank .17 ✓✓; touched .86Xiao Pang · Max Rockbank ✓✓; outsider Keynes; proxy Xiao Pang ~ (one seat, not two)
无锡 8月新房均价涨、成交跌 (zh)Xiao Pang; proxy the streetXiao Pang ; Max Rockbank; proxy Chen Huilan Xiao Pang .59 · Li Ka-shing .35 ; touched .86Xiao Pang (a Chinese licence) · Max Rockbank ; Keynes; Chen Huilan
Miami-Dade, Palm Beach $10 m+ recordsMax Rockbank, Li Ka-shing; no proxyMax Rockbank · Xiao Pang; outsider Seneca ; proxy a watch repairer Max Rockbank .96 ; touched .75 Max Rockbank · Xiao Pang ; Seneca; proxy Xiao Pang
Antitrust suit targets the AI slowdown pactAmbika Kumar (litigator), Taku, HintonMoritz · Karpathy no lawyer on the shortlist; outsider chen-huilan-test Musk .33 (Musk v. Altman) · Kumar .30 · Hinton .19 Hinton · Ambika Kumar ✓✓; outsider Chen Huilan
Silversmith backs Vantora, $100 mMoritz, Taku, SonMoritz · Sean Ellis ~Moritz .67 · Taku .21 · Son .05 Moritz · Taku ✓✓; outsider Chen Huilan
Lunar water too scarce for a Moon cityMusk (the subject), Hawking; outsider Liu CixinMusk · Porter ; Holmes; proxy Xiao Pang Musk .91 · Hawking .05 ; touched .11 Musk · Hawking ✓✓; Buffett; proxy Xiao Pang
Historic duck club listed, $1.5 mMax RockbankMax Rockbank · Li Ka-shing; outsider Charlie Rose Max Rockbank .86 ; touched .28Max Rockbank · Xiao Pang ; outsider Darwin (the marsh)
Anthropic may release a model before the IPOKarpathy, HintonKarpathy (licence real:「he now works at Anthropic」) · Jensen; Holmes; proxy Xiao Pang Karpathy .27 · Hinton .21 · Musk .18 ; passage-finder .80 「at Anthropic since 19 May 2026」Karpathy · Hinton ✓✓; Buffett; proxy Xiao Pang
Ukraine's largest drone attack on MoscowFeddo, Dan Wang (analysts); Churchill (history)FDR · Churchill ~ (FDR's licence:「he cannot stand or walk unaided」); Cameron; proxy Xiao Pang Churchill .29 · Feddo .18 · Harari .15 ; passage-finder: nothing above .13 — no licence, honestlyFeddo · Keynes ; Chen Huilan; proxy Xiao Pang
Tally (my read)A · tags → FlashB · Jev aloneC · Jev → Flash
the expert seat holds the field (of 12)812 (top-2)11
the outsider is a real lens, not a stranger (of 12)315
the proxy seat on only when it should be (of 12)6 (code test too loose)10 (Noul ≥ 0.8)7 (same code test)
licence found in the profile, by string (seats)39 / 4044 / 44
licence truly speaks to the subject (my read)≈ half3 of 12 tops; honest lows≈ two thirds

4 · What Jev did well, and what it did not

The expert ranking is the win.
One Choice over all 105 cards — name, role, tagline, eight tags — put the field's own persona first or second on every brief: Hawking on the crater, Xiao Pang on both housing stories, Max Rockbank on the three property listings, Moritz on both funding stories, Karpathy and Hinton on both Anthropic stories, Ambika Kumar second on the antitrust suit, Feddo second on Moscow. The tag shortlist, by contrast, put Shreyas Doshi, Jack Dorsey and Jony Ive on the mortgages, Stan Lee and Jeff Koons on Moscow, and no lawyer on the lawsuit — a word-overlap on 105 sets of tags is noisy, and Flash can only pick from what it is shown. Where Jev's list was clean, Flash's pick was clean (arm C).
The outsider question fails.
Asked「which persona from a field this story does NOT belong to would see it differently」, Jev names the retired watch repairer, Roger Federer and Rafael Nadal on nearly every brief, at confidence 0.1. It reads the question literally — the most unrelated card — and「a lens on this one」is not a thing it can weigh. Liu Cixin on the crater (0.14) was the one hit. A far seat needs the far candidates drawn in code and Flash asked whether any of them would add.
The reader-touched Noul is the best proxy gate tried.
0.86 on both housing stories, 0.08 on the crater, 0.11 on the Moon city, 0.20 on both funding rounds; the misses are Miami's $10 m sales (0.75) and the antitrust suit (0.76 — the subscription price, arguably). At ≥ 0.8 it switches the seat on for exactly the two stories it should. The code test in the baseline (a「my」or「should I」in the questions) switched it on for nine of twelve, and Flash then sat Xiao Pang on the Moon, on Moscow and on the Anthropic IPO.
The proxy Choice is weak.
Among the everyday personas Jev has no ground to stand on: a Tokyo student tops most briefs, and on the Bay Area mortgages Xiao Pang (0.30) ties with a lingerie company secretary (0.29). Keep the seat's candidates a curated list and let Flash pick, or default to the reader's own persona.
The passage-finder is half a tool.
One Noul per profile paragraph (24–34 a persona, $0.0003, 0.7 s) finds a real licence when one exists — Karpathy「at Anthropic since 19 May 2026」0.80, Musk's own trial 0.53, Xiao Pang「what changed him was the market」0.57 — and scores honestly low when none does: Churchill on Moscow 0.13, Hawking on a crater 0.35 (his top passage is the wheelchair over the toes). It also likes biography tables (Moritz's birthplace row, 0.71). As a second gate under Flash's quoted licence it would catch the far-fetched ones; as the sole licence it cannot, because it returns a number and the rule wants the words.
The string check is not the licence.
Both Flash arms pass the code check almost always because the model copies from the passages it was shown — and still cites「he cannot stand or walk unaided」to seat Roosevelt on drone warfare, or「the only child of a country store」to seat Charlie Rose on a duck club. Found in the profile ≠ speaks to the subject. The passage-finder's number is the cheapest second opinion.
Two smaller things.
Jev is not deterministic: the same four-option probe returned Hawking 0.53 / Einstein 0.47 twice and Einstein 0.51 / Hawking 0.49 the third time — take the top three, never only the top one. And a Choice caps at 255 options, so the roster fits in one question until it passes 255 personas; after that, chunk and merge.

5 · Cost and seconds, per brief

ArmCallsTokens in$ a brief¢ a morning (12)Seconds
A · tags → Flash1 Flash≈ 8,7000.00202.52.1
B · Jev alone1 Jev pick + 1 Jev passage-finder15,600 + 6,3000.00091.11.6
C · Jev → Flash1 Jev pick + 1 Flash15,600 + 8,3000.00253.03.0

Flash measured at the off-peak sheet ($0.22 / M in, $0.66 / M out; the calls ran at 17:26 CST, inside the 2× window, so the morning's real figure is what is shown). Jev's cost is OpenRouter's own usage.cost. The Jev pick call carries the roster twice (once per Choice); dropping the outsider Choice halves it to ~9k tokens. Nothing here is a cost question: any arm is under a nickel a morning.

6 · The recommendation

recommendUse Jev for the shortlist and the proxy gate; Flash for the pick and the licence. Concretely: ONE Jev call a brief — the expert Choice over the whole roster (top-6 → the shortlist) and the reader-touched Noul (≥ 0.8 → the proxy seat is on); drop Jev's outsider Choice and draw the three far candidates in code as §1b already says; then the ONE Flash call as planned, with the code-picked passages and the quoted licence, checked in code. Cost ≈ 0.2¢ a brief, three seconds. The subject seat stays a code match on the notebook's people. Owed before any build: the far-seat prompt, and whether the passage-finder's number (≥ 0.5) becomes a second gate under the quoted licence.

Why not Jev alone: it cannot write the licence, the rule the owner set so a match is never a stretch; and it cannot find the outsider. Why not the baseline alone: its tag shortlist is the weak link — it never showed Flash the lawyer, or Feddo, or Max Rockbank on the Bay Area — and Flash cannot seat who it does not see. The hybrid is the same one Flash call the plan already budgets, with a cleaner list in front of it, for six-hundredths of a cent more.

7 · Open questions

QuestionWhy it matters
The licence rule against the obvious seatHawking's profile has no passage on impacts or craters; by「no passage, no seat」he does not speak on the crater the owner named him for. The rule or the seat — the owner's call. (The same for Churchill on Moscow.)
Which keyOpenRouter works today on the owner's existing key, but the endpoint is marked alpha and has no SDK; a TypeSafe account ($5 free credit reported) is the direct road. Either way, pin jev-1.13.0.
Chinese morningsTypeSafe says CJK works with reduced accuracy; one zh-titled brief with English blocks ranked fine. A brief whose blocks are Chinese is untested.
The reader's-own-first rule§1b says a persona the reader chats with often takes the seat over an equal stranger. That is a re-rank in code on Jev's probabilities (a bonus to the reader's own), not tested here.
Yield after two daysSame: a code re-rank on the ledger, after Jev, before Flash.
From the boxUntested by design. Singapore → US, no wall; expect well under a second. Watch OpenRouter's alpha endpoint for shape changes.

8 · Status

WhenWhat
2026-09-22 · the benchThe owner asked whether Jev, just launched, could pick who comments on a brief. Found and read (vendor docs, the API reference, the jaggedness page, three independent benches); reached from the dev PC; run through the owner's OpenRouter key on the twelve 09-22 dev briefs in three arms, every call cached, ≈6¢ spent. Verdict above. No code in lib/ changed; the bench script and cache live in the session scratchpad, not the repo.

Sources

TypeSafe: the launch post · System One · the API reference · models (limits, languages) · jev-1.13 rough edges · the 182-option cookbook. Independent: Simon Willison · liteLLM's bench · awesome-typesafe-jev (the ordering and calibration audits) · DataCamp. Gateways: OpenRouter · Vercel.

bench 2026-09-22 · siblings: reading with the masters · the brief line · the Background page