Where Jev fits — speed · cost · accuracy · UX research · 2026-09-22

Short answer: Jev helps -ish in two places, and it is not a cost story. Speed: every room reply today waits behind two decisions written as prose — the prop master (1.05 s at the median) and the floor producer (3.3 s) — before the panel says a word. Jev answers a decision in 0.6 s through OpenRouter from here, about 0.15 s direct; the prop master alone would take a second off every turn, and the floor producer's who speaks half could go the same way if the floor producer were split into its number half and its prose half. UX: the persona search — today a keyword filter over tags — becomes「describe who you want」: eight lay queries, eight right personas, where the filter missed three outright and ranked two badly. Cost: no — the decision-shaped calls are about a tenth of the app's bill; the writer, the panel and the pictures are the rest and Jev cannot write. Accuracy: level with Flash on the prop master (53 of 60 agree, the misses where Jev was not shown the table), no better on Ink's relevance read, and not a drop-in for the floor producer's judgement. The old ways stay; this page is the map of where the new one would slot in, in the order it should.

the terms Jev is TypeSafe AI's decision model (2026-09-15; the fact sheet is on the Ink bench page): text in, numbers out — pick one of up to 255 options with a probability on each, a score on a described scale, or a yes/no as a 0–1 probability; all questions in one pass, no prose, $0.042 per million input tokens, output free. A decision-shaped call is any call the app makes today whose answer is really a label, a rank or a yes/no, even if the model returns it in sentences. The prop master decides whether a turn wants an object on the table (a vote, a die, a board); the floor producer stages every panel turn — who speaks, in what order, at what length, in what manner. The ledger is call_events, one row per provider call; the figures below are the dev ledger and the dev event logs (June–September), not the box. Access: TypeSafe's own sign-up is full; the probes ran on the OpenRouter key already in .env, which serves Jev today.

1 · The four axes at a glance

AxisWhereWhat changesEvidenceVerdict
Speedthe front of every room turnprop master 1.05 s → ~0.2 s; the floor producer's who/hold/length-tier 3.3 s → ~0.3 s, its manner prose kept or made a per-host default2,482 prop calls p50 1.05 s / p95 1.5 s; 6,669 floor stagings (v1p) p50 3.31 s / p95 5.26 s; Jev p50 0.6 s via OpenRouter herethe prize
Costthe whole ledgerthe decision-shaped kinds ≈ 11% of dev spend; Jev would make them ~free$2.7 of $24.4 over 9,869 callssmall
Accuracyprop master · Ink relevance · floor producerprop: level with Flash; scope: stricter, not better; floor: ranks the right host, cannot draw the line§5, four probes, 178 calls, $0.0085even
UXpersona search · New chat · Ink's companya search that understands a sentence; a curated panel that appears in half a second; the proxy seat gated right8/8 queries; the Ink benchreal

2 · The map — every decision the app makes today

The ledger's own vocabulary, one row per kind that decides something, with what runs it now and the Jev shape it would take. a fit worth building · ~ a fit with a catch · not Jev's (it needs words, or it is code already).

KindWhat it decidesTodayJev shapeFitWhy
tool.propdoes this turn want an object on the table; set-up or operateFlash, 1.05 s, 4–12¢ a thousand turnsone Choice {none · setup · operate}, the table state in the state53/60 agree without seeing the table; the 7 misses were turns the prop master knew a vote had closed or a score was owed — give Jev the open instruments and it has the same facts. 0.6 s here, ~0.15 s direct.
fp.midwho speaks · order · length tier · manner · holdFlash, 3.3 s p50, 5.3 s p95a Noul per host + a Choice for the lead + a Score for the length tier + a Noul for hold; the manner stays prose~Jev's top host was among the floor producer's speakers 19/25, but its yes/no per host hovers at 0.4–0.6 and splits badly (the floor staged 48 of 87 seats, Jev 21). The room read — temperature, rhythm, asks — is the floor producer's real job and it is prose. A split (numbers from Jev, manner a per-host default from the voice: note, Flash only when needed) is a redesign with 3 s on the table, not a swap.
tool.promisedthe wake-door check: did the host promise an act it has not madeFlashone Noul~a plain yes/no on a reply; untested here, the shape fits.
photo.gateSKIP / LOOK / STUDY, which photo, what to ask the eyesFlashthe verdict a Choice, the photo a Choice; the ASK line stays prose~two of three outputs are labels; the third is a sentence for the vision model. Half a call saved, not a whole one.
kit.routethe thread router: which thread a line belongs toFlashone Choice over the open threadsrouting is the textbook case (TypeSafe's own intent-routing pattern). Untested here.
kit.ruling · kit.voice · kit.interpretera side ruling · the host's words · reading a moveFlashrulings are said out loud; a device moment is words. Jev has none.
persona.followupthe ping's judge: does this deserve a follow-upFlashone Noula judge that returns a verdict; the cheapest swap on the list.
persona.memory · persona.dreamwhat to remember; the dreamFlasha salience Score in front of the writer; the writer stays~extraction is words; deciding which lines are worth extracting is a number. A gate, not a replacement.
room.translate · room.titletranslate on divergence · the auto-titlecode + Flash · Flashthe divergence test is code; a title is prose.
app.convenethe panel curator: three panels for a situation, each seat with a whyV4 Pro / Flash, secondsa Choice over the roster ranks in 0.6 s; Flash writes the whys afterthe same shape as Ink's company (benched 12/12 on the expert seat). UX: the list is on screen before the whys arrive.
the picker searchwhich personas match what was typeda token filter over tags · role · name · tagline, in the browsera Choice over the roster on submit§6 — the clearest UX win on the page.
app.ink_scopethe relevance read: in/out · theme · closeness 0–3, ~40 stories a callFlash, $0.22 over 202 callsa Score (4 levels) + a theme Choice per story~within ±1 on 68/80, in/out 60/80, theme 52/80 — stricter than Flash, not shown better; and a brew is not latency-bound. No reason to move.
app.ink_pitchevery persona reads the pool and pitches — 357 calls at 8.6k tokens eachFlash, $0.42a Score per persona × story, fanned outthe「hundred calls a morning for the same answer」problem: Jev answers a hundred pitches for a cent. The company shortlist already replaces most of it.
app.ink_judgethe gate: the rubric read of a pieceFlash / Gemini, $1.36 over 651 calls — the largest decision-shaped billcomposite Score per rubric linethe owner's ruling stands (09-07): no model score at the register floor; the gate is advisory fine print, and the owner's read is the measure. Not a Jev question.
app.ink_metaphorthe metaphor rate: which sentences carry a live metaphorFlashone Noul per sentence, then code counts~Jev「does not count」— but a Noul a sentence sidesteps counting. The annotator's scale is ours; Jev would need calibrating to it.
app.ink_archivewhich archive passages fit the storyFlash + Serper, 1,864 calls, $3.34a Score per passage — TypeSafe's rerank recipethe biggest call count on the ledger, and reranking is what Jev's own cookbook is for. Untested here; the passage-finder on the Ink bench is the same move.
the companywho comments on a briefplanned: tags → FlashJev shortlist + reader-touched gate → Flash licencethe bench.

3 · Speed — the room turn

today human line ─▶ prop master 1.05 s ─▶ floor producer 3.3 s ─▶ panel starts writing ≈ 4.4 s before the first word step 1 human line ─▶ Jev prop 0.2–0.6 s ─▶ floor producer 3.3 s ─▶ panel ≈ 3.5–3.9 s step 2 human line ─▶ Jev: prop · who · hold · tier, one call 0.2–0.6 s ─▶ manner from the voice: note ─▶ panel ≈ 0.5 s (step 2 is a redesign of the floor producer: its numbers move to Jev, its prose becomes a default per host, and Flash is called only when the read is hard — the accepted 3 s debt of the flatness study, paid down)

Two cautions. The 0.6 s measured here is OpenRouter's proxy plus the TUN from this desk; TypeSafe direct runs 70–500 ms by their claim and 127 ms at the median in liteLLM's bench, and the box in Singapore has neither the proxy nor the wall. And a Jev prop master must be shown the table — the open ballot, the board, the score owed — as facts in its state; every miss in the probe was a turn where the prop master knew the device's state and Jev did not.

4 · Cost — the dev ledger

WhatCalls$Share
the whole dev ledger, 2026-08-21 → 09-229,86924.36100%
writing and speaking (the Ink writer, the panel, renditions, the bank, copy, Seen, the studio)≈ 5,300≈ 17.572%
pictures (Ink art, the eyes)≈ 150≈ 2.08%
decision-shaped (prop · floor · gate · scope · pitch · judge · editor · followup · convene)≈ 2,700≈ 2.711%
search fees and the rest (Serper, Brave, Bocha, STT, embeddings)≈ 1,700≈ 2.29%

Even if every decision-shaped call became free, the bill drops by a tenth. The money is in the words, and Jev has none. The one line where Jev changes the arithmetic is the pitches: a hundred Flash reads of the pool for a hundred pitches is the shape Jev exists for, and the company shortlist retires most of it anyway.

5 · Accuracy — four probes against the app's own recorded decisions

ProbeSetJev askedResultRead
the prop master60 real human turns from the dev rooms, 30 PROP · 30 NONE as Flash judged them, six lines of contextChoice {none · setup · operate}PROP/NONE agree 53/60; three-way 51/60five misses were a game answer or a closed vote where the score was owed — Jev never saw the table. Two were NONEs it made PROP (a mid-round question the host holds the answer to; 「end the game」after the table was already cleared). Level with Flash on what it was shown.
the floor producer, who speaks30 multi-host turns (v1p mostly)a Noul per host, a Choice for the leadtop host among the floor's speakers 19/25; per-seat yes/no agree 44/87; lead agree 12/30the ranking is right, the line is not: Jev's yes/no per host sits at 0.4–0.6 and a 0.5 threshold seats nobody on half the turns. The @-mention and named-host routes are code already. Not a swap.
Ink's relevance read80 stories of 09-22, 20 at each closeness 0–3Score on the brief's four levels + a theme Choiceexact 27/80 · within ±1 68/80 · in/out 60/80 · theme 52/80Jev is the stricter reader: 18 stories Flash kept it dropped, 2 the other way. No ground truth either way; nothing to gain in a brew that runs at night.
the persona search8 lay queries, English and Chinesea Choice over the 105 cards8/8 the right persona first, 0.6–0.9 s, 7.2k tokens (≈ $0.0003) a query§6.

178 calls, $0.0085, p50 0.61 s, p95 0.83 s, all through OpenRouter. Jev is not deterministic near a threshold (a 0.53 / 0.47 flip on identical input on the Ink bench) — a yes/no at 0.5 is a coin where the model is unsure; read the number, not the label.

6 · UX — the search that understands a sentence

The picker's box today is a token filter: a query is split into words, each word is looked for in the tags, the role, the name and the tagline, and a card scores by hits. It is fast and it is why tags.txt carries lay synonyms. It cannot read a sentence, and a Chinese sentence is one token to it. The same eight queries, the filter's first four hits against Jev's first pick:

TypedThe filter today (hits · first four)Jev (p)
someone to talk to about my divorce29 · Nora Ephron, Lu Xun, Shreyas Doshi, Yu HuaEsther Perel .85, Gary Chapman .12
帮我看看上海的房子该不该现在买0 hitsXiao Pang 1.0
who can explain black holes to my kid45 · Hawking, Hamaguchi, the Buddha, Lu XunHawking .84, Feynman .14
a coach for my tennis serve51 · Mouratoglou, Chestnut, Nadal, DoshiMouratoglou 1.0
help me name my startup and pitch investors57 · Moritz, Liang Ning, Nora Ephron, MungerMoritz .31, Sean Ellis .26, Taku .22
I want to argue about whether AI will take my job58 · Nora Ephron, Son, Wolfram, KarpathyHinton .48, Karpathy .21, Harari .14
写一封给老同学的信,文笔要好0 hitsYu Hua .74, Lu Xun .15
someone who knows Singapore politics25 · Augustus, Confucius, Orwell, DidionLee Kuan Yew 1.0

The shape that fits the app: the filter stays for keystrokes (free, instant, offline), and a sentence — three words or more, or any Chinese — goes to Jev on submit, with the filter's list re-ordered by Jev's probabilities. Three more UX places with the same shape: New chat (the curator's three panels appear in half a second, the whys stream in after); Ink's company (the reader's proxy seat switched on only when the story touches their money, home, health or safety — the gate the bench found); and the room's first word, which §3 is about.

7 · What Jev must not do here

Anything said.
The manner cue, the licence passage, a ruling, a title, a why. Jev returns numbers; the words stay with Flash, Sonnet or the panel model.
The floor producer's judgement.
The firewall (staging, never content) is safe with Jev by construction — a number cannot carry a subject — but the read of the room is not a number yet; the probe says Jev ranks hosts well and draws the line badly.
The gate's score.
The owner ruled no model score at the register floor. A cheaper model score is still a model score.
A decision without its facts.
Jev reads what it is given and「answers the question you wrote」. The prop master's misses were missing facts, not missing sense; the table state, the open threads, the ledger of who spoke go into the state or the answer is a guess.

8 · The recommendation, in order

recommend 1 · The persona search on submit — standalone, no room risk, the largest UX gain per line of code; the filter stays for keystrokes. 2 · The prop master — the same Choice as the probe with the table state in the state, behind the existing fail-open door; a second off every turn. 3 · Ink's company shortlist and proxy gateas benched. 4 · The ping's judge and the thread router — plain yes/no and routing, cheap to try. Later, and as a design: the floor producer split (§3 step 2) and the archive rerank. Not: the relevance read, the gate, anything that writes. Prerequisite: a key — the OpenRouter route works today on an alpha endpoint; TypeSafe direct when the sign-up opens, pinned to jev-1.13.0.

9 · Open questions

QuestionWhy it matters
Direct latency from the box0.6 s here is the proxy's number. If Singapore → TypeSafe direct is ~0.2 s, step 1 of §3 is worth a full second; if it is 0.6 s, half of one.
The prop master with the tablethe probe withheld the device state; the seven misses need re-running with it before「level with Flash」becomes「better than」.
The floor producer splithow much of the staging's value is the manner cue? If most, step 2 saves the seconds and loses the room. An A/B on the exam battery (exam/run.py --quick) is the gate.
ChineseTypeSafe says CJK works with reduced accuracy; the search probe's two Chinese queries were right, the Ink bench's one Chinese brief was right. A Chinese-only room is untested.
The key and the endpointOpenRouter's Decisions endpoint is marked alpha and has no SDK; a shape change breaks a caller silently. TypeSafe direct is the road once the waitlist opens.
Thresholds driftJev's probabilities are the model's, not ours; a 0.8 gate tuned on jev-1.13 must be re-read on jev-1.14. Pin the version.

10 · Status

WhenWhat
2026-09-22 · researchThe owner: TypeSafe's sign-in is full, the old ways stay — but where could Jev improve the app: speed, cost, accuracy, UX? This page: the ledger's decision-shaped kinds mapped to Jev shapes; the recorded latencies of the prop master and the floor producer from 2,482 and 6,669 dev calls; four probes (178 Jev calls through the OpenRouter key, $0.0085) against the app's own recorded verdicts; the order to build in. No code changed.
research 2026-09-22 · siblings: Jev, the decision model (the Ink bench) · the usage dashboard · models by function · the floor producer · omni search