Short answer: Jev helps -ish in two places, and it is not a cost story. Speed: every room reply today waits behind two decisions written as prose — the prop master (1.05 s at the median) and the floor producer (3.3 s) — before the panel says a word. Jev answers a decision in 0.6 s through OpenRouter from here, about 0.15 s direct; the prop master alone would take a second off every turn, and the floor producer's who speaks half could go the same way if the floor producer were split into its number half and its prose half. UX: the persona search — today a keyword filter over tags — becomes「describe who you want」: eight lay queries, eight right personas, where the filter missed three outright and ranked two badly. Cost: no — the decision-shaped calls are about a tenth of the app's bill; the writer, the panel and the pictures are the rest and Jev cannot write. Accuracy: level with Flash on the prop master (53 of 60 agree, the misses where Jev was not shown the table), no better on Ink's relevance read, and not a drop-in for the floor producer's judgement. The old ways stay; this page is the map of where the new one would slot in, in the order it should.
call_events, one row per provider call; the figures below are the dev ledger and the dev event logs (June–September), not the box. Access: TypeSafe's own sign-up is full; the probes ran on the OpenRouter key already in .env, which serves Jev today.| Axis | Where | What changes | Evidence | Verdict |
|---|---|---|---|---|
| Speed | the front of every room turn | prop master 1.05 s → ~0.2 s; the floor producer's who/hold/length-tier 3.3 s → ~0.3 s, its manner prose kept or made a per-host default | 2,482 prop calls p50 1.05 s / p95 1.5 s; 6,669 floor stagings (v1p) p50 3.31 s / p95 5.26 s; Jev p50 0.6 s via OpenRouter here | the prize |
| Cost | the whole ledger | the decision-shaped kinds ≈ 11% of dev spend; Jev would make them ~free | $2.7 of $24.4 over 9,869 calls | small |
| Accuracy | prop master · Ink relevance · floor producer | prop: level with Flash; scope: stricter, not better; floor: ranks the right host, cannot draw the line | §5, four probes, 178 calls, $0.0085 | even |
| UX | persona search · New chat · Ink's company | a search that understands a sentence; a curated panel that appears in half a second; the proxy seat gated right | 8/8 queries; the Ink bench | real |
The ledger's own vocabulary, one row per kind that decides something, with what runs it now and the Jev shape it would take. ✓ a fit worth building · ~ a fit with a catch · ✗ not Jev's (it needs words, or it is code already).
| Kind | What it decides | Today | Jev shape | Fit | Why |
|---|---|---|---|---|---|
| tool.prop | does this turn want an object on the table; set-up or operate | Flash, 1.05 s, 4–12¢ a thousand turns | one Choice {none · setup · operate}, the table state in the state | ✓ | 53/60 agree without seeing the table; the 7 misses were turns the prop master knew a vote had closed or a score was owed — give Jev the open instruments and it has the same facts. 0.6 s here, ~0.15 s direct. |
| fp.mid | who speaks · order · length tier · manner · hold | Flash, 3.3 s p50, 5.3 s p95 | a Noul per host + a Choice for the lead + a Score for the length tier + a Noul for hold; the manner stays prose | ~ | Jev's top host was among the floor producer's speakers 19/25, but its yes/no per host hovers at 0.4–0.6 and splits badly (the floor staged 48 of 87 seats, Jev 21). The room read — temperature, rhythm, asks — is the floor producer's real job and it is prose. A split (numbers from Jev, manner a per-host default from the voice: note, Flash only when needed) is a redesign with 3 s on the table, not a swap. |
| tool.promised | the wake-door check: did the host promise an act it has not made | Flash | one Noul | ~ | a plain yes/no on a reply; untested here, the shape fits. |
| photo.gate | SKIP / LOOK / STUDY, which photo, what to ask the eyes | Flash | the verdict a Choice, the photo a Choice; the ASK line stays prose | ~ | two of three outputs are labels; the third is a sentence for the vision model. Half a call saved, not a whole one. |
| kit.route | the thread router: which thread a line belongs to | Flash | one Choice over the open threads | ✓ | routing is the textbook case (TypeSafe's own intent-routing pattern). Untested here. |
| kit.ruling · kit.voice · kit.interpreter | a side ruling · the host's words · reading a move | Flash | — | ✗ | rulings are said out loud; a device moment is words. Jev has none. |
| persona.followup | the ping's judge: does this deserve a follow-up | Flash | one Noul | ✓ | a judge that returns a verdict; the cheapest swap on the list. |
| persona.memory · persona.dream | what to remember; the dream | Flash | a salience Score in front of the writer; the writer stays | ~ | extraction is words; deciding which lines are worth extracting is a number. A gate, not a replacement. |
| room.translate · room.title | translate on divergence · the auto-title | code + Flash · Flash | — | ✗ | the divergence test is code; a title is prose. |
| app.convene | the panel curator: three panels for a situation, each seat with a why | V4 Pro / Flash, seconds | a Choice over the roster ranks in 0.6 s; Flash writes the whys after | ✓ | the same shape as Ink's company (benched 12/12 on the expert seat). UX: the list is on screen before the whys arrive. |
| the picker search | which personas match what was typed | a token filter over tags · role · name · tagline, in the browser | a Choice over the roster on submit | ✓ | §6 — the clearest UX win on the page. |
| app.ink_scope | the relevance read: in/out · theme · closeness 0–3, ~40 stories a call | Flash, $0.22 over 202 calls | a Score (4 levels) + a theme Choice per story | ~ | within ±1 on 68/80, in/out 60/80, theme 52/80 — stricter than Flash, not shown better; and a brew is not latency-bound. No reason to move. |
| app.ink_pitch | every persona reads the pool and pitches — 357 calls at 8.6k tokens each | Flash, $0.42 | a Score per persona × story, fanned out | ✓ | the「hundred calls a morning for the same answer」problem: Jev answers a hundred pitches for a cent. The company shortlist already replaces most of it. |
| app.ink_judge | the gate: the rubric read of a piece | Flash / Gemini, $1.36 over 651 calls — the largest decision-shaped bill | composite Score per rubric line | ✗ | the owner's ruling stands (09-07): no model score at the register floor; the gate is advisory fine print, and the owner's read is the measure. Not a Jev question. |
| app.ink_metaphor | the metaphor rate: which sentences carry a live metaphor | Flash | one Noul per sentence, then code counts | ~ | Jev「does not count」— but a Noul a sentence sidesteps counting. The annotator's scale is ours; Jev would need calibrating to it. |
| app.ink_archive | which archive passages fit the story | Flash + Serper, 1,864 calls, $3.34 | a Score per passage — TypeSafe's rerank recipe | ✓ | the biggest call count on the ledger, and reranking is what Jev's own cookbook is for. Untested here; the passage-finder on the Ink bench is the same move. |
| the company | who comments on a brief | planned: tags → Flash | Jev shortlist + reader-touched gate → Flash licence | ✓ | the bench. |
Two cautions. The 0.6 s measured here is OpenRouter's proxy plus the TUN from this desk; TypeSafe direct runs 70–500 ms by their claim and 127 ms at the median in liteLLM's bench, and the box in Singapore has neither the proxy nor the wall. And a Jev prop master must be shown the table — the open ballot, the board, the score owed — as facts in its state; every miss in the probe was a turn where the prop master knew the device's state and Jev did not.
| What | Calls | $ | Share |
|---|---|---|---|
| the whole dev ledger, 2026-08-21 → 09-22 | 9,869 | 24.36 | 100% |
| writing and speaking (the Ink writer, the panel, renditions, the bank, copy, Seen, the studio) | ≈ 5,300 | ≈ 17.5 | 72% |
| pictures (Ink art, the eyes) | ≈ 150 | ≈ 2.0 | 8% |
| decision-shaped (prop · floor · gate · scope · pitch · judge · editor · followup · convene) | ≈ 2,700 | ≈ 2.7 | 11% |
| search fees and the rest (Serper, Brave, Bocha, STT, embeddings) | ≈ 1,700 | ≈ 2.2 | 9% |
Even if every decision-shaped call became free, the bill drops by a tenth. The money is in the words, and Jev has none. The one line where Jev changes the arithmetic is the pitches: a hundred Flash reads of the pool for a hundred pitches is the shape Jev exists for, and the company shortlist retires most of it anyway.
| Probe | Set | Jev asked | Result | Read |
|---|---|---|---|---|
| the prop master | 60 real human turns from the dev rooms, 30 PROP · 30 NONE as Flash judged them, six lines of context | Choice {none · setup · operate} | PROP/NONE agree 53/60; three-way 51/60 | five misses were a game answer or a closed vote where the score was owed — Jev never saw the table. Two were NONEs it made PROP (a mid-round question the host holds the answer to; 「end the game」after the table was already cleared). Level with Flash on what it was shown. |
| the floor producer, who speaks | 30 multi-host turns (v1p mostly) | a Noul per host, a Choice for the lead | top host among the floor's speakers 19/25; per-seat yes/no agree 44/87; lead agree 12/30 | the ranking is right, the line is not: Jev's yes/no per host sits at 0.4–0.6 and a 0.5 threshold seats nobody on half the turns. The @-mention and named-host routes are code already. Not a swap. |
| Ink's relevance read | 80 stories of 09-22, 20 at each closeness 0–3 | Score on the brief's four levels + a theme Choice | exact 27/80 · within ±1 68/80 · in/out 60/80 · theme 52/80 | Jev is the stricter reader: 18 stories Flash kept it dropped, 2 the other way. No ground truth either way; nothing to gain in a brew that runs at night. |
| the persona search | 8 lay queries, English and Chinese | a Choice over the 105 cards | 8/8 the right persona first, 0.6–0.9 s, 7.2k tokens (≈ $0.0003) a query | §6. |
178 calls, $0.0085, p50 0.61 s, p95 0.83 s, all through OpenRouter. Jev is not deterministic near a threshold (a 0.53 / 0.47 flip on identical input on the Ink bench) — a yes/no at 0.5 is a coin where the model is unsure; read the number, not the label.
The picker's box today is a token filter: a query is split into words, each word is looked for in the tags, the role, the name and the tagline, and a card scores by hits. It is fast and it is why tags.txt carries lay synonyms. It cannot read a sentence, and a Chinese sentence is one token to it. The same eight queries, the filter's first four hits against Jev's first pick:
| Typed | The filter today (hits · first four) | Jev (p) |
|---|---|---|
| someone to talk to about my divorce | 29 · Nora Ephron, Lu Xun, Shreyas Doshi, Yu Hua | Esther Perel .85, Gary Chapman .12 |
| 帮我看看上海的房子该不该现在买 | 0 hits | Xiao Pang 1.0 |
| who can explain black holes to my kid | 45 · Hawking, Hamaguchi, the Buddha, Lu Xun | Hawking .84, Feynman .14 |
| a coach for my tennis serve | 51 · Mouratoglou, Chestnut, Nadal, Doshi | Mouratoglou 1.0 |
| help me name my startup and pitch investors | 57 · Moritz, Liang Ning, Nora Ephron, Munger | Moritz .31, Sean Ellis .26, Taku .22 |
| I want to argue about whether AI will take my job | 58 · Nora Ephron, Son, Wolfram, Karpathy | Hinton .48, Karpathy .21, Harari .14 |
| 写一封给老同学的信,文笔要好 | 0 hits | Yu Hua .74, Lu Xun .15 |
| someone who knows Singapore politics | 25 · Augustus, Confucius, Orwell, Didion | Lee Kuan Yew 1.0 |
The shape that fits the app: the filter stays for keystrokes (free, instant, offline), and a sentence — three words or more, or any Chinese — goes to Jev on submit, with the filter's list re-ordered by Jev's probabilities. Three more UX places with the same shape: New chat (the curator's three panels appear in half a second, the whys stream in after); Ink's company (the reader's proxy seat switched on only when the story touches their money, home, health or safety — the gate the bench found); and the room's first word, which §3 is about.
jev-1.13.0.| Question | Why it matters |
|---|---|
| Direct latency from the box | 0.6 s here is the proxy's number. If Singapore → TypeSafe direct is ~0.2 s, step 1 of §3 is worth a full second; if it is 0.6 s, half of one. |
| The prop master with the table | the probe withheld the device state; the seven misses need re-running with it before「level with Flash」becomes「better than」. |
| The floor producer split | how much of the staging's value is the manner cue? If most, step 2 saves the seconds and loses the room. An A/B on the exam battery (exam/run.py --quick) is the gate. |
| Chinese | TypeSafe says CJK works with reduced accuracy; the search probe's two Chinese queries were right, the Ink bench's one Chinese brief was right. A Chinese-only room is untested. |
| The key and the endpoint | OpenRouter's Decisions endpoint is marked alpha and has no SDK; a shape change breaks a caller silently. TypeSafe direct is the road once the waitlist opens. |
| Thresholds drift | Jev's probabilities are the model's, not ours; a 0.8 gate tuned on jev-1.13 must be re-read on jev-1.14. Pin the version. |
| When | What |
|---|---|
| 2026-09-22 · research | The owner: TypeSafe's sign-in is full, the old ways stay — but where could Jev improve the app: speed, cost, accuracy, UX? This page: the ledger's decision-shaped kinds mapped to Jev shapes; the recorded latencies of the prop master and the floor producer from 2,482 and 6,669 dev calls; four probes (178 Jev calls through the OpenRouter key, $0.0085) against the app's own recorded verdicts; the order to build in. No code changed. |