Jev on the floor — does a decision model help the producer? built · opt-in · 2026-09-23

Short answer: a little, and not where it was expected. The floor producer now has a fourth arm, Jev (v1j): the decision model answers the producer's questions — how many voices, who leads, should each host speak, how long, what kind of turn — and code writes the staging line from the numbers. It works, it fails safely, it is a quarter cheaper and half a second faster than today's Flash arm. But the three-second prize the research page counted on is gone: that was the retired Pro arm. Today's default (v1x, V4.1 Flash) stages in 1.4 s; Jev in 0.9 s. And in a blind quality read of four rooms per arm the two are a wash with opposite faults: Flash seats two and one of them often only agrees; Jev seats one and the second angle never comes. It stays opt-in. The default stays v1x.

the terms The floor producer stages every panel turn before the personas speak: who speaks, in what order, at what length, in what manner, or whether the panel holds. Its arms: Code (v0, no model), V4.1 Flash (v1x, today's default, a prose call), V4 Pro (v1p, retired), and now Jev (v1j). Jev is TypeSafe's decision model (the fact sheet): text in, typed numbers out, never a sentence. The staging line is the one thing the panel reads. The firewall is the rule that the producer emits staging, never content. The flatness study (the floor producer page, Part 9) is the yardstick: a blind judge scores redundancy, concision, conflict and repeated gestures against the users' own complaints.

1 · What was built

A fourth segment on the chat's floor picker and on the console's floor default (Models → fp mid), opt-in, greyed when the server has no Jev key. When a room runs v1j, the producer's whole brief — the roster, the recency row, the mentions, the furniture, the transcript tail, this turn's words — goes to Jev as one call with these questions:

voices Choice none · one · two · three-or-more how many hosts speak lead Choice over the seated hosts who speaks first and gets the length speak_X Noul one per host should X speak at all (the ranking) tier Score minimal · tight · medium · generous · solo-deep the lead's ceiling mode Choice quick fact · decide · vent · teach · banter · challenge · game move · aside · plan · story · contention · deep self-work distress Noul warmth and space, never a push

Code then writes the line in v0's grammar: the lead at the tier's ceiling (≤65 words for medium, ≤140 字 in a Chinese room), a second voice on one tight line, the rest silent, a tone drawn from a closed table keyed by the mode and the host's own voice: note. The rules stay code's: an @-mention or a typed name seats the addressed host only, whatever Jev ranked; never everyone in a room of three or more; solo-deep only when one host holds the room; and a hold must be said twice — Jev's「none」and every host's own speak score under 0.35, with two or more humans and no object wanted. Any failure falls to v0, exactly as the Flash arm does. The words of the questions live in lib/prompts/floor_producer.jev.json.

The firewall holds by construction: a number cannot carry a subject, so no manner cue can script a host's point. That is the one place Jev is not merely as safe as the prose arm but safer.

2 · The probe — 80 recorded turns

Before building, the same questions were asked of Jev on 80 multi-host turns the Flash producer had actually staged in the dev rooms, with a brief shaped like the producer's own.

DecisionAgreement with the live producerRead
who leads44 of 60where the producer's lead could be matched by name (the other 20 were written in Chinese); its top-ranked host was among the staged speakers on most of the rest
the length tier67 of 80 within one stepJev runs a step longer than Flash (medium and generous where Flash wrote tight)
how many voices19 of 80the live producer seated two on 76 of 80; Jev said one on 37 and nobody on 25 — 14 of those in single-human rooms where a hold is illegal. The voices question reads literally.
the kind of turnquick fact 16 · banter 15 · contention 9 · challenge 9 · plan 9 · teach 7 — sane labels, no ground truth
the call0.78 s median, 1.96 s p952,067 tokens a call, $0.0001

3 · The A/B — four rooms, two arms, one blind judge

The flatness study's own driver ran four of its English scenarios — a founder deciding on a term sheet (3 hosts), dinner-table banter (3), a wife's vent (2), and a cold multi-human product room (3) — twice: the producer on v1x and on v1j, the panel on V4.1 Flash in both, the opening turn staged by the same prose producer in both (v1j is a mid-chat arm). Nineteen staged turns per arm.

Measured (mechanical)v1x · Flashv1j · Jev
producer seconds, median / p951.43 / 2.160.91 / 1.64
falls to v000
hosts speaking per turn, mean2.051.11
three-host turns where all three spoke40
turns nobody answered00 (2 before the two-part hold rule)
the lead's length over its cap, median0.600.68
a second voice's words, median3536
the four rooms' bill$0.0228$0.0173
The blind judge (the study's spec, one sub-agent a room, arms hidden)v1x · Flashv1j · Jev
non-lead lines judged268
of them redundant or filler7 (27%)3 (38%)
distinct angles per multi-host turn, mean0.870.77
lines judged tight82%70%
lines judged padded8%9%
conflict under the natural dial (gap, 0 = as promised)+0.25+0.25
manufactured conflict · repeated gestures0 · 00 · 0

What the judges wrote, room by room. Flash: 「the closing turn's second host re-paraphrases the lead」 (vent) · 「one non-lead re-asks the lead's condition in new words」 (decide) · 「the panel converges on the same idea three times; the decision-rule turn is the worst offender」 (multi-human) · 「the opening wastes two of three non-lead bubbles」 (banter). Jev: 「five of six turns single-host, so little to be redundant against; both hosts in lockstep」 (vent) · 「every later turn collapses to one responder」 (decide) · 「a three-host panel produced one real opinion」 (banter; the two empty lines are the shared opening's) · 「mostly single-host; the one second voice restates the lead — but Steve's 'I disagree with most of that' is real friction」 (multi-human).

Read together: Jev's three「redundant」lines out of eight are mostly the opening turn's, which both arms share; on its own staged turns Jev put a second voice on the table twice and one of those restated the lead. Flash put a second voice on the table on nearly every turn and one in four of them only agreed. Jev cures the pile-up by not holding one; the second angle the study prizes goes with it. Four rooms a side is a read, not a verdict.

4 · What this settles

The latency prize was smaller than the research page said.
The 3.3 s median on the map page was the Pro arm's over the summer; the Pro arm retired on 09-09 and v1x on V4.1 Flash stages in 1.4 s. Jev saves half a second a turn, not three. Real, not decisive.
Jev ranks well and counts badly.
Who leads and how long: usable. How many voices: the question reads literally and「none」comes out where a person would never hold. The two-part rule fixed the held turns; the one-voice bias stays, because Jev's second-host scores sit at 0.4–0.6 and the count question says one.
The manner cue survives being a table.
The judges found no manufactured conflict, no repeated gestures, and the same padding rate as the prose arm. A closed tone vocabulary keyed by the kind of turn did the prose cue's job on these rooms.
Cheaper, by a quarter, on the whole room.
$0.0173 against $0.0228 for four rooms — the producer's own call is a rounding error either way; the saving is the second host not speaking.

5 · The recommendation

recommendKeep v1j as an opt-in arm; the default stays v1x. Use it where the room wants one voice and speed — a solo host, a quick-fact room, a game table where the referee's silence matters — and to test the firewall-by-construction idea further. Do not make it the default until the second voice comes back: the next design is v1j with v0's rhythm — the count from Code's shape deck (which already varies one · two · full house across turns), the ranking and the tier from Jev, so the question Jev answers worst is never asked of it. That is a small change to the composer and its own A/B.

6 · Open questions

QuestionWhy it matters
Four rooms a sideThe mechanical numbers are stable; the judge's are not at this n. The study ran nine scenarios in three languages; a Chinese room and a five-host room are untested on v1j.
The openingv1j is mid-chat only; the opening producer stays prose (v1x). A decision-model opening is a different question — composition, not a turn read.
Return rateThe v2 manners variant is judged on return rate, not on a judge. v1j could ride the same stamp if it ever goes live for a slice.
The key on the boxSame as the prop master: no OpenRouter or TypeSafe key in the box's environment yet; the seat greys there until one is added.

7 · Status

WhenWhat
2026-09-23 · built + A/BThe owner:「let's see if Jev helps the fp」. The 80-turn probe; the v1j arm (the questions file, the composer, the four-segment picker, the console default, 15 smoketest checks incl. the mention rule, the two-part hold, the never-everyone rule, the solo-deep cap, fail-open); the four-room A/B on the study's driver with a blind judge on nine transcripts (≈$0.05 of rooms, ≈$0.01 of Jev); the two-part hold rule from the first run's two held turns. On main, opt-in, not the default, not on the box.
built 2026-09-23 · siblings: the floor producer · where Jev fits · Jev, the decision model · conversation quality