Dialogue · Design notes · Feedback widget
← Design notes

Feedback widget — the v1 instrument

How we measure the floor producer when there's no user base to A/B and the real conversations are full of personal information. The answer: collect the signal in-product, from the user, about their own turns — so nothing private ever leaves the room. This is the measurement loop for a 3-tester world, and the first thing to build. A working mock-up — click the thumbs and tags.

Per bubble, one rating eachA rail of bare icons inside each reply's bubble — every host's turn is rated on its own. The tag panel is born folded, so a panel of three isn't three open cards.
Three armsThe same widget runs under no FP v0 v1 — blind to the user — so feedback also tells us if FP helps at all.
A survey = the long formEvery N turns (admin-set; 0 = off) a longer version of the same widget pops up for the cumulative read.

In the room

The rating rides inside each reply's bubble: three bare icons — 👍 👎 ⌄ — in one quiet right-aligned row under the text, and nothing below the bubble until it's needed (the fold-out panel is born folded and costs zero height, so the stream keeps its constant 15px rhythm). A clicked thumb becomes a neutral pill — deliberately not green/red, which drown against the bubble's own background — read per theme: a quiet translucent ink-wash in dark mode, a lifted translucent-white keycap (with the idle rail faded further back) in light mode, where an ink wash looked dirty on cream — the same material both ways, a wash over the bubble rather than a solid chip. 👍 is one tap, no expand; 👎 lights and opens the tag panel below the bubble, where the lever-tags sit in fixed columns so the user learns where each one lives. The chevron folds and unfolds. Opening the panel scrolls it fully into view — clear of the frosted top bar and composer, under which the chat scrollport runs — and Add a note focuses the box with the caret at the end, then nudges it above the composer / mobile keyboard once the shell's keyboard resize settles. While the note box has focus the composer hides — never two text areas on screen at once — and returns the moment focus leaves the note. In a bubble that carries a link chip (a DOC / diagram hand-over) the rail sits below the chip: the thumbs are always the bubble's last row. An open panel auto-folds when the next message lands (releasing the note's focus with it) — so even in a busy panel the cards never pile up open. (This rail-in-bubble form is CW's ui/sandbox design, adopted v266 with two changes: one button size across iOS/Android/PC, and the ink-wash pressed state in place of her green/red icon tint.)

Founders' panelarm: v1 · admin-only, hidden from the user
Should I raise VC money, or bootstrap?
Warren Buffett
Bootstrapping keeps you in control and forces real margins. If the business can't fund itself early, raising money mostly delays the reckoning — I'd start lean and let the numbers earn the next dollar.
Jensen Huang
It depends how big the prize is. In a land-grab market, capital is speed, and speed wins — sometimes raising to move first is the whole game.
LENGTH
TONE
CONTENT
— Munger is sitting this one out —

Two replies: Buffett's rail idle (three bare icons in the bubble's corner, nothing below), Jensen's panel open — what ⌄ or 👎 reveals: lever-tags in fixed columns, LENGTH · TONE · CONTENT. (Try it — thumbs and chips toggle; the chevron folds.)

The widget, up close

A thumb is the cheap signal; the tags are the actionable one — each maps to a lever the producer controls, so a complaint points straight at what to change. Here it is mid-use:

Jensen Huang
Capital is speed, and speed wins — sometimes raising to move first is the whole game.
LENGTH
TONE
CONTENT
grouptagleverwhat it tells the producer
LENGTHtoo long / too shortlengthshrink / grow the budget for this kind of turn
TONEtoo softchallenge ↑wanted to be pushed, got coddled (host → you)
TONEtoo harshchallenge ↓over-grilled — back off
CONTENTnot relevantfocusdrifted off the real ask
CONTENTechoes othersfreshnessthis reply added no angle — it rehashed the panel or an earlier turn (the #1 pain: the redundancy)
✎ note + 👍/👎the catch-all (takeaway, pile-on, anything…) + the overall read

The two CONTENT tags are independent — a reply can be both off-ask and a rehash. Each opposed pair — too long / too short and too soft / too harsh — is mutually exclusive: tapping one clears the other. The old PANEL group (no-debate / repeated-each-other / wrong-person) was dropped when the card went per-bubble: those are turn-level judgements about how the hosts played off one another, and they don't fit a rating attached to a single reply.

The panel opens on demand. Every reply is born with just its in-bubble rail — 👍 👎 ⌄ — and zero rows below the bubble. A clicked thumb reads as the ink-wash pill; 👎 also unfolds the tags below; the chevron toggles. And the moment the next message lands, any open panel auto-folds — so a panel of three never stacks three open bodies down the screen:

Warren Buffett
Start lean; let the numbers earn the next dollar.
Jensen Huang
In a land-grab market, raising to move first is the whole game.

The folded default: two replies, two rails, nothing below either bubble (the first already 👍'd — the ink-wash pill). Tap ⌄ or 👎 to open one — the rest stay out of the way.

The check-in survey

Per-turn taps catch the local read; they miss the cumulative one — "this whole room is wearing me out," the thing the machine judge is blindest to. So every N turns — a cadence the admin sets in console → System → Feedback (0 = off) — the same vocabulary appears in long form, and the user can summon it any time via Take Survey in any turn's card (or the room menu). The counter is server-truth: it counts rendered panel rounds, so in a multi-user room it reaches every present person (a lurker who never types is asked too) and it survives a reload. A qualified submit resets that user's clock — answer the survey and the next one is N rounds out, not at the next fixed multiple. Pushed when we want the read; pullable when they have more to say:

Survey

~20 turns in · takes 10 seconds · helps us tune the room

How's this room going for you?
Length, overall?
When you wanted them to push back or disagree, did they?
Anything that'd make it better?

The long form. Same vocabulary as the per-turn tags, asked once about the whole session.

The arms — no FP v0 v1

The widget is arm-agnostic, so it doubles as the coarse comparison you can't get from a real-user A/B. Each room is assigned one arm at creation — blind to the user (no badge, FP-bubble off) — and its feedback accumulates under that arm:

not a classical A/BWith 3 testers there's no population to randomize and no "return" to measure — so this isn't a powered experiment. It's a direct read: tune v1's prompt → the next sessions' tags improve (hill-climbing on human labels, tighter than A/B since the user rates exactly what they got). The no FP/v0 rooms are just an occasional "are we still beating the floor?" spot-check, not the main signal.
shipped 2026-07-03 — ratings now steer the producer liveBeyond the offline tuning read above, the ratings + latest survey now feed directly back into the smart floor producer per room: a GUEST FEEDBACK block in f's brief, read as panel-level standing asks (the room's taste, never a verdict on one host). Fresh or repeated tags carry near-ask force; a rating drops out after an admin-set lifetime (default 8 turns); a survey stands until that user's next survey. Admin On/Off + lifetime live in console → Settings → Feedback — ratings are always recorded either way. Design + mechanics: the floor-producer page.

What each tap logs

The point of per-reply is attribution: a tag is stored against the exact reply that drew it, with the staging that produced that turn + the signals the producer saw — so "too long" isn't a vague gripe, it's a labelled example of a specific decision. That same record is the seed of a trained policy later (a 👍/👎 pair is exactly what DPO trains on).

// one tap → one row. nothing private leaves the room. feedback { lid: "L_8f3a" // the single reply this rates (its line id) room_id: "r_4c1e" user_id: "u_2" arm: "v1" // no_fp | v0 | v1 (blind to the user) thumb: "down" // up | down | null tags: ["too_long", "echoes_others"] // the lever-mapped chips, grouped LENGTH·TONE·CONTENT note: "Jensen restated Buffett, longer" // ── attribution: what the producer did + saw ── directive: "<the call-list f emitted>" // v1 only signals: { engagement, asks, roster, substance, … } ts: … } survey { // the long form — every N rounds (admin-set; 0=off) or on Take Survey room_id, user_id, arm, turn_n: 20, // turn_n = the round at submit; a qualified one resets this user's cadence overall: "fine", length: "about_right", pushback: "sometimes", better: "<free text>", ts: … }

Why it's built this way

build orderThis widget is the first piece of v1 — the instrument every later tuning step depends on. Ship it under v0 first (it's arm-agnostic), so the loop is collecting before the smart producer even exists. Built and live on prod (gated by the Who-sees-what feedback flag) — a rating rail inside every reply bubble (👍 👎 ⌄ bare icons, right-aligned under the text; clicked = the ink-wash pill; 👎 opens), the grouped tags LENGTH·TONE·CONTENT, an inline note [save], auto-fold on the next message, server-truth round cadence with reset-on-qualified-submit, and Take Survey. The v266 rail redesign is on main, awaiting the next box deploy.
Feedback-widget reference · drafted 2026-06-27 · revised 2026-07-03 for the rail-in-bubble redesign (v266, from CW's ui/sandbox) · the v1 measurement instrument for the floor producer. A per-reply rating — a bare-icon rail inside the bubble (👍 👎 ⌄, ink-wash pressed state, one size across platforms) + grouped lever-tags (LENGTH · TONE · CONTENT) + an inline note-save, auto-collapsing when the next message lands · the admin-set Survey (server-truth round cadence, reset on qualified submit) · three arms (no-FP / v0 / v1). Status: live on prod; the v266 rail redesign + v267 reveal-and-note-focus + v268 chip-order-and-composer-hide + v270–v271 light-mode-keycap polish on main, awaiting the next deploy.