← Design notes

Local rooms — the solo mode shipped 2026-08-24 · dev-only

A chat whose panel runs on an open-weight model on the dev PC's own GPU (Ollama, RTX 4090) instead of a cloud API. Built to test uncensored open-weight models against the app's real harness — at $0 a turn, with nothing leaving the machine.

The invariant: a local room makes no external provider call, ever. Not "fewer" — none. Every cloud door either refuses with a line or degrades deterministically, and the smoketest pins it: a scripted local turn's only wire is localhost.
the door that was openThe OPENING producer escaped it until 2026-08-25. fp.open is a separate console arm — independent of the room's turn arm and of self.floor, so it fired on a room whose Floor producer read Off — and it runs on DeepSeek. A local room's very first beat was posting its cast and the guest's opening note to the cloud: the one call that escaped the promise, on the turn that carries the room's whole premise. Now blanked before it can be made; the deterministic kickoff cue is the opening here, exactly as when the arm ships dark. The lesson generalises — a switch that is「independent of the room」is a switch the room's own contract does not reach, so every future one has to be walked past this page.

Born local

Localness is chosen at birth and never switched into or out of. The new-chat Model dropdown offers the local seats; picking one creates the room local for life. This deletes every seam a live crossing would open — a mounted game with nowhere to run, a memory half-state (read this morning, off tonight), a + menu popping in and out — and it keeps incognito what it already is: a creation-time stamp.

The SEAT is not part of that (owner, 2026-08-25). Every local seat keeps the same contract, so a local chat may move freely between them — Qwen3 ↔ Magidonia ↔ Cydonia — and the room is exactly as solo after as before. What stays refused, server-side and both ways, is the crossing: local→cloud would break a promise this room has already made, and cloud→local would claim one its history never kept.

What a local room runs, and what it refuses

surfaceon a local roomwhy
Panel (the voices)the local modelthe point
Floor producerOff · Codethe FP never runs on the local model (owner, 2026-08-24 pm — the v1l arm lived v995–v997 and is retired; a saved/stale v1l lands on Code). Code is deterministic, no call at all.
Split run (prop + act), kits, toysoffthe device's clerks are cloud calls; a game the device cannot run does not go on the table
+ menu (Photo · Camera · Kits · Toys)not renderedeverything behind it is cloud-backed; the photo door also refuses server-side (the eyes are cloud)
@Assistant dispatchrefused with a lineweb search is inherently non-local
Translate (Layer B)offthe translator is a cloud call
Auto-titleoff — plain truncation titlesowner's call: one less call, one less surprise on the GPU
Persona memoryauto-incognitoforced at Room.__init__, so every creation path (create · fork · resume) is covered at once
Costbooked at $0the ledger keeps tokens + latency; the price table rows are zero

The plumbing

MAD_LOCAL_LLM_URL=http://localhost:11434/v1/chat/completions ← set ONLY by the dev launchers unset (the box) → the rows do not exist: no picker entries, no routing, no mode

The seats

seatwhat it isreach for it when
Qwen3 30B
local-qwen3
Qwen3-30B-A3B-Instruct-2507, stock. MoE — ~3B active, so it is the fast one (~80–120 tok/s).Chinese rooms, long prompts, anything where following the directive matters. The daily driver.
Magidonia 24B
local-magidonia
TheDrummer’s RP finetune of Magistral Small. Dense, ~35–45 tok/s.English persona rooms — the least assistant-sounding voice of the original three seats.
Broken-Tutu 24B
local-brokentutu
ReadyArt's MS3.2 merge (Transgression v2.0), retrained on a 43M-token unslopped set to keep coherence.The bet against the gap the others left: Magnum's intensity with instruction-following. Eager and fully uncensored. ⚠ In testing it lived up to the eagerness but proved the leakier of the two — an occasional language drift (answered Chinese unprompted once) and a staging-line echo into a bubble. Steerable with an explicit language + system prompt, but watch it.
Mistral 3.2 uncensored
local-huihui
Huihui's abliteration of the plain Mistral Small 3.2 instruct model — refusals stripped, not an RP finetune.The smartest, most steerable local seat. In testing it was the clean harness win — obedient English, articulate, no costume, no leak — exactly the Magidonia-brain-with-Magnum-permissions the research aimed at. Eagerness comes from the system prompt + FP=Code, not baked in.
Goetia 24B
local-goetia
Naphula's Goetia v1.3, carried at Q5_K_M — the only seat above Q4, because the card can hold ~17GB of weights beside a 32k window and nothing more.The register the others only reach by ablation: Goetia is finetuned on darker material rather than stripped of refusals, so it writes it rather than merely failing to decline it.
Magnum v4 22B
local-magnum
anthracite's Claude-prose tune on the older Mistral Small 22B (Oct 2024), at Q5_K_M.Prose texture — it was tuned to imitate Claude's sentence rhythm. A style sidegrade, not extra looseness; a generation behind the other Mistral seats. ⚠ Its age shows: no function-calling training (the seat opts out of the tool — "tools": False — after it echoed the schema back as its answer), and the opening beat still garbles in roughly a third of rooms; a garbled kickoff heals on your first message.
Cydonia 24B uncensored
local-cydonia
Cydonia v4.3 (the Mistral-Small sibling of Magidonia) with the automated heretic refusal-strip on top.A stock seat declines something the room legitimately needs. Ablation costs a little coherence — it is the fallback, not the default.
tradeOne GPU, one model at a time. Alternating Qwen3 and Magidonia rooms pays a 10–30s model swap per alternation — A/B them in separate stretches. (The slot-eviction worry that came with the v1l FP went with it: the local model now serves the panel alone.)

Owed

-ish · local rooms · born local, solo, $0 · see also Dev loop · Floor producer · Chat start