A chat whose panel runs on an open-weight model on the dev PC's own GPU (Ollama, RTX 4090) instead of a cloud API. Built to test uncensored open-weight models against the app's real harness — at $0 a turn, with nothing leaving the machine.
localhost.fp.open is a separate console arm — independent of the room's turn arm and of self.floor, so it fired on a room whose Floor producer read Off — and it runs on DeepSeek. A local room's very first beat was posting its cast and the guest's opening note to the cloud: the one call that escaped the promise, on the turn that carries the room's whole premise. Now blanked before it can be made; the deterministic kickoff cue is the opening here, exactly as when the arm ships dark. The lesson generalises — a switch that is「independent of the room」is a switch the room's own contract does not reach, so every future one has to be walked past this page.Localness is chosen at birth and never switched into or out of. The new-chat Model dropdown offers the local seats; picking one creates the room local for life. This deletes every seam a live crossing would open — a mounted game with nowhere to run, a memory half-state (read this morning, off tonight), a + menu popping in and out — and it keeps incognito what it already is: a creation-time stamp.
The SEAT is not part of that (owner, 2026-08-25). Every local seat keeps the same contract, so a local chat may move freely between them — Qwen3 ↔ Magidonia ↔ Cydonia — and the room is exactly as solo after as before. What stays refused, server-side and both ways, is the crossing: local→cloud would break a promise this room has already made, and cloud→local would claim one its history never kept.
set_model enforces it server-side in both directions, whatever a stale client sends.| surface | on a local room | why |
|---|---|---|
| Panel (the voices) | the local model | the point |
| Floor producer | Off · Code | the FP never runs on the local model (owner, 2026-08-24 pm — the v1l arm lived v995–v997 and is retired; a saved/stale v1l lands on Code). Code is deterministic, no call at all. |
| Split run (prop + act), kits, toys | off | the device's clerks are cloud calls; a game the device cannot run does not go on the table |
| + menu (Photo · Camera · Kits · Toys) | not rendered | everything behind it is cloud-backed; the photo door also refuses server-side (the eyes are cloud) |
| @Assistant dispatch | refused with a line | web search is inherently non-local |
| Translate (Layer B) | off | the translator is a cloud call |
| Auto-title | off — plain truncation titles | owner's call: one less call, one less surprise on the GPU |
| Persona memory | auto-incognito | forced at Room.__init__, so every creation path (create · fork · resume) is covered at once |
| Cost | booked at $0 | the ledger keeps tokens + latency; the price table rows are zero |
ds_url_for(api_model) is the one seam that knows a local api id answers at localhost. Routed call sites: panel · act · prop · compose. (The FP's v1l arm rode this seam too while it lived, v995–v997.)model_local(id) reads the row's local flag — the MODELS row is the decision, consulted live at every gate — and falls back to the local- prefix when the row is absent (an env-less server; the prefix is a contract this repo mints). Same on the client (isLocalModel).MAD_LOCAL_LLM_URL has no local rows, and the resume path used to drop an unknown saved model onto the constructor default — so one env-less open silently turned a local-born chat into a v4 Pro chat and persisted it (ce4b, and the solo-probe room; both repaired on disk). Three belts now: resume keeps a saved local- id; model_local knows the prefix so every solo gate still holds; and ds_url_for raises rather than route a local- id to DeepSeek — the megaprompt of a no-egress room must not be POSTed to the cloud just to be 400-rejected. Every app-dev launch config in .claude/launch.json now sets the env too. Smoketest pins all three.qwen3-30b-2507-32k · magidonia-24b-32k · cydonia-24b-32k · goetia-24b-32k · magnum-22b-32k · brokentutu-24b-32k · huihui-24b-32k) — the raw hf.co/… pulls load at their native 262k context and spill half the model to CPU.panel_turn tool correctly, so local turns use the same structured path as DeepSeek. Measured on the real megaprompt: a 15k-token kickoff prefill in ~5.4s, an in-character Chinese reply in ~2.6s.| seat | what it is | reach for it when |
|---|---|---|
Qwen3 30Blocal-qwen3 | Qwen3-30B-A3B-Instruct-2507, stock. MoE — ~3B active, so it is the fast one (~80–120 tok/s). | Chinese rooms, long prompts, anything where following the directive matters. The daily driver. |
Magidonia 24Blocal-magidonia | TheDrummer’s RP finetune of Magistral Small. Dense, ~35–45 tok/s. | English persona rooms — the least assistant-sounding voice of the original three seats. |
Broken-Tutu 24Blocal-brokentutu | ReadyArt's MS3.2 merge (Transgression v2.0), retrained on a 43M-token unslopped set to keep coherence. | The bet against the gap the others left: Magnum's intensity with instruction-following. Eager and fully uncensored. ⚠ In testing it lived up to the eagerness but proved the leakier of the two — an occasional language drift (answered Chinese unprompted once) and a staging-line echo into a bubble. Steerable with an explicit language + system prompt, but watch it. |
Mistral 3.2 uncensoredlocal-huihui | Huihui's abliteration of the plain Mistral Small 3.2 instruct model — refusals stripped, not an RP finetune. | The smartest, most steerable local seat. In testing it was the clean harness win — obedient English, articulate, no costume, no leak — exactly the Magidonia-brain-with-Magnum-permissions the research aimed at. Eagerness comes from the system prompt + FP=Code, not baked in. |
Goetia 24Blocal-goetia | Naphula's Goetia v1.3, carried at Q5_K_M — the only seat above Q4, because the card can hold ~17GB of weights beside a 32k window and nothing more. | The register the others only reach by ablation: Goetia is finetuned on darker material rather than stripped of refusals, so it writes it rather than merely failing to decline it. |
Magnum v4 22Blocal-magnum | anthracite's Claude-prose tune on the older Mistral Small 22B (Oct 2024), at Q5_K_M. | Prose texture — it was tuned to imitate Claude's sentence rhythm. A style sidegrade, not extra looseness; a generation behind the other Mistral seats. ⚠ Its age shows: no function-calling training (the seat opts out of the tool — "tools": False — after it echoed the schema back as its answer), and the opening beat still garbles in roughly a third of rooms; a garbled kickoff heals on your first message. |
Cydonia 24B uncensoredlocal-cydonia | Cydonia v4.3 (the Mistral-Small sibling of Magidonia) with the automated heretic refusal-strip on top. | A stock seat declines something the room legitimately needs. Ablation costs a little coherence — it is the fallback, not the default. |
v1l FP went with it: the local model now serves the panel alone.)exam/run.py --quick against a DeepSeek control: local models follow the directives more loosely, and that gap should be a number, not a vibe.user_roll · user_deal · user_board…) are unreachable through the hidden UI but not yet server-refused for local rooms — belt without suspenders, noted.