# Review: opencode's "magic" brainstorm (closed-loop optimizer, local drafting, context-aware pinch) Status: done -- review of shipped work **What it was reviewing:** three feature proposals opencode (via the llm-router `auto` model) generated when asked to brainstorm "high-magic" next steps for this router — a `quality_monitor.py` closed-loop optimizer, a local-draft/cloud-refine hybrid ("Local-Cloud Hybrid Synergy"), and a model-aware upgrade to the disabled `context_prune.py` ("pinch") module. Not a diff — nothing here had been implemented — so this checks the pitch's claims against the actual current code (`feedback.py`, `context_prune.py`, `iteration.py`, `dispatcher.py`) rather than against the pitch's own framing. ## Verdict: none as pitched. One is nearly free, one has a broken cost ## mechanism, one is real but narrower than claimed. ### 1. Closed-loop quality monitor — mostly already built; the risky half fights the project's own design **Claim:** `feedback.py` is "on-demand," so a new `quality_monitor.py` background service is needed to detect drift via sliding-window failure rates and auto-dampen a model's score or flag it for re-eval. **Finding.** `feedback.py:79-114` already does the sliding-window aggregation this proposes: `unapplied_failures` groups `verifications` rows by `(model, category)`, and `apply_failures` folds them into `proficiency` via `add_self_eval`. The only missing piece is a schedule — turning it into a background service is a `llm-router-feedback.timer` unit identical in shape to the two already shipped (`deploy/llm-router-poller.timer`, `deploy/llm-router-seed.timer`), not new logic. The "detect drift and auto-dampen or flag for re-eval" half is the part actually being proposed, and it cuts against this project's own history: three separate harness bugs (CLAUDE.md's "Harness bugs this shook out" — the shared token budget, kimi's leading-space indentation error, unparseable judge output) each looked exactly like a quality signal and were only caught by a human reading per-task detail, never by a threshold. `add_self_eval`'s running mean is deliberately smoothing for the same reason. An automated dampening layer on top would fight that design, not extend it. **Recommendation.** Ship the timer — near-zero cost, real value. Treat "auto-dampen" as "flag for human review," not "silently adjust routing." ### 2. Local drafting ("Local-Cloud Hybrid Synergy") — the savings mechanism doesn't check out, and it reopens a gap this project already got burned by once **Claim:** a local model drafts a skeleton/plan, a cloud model "refines" it, cutting billed `completion_tokens` — the expensive side of the bill. **Finding.** The draft becomes part of the cloud call's *prompt*, not its completion — "refine this draft" still requires the cloud model to generate a full final answer as completion tokens, unless doing constrained diff-editing, which was not proposed and is not a small addition. Given CLAUDE.md's own measurement that a completion token costs **201x** a prompt token ("Verification: what local compute is actually good for"), a scheme that's really "local draft + full prompt + full cloud completion" can easily cost *more*, not less. The pitch's central cost claim was never checked against the codebase's own numbers. It also reopens open item #5 ("Local energy is not on the ledger") at a much larger scale — full local generation instead of just classification — before the metering that would let anyone tell if it's actually cheaper exists. This project already spent real effort discovering "local is free" was wrong once (the hosted classifier beat the local one on speed, accuracy, *and* attributed energy); this proposal reintroduces the same unverified assumption in a bigger, less reversible form. It also serializes two model calls on the interactive path, against `iteration.py`'s explicit stance that every retry/extra hop is a latency cost. **Recommendation.** Do not build. If local drafting comes back, item #5 (real local energy metering) needs to exist first, so the cost claim can be checked instead of assumed. ### 3. Context-aware pinch — the most grounded idea, but narrower than pitched **Claim:** make `context_prune.py` ("pinch," shipped `enabled: false`) model-aware — size the prune to the *selected* model's context window, so a request can trade context precision for routing to a cheaper/smaller model. **Finding.** `context_prune.py` and `PinchConfig` are fully built and wired into both dispatch call sites (`dispatcher.py:2107`, `:2334`, confirmed by direct read). Making it model-aware is real but non-trivial: today pinch runs *before* model selection with one static `budget_tokens`, and its output already feeds the tier/cost decision (`dispatcher.py:2099-2123`) — so "pinch to fit the chosen model" needs either a provisional-selection → pinch-to-fit → reselect loop, or moving pinch to after a first-pass candidate pick. It reuses tested code rather than inventing a new subsystem, which is the strongest thing in its favor. One overclaim in the pitch: pinch only ever trims *tool-result* messages — user/assistant/system content is always kept verbatim by explicit design (`context_prune.py:9`). So "massive cost drop by shifting to a cheaper small-window model" only applies to long agent sessions with tool-call history. A single huge pasted document in a user message — a common way to block a cheap model — gets none of this benefit today, and the pitch didn't note that limit. **Recommendation.** If one of the three ships, this is it — but scope it explicitly to agent/tool-heavy sessions, not "any long conversation," since that's what the underlying mechanism can actually deliver. ## Addendum: what this session actually shipped Reviewing #2 raised a GPU-capacity question that turned out to be real and already live, not hypothetical. Measured on this box's Quadro RTX 6000 (24GB): `mistral-nemo:12b` (classifier) and `qwen3-vl:4b` (`local_vision` fallback), both loaded at Ollama's untagged default context (32768), together used 22.1GB — 1.9GB free, one browser tab of GPU use away from contention. Neither model needed anywhere near that: `max_input_chars` + `system_prompt` + `max_output_tokens` bound the classifier well under 4k tokens, and 32768 was never a value anyone chose — it's just what the base model's library Modelfile defaults to. Fixed by baking a right-sized context into a Modelfile-tagged variant of each model (`mistral-nemo-router:12b` @ 8192, `qwen3-vl-router:4b` @ 16384) and pointing `classifier.model` / `verification.model` / `local_vision.model` at the tags instead of the base ones. Verified live: 8.6GB and 6.9GB resident respectively, 15.9GB combined worst case (both hot at once) versus 22.1GB before — free VRAM went from 1.9GB to 8.1GB in that case, and from 10.8GB to 14.6GB in the common case (classifier alone). Confirmed empirically along the way, not assumed: Ollama's OpenAI-compatible endpoint (0.22.0) silently ignores a per-request `num_ctx` or `keep_alive` override under every field shape tried (`options.num_ctx`, top-level `num_ctx`, `context_length`, top-level `keep_alive`) — a 200 comes back and nothing changes. Only the native `/api/chat` endpoint honors either, which is why the fix has to live in the model tag rather than in a request parameter, and why a shorter `keep_alive` for the vision fallback specifically was left as a follow-up (it would need that path rewritten onto the native API, including its image format). This doesn't change the verdict on #2: headroom existing now is not the same as local generation being cheaper than a cloud completion, which is still unmeasured. It does mean the box has real spare VRAM again, which is one less reason to reach for anything drastic to free some. ## Overall None of the three should be built as pitched. #1's real, low-cost half (the timer) is worth doing now. #2's cost mechanism doesn't hold up against this project's own numbers and depends on an open item (#5) that isn't done yet — hold it. #3 is the one worth real design time, once scoped down to what pinch actually touches.