Files
6krrt/plans/magic-brainstorming-review.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

8.0 KiB

Review: opencode's "magic" brainstorm (closed-loop optimizer, local drafting, context-aware pinch)

Status: done -- review of shipped work

What it was reviewing: three feature proposals opencode (via the llm-router auto model) generated when asked to brainstorm "high-magic" next steps for this router — a quality_monitor.py closed-loop optimizer, a local-draft/cloud-refine hybrid ("Local-Cloud Hybrid Synergy"), and a model-aware upgrade to the disabled context_prune.py ("pinch") module. Not a diff — nothing here had been implemented — so this checks the pitch's claims against the actual current code (feedback.py, context_prune.py, iteration.py, dispatcher.py) rather than against the pitch's own framing.

Verdict: none as pitched. One is nearly free, one has a broken cost

mechanism, one is real but narrower than claimed.

1. Closed-loop quality monitor — mostly already built; the risky half fights the project's own design

Claim: feedback.py is "on-demand," so a new quality_monitor.py background service is needed to detect drift via sliding-window failure rates and auto-dampen a model's score or flag it for re-eval.

Finding. feedback.py:79-114 already does the sliding-window aggregation this proposes: unapplied_failures groups verifications rows by (model, category), and apply_failures folds them into proficiency via add_self_eval. The only missing piece is a schedule — turning it into a background service is a llm-router-feedback.timer unit identical in shape to the two already shipped (deploy/llm-router-poller.timer, deploy/llm-router-seed.timer), not new logic.

The "detect drift and auto-dampen or flag for re-eval" half is the part actually being proposed, and it cuts against this project's own history: three separate harness bugs (CLAUDE.md's "Harness bugs this shook out" — the shared token budget, kimi's leading-space indentation error, unparseable judge output) each looked exactly like a quality signal and were only caught by a human reading per-task detail, never by a threshold. add_self_eval's running mean is deliberately smoothing for the same reason. An automated dampening layer on top would fight that design, not extend it.

Recommendation. Ship the timer — near-zero cost, real value. Treat "auto-dampen" as "flag for human review," not "silently adjust routing."

2. Local drafting ("Local-Cloud Hybrid Synergy") — the savings mechanism doesn't check out, and it reopens a gap this project already got burned by once

Claim: a local model drafts a skeleton/plan, a cloud model "refines" it, cutting billed completion_tokens — the expensive side of the bill.

Finding. The draft becomes part of the cloud call's prompt, not its completion — "refine this draft" still requires the cloud model to generate a full final answer as completion tokens, unless doing constrained diff-editing, which was not proposed and is not a small addition. Given CLAUDE.md's own measurement that a completion token costs 201x a prompt token ("Verification: what local compute is actually good for"), a scheme that's really "local draft + full prompt + full cloud completion" can easily cost more, not less. The pitch's central cost claim was never checked against the codebase's own numbers.

It also reopens open item #5 ("Local energy is not on the ledger") at a much larger scale — full local generation instead of just classification — before the metering that would let anyone tell if it's actually cheaper exists. This project already spent real effort discovering "local is free" was wrong once (the hosted classifier beat the local one on speed, accuracy, and attributed energy); this proposal reintroduces the same unverified assumption in a bigger, less reversible form. It also serializes two model calls on the interactive path, against iteration.py's explicit stance that every retry/extra hop is a latency cost.

Recommendation. Do not build. If local drafting comes back, item #5 (real local energy metering) needs to exist first, so the cost claim can be checked instead of assumed.

3. Context-aware pinch — the most grounded idea, but narrower than pitched

Claim: make context_prune.py ("pinch," shipped enabled: false) model-aware — size the prune to the selected model's context window, so a request can trade context precision for routing to a cheaper/smaller model.

Finding. context_prune.py and PinchConfig are fully built and wired into both dispatch call sites (dispatcher.py:2107, :2334, confirmed by direct read). Making it model-aware is real but non-trivial: today pinch runs before model selection with one static budget_tokens, and its output already feeds the tier/cost decision (dispatcher.py:2099-2123) — so "pinch to fit the chosen model" needs either a provisional-selection → pinch-to-fit → reselect loop, or moving pinch to after a first-pass candidate pick. It reuses tested code rather than inventing a new subsystem, which is the strongest thing in its favor.

One overclaim in the pitch: pinch only ever trims tool-result messages — user/assistant/system content is always kept verbatim by explicit design (context_prune.py:9). So "massive cost drop by shifting to a cheaper small-window model" only applies to long agent sessions with tool-call history. A single huge pasted document in a user message — a common way to block a cheap model — gets none of this benefit today, and the pitch didn't note that limit.

Recommendation. If one of the three ships, this is it — but scope it explicitly to agent/tool-heavy sessions, not "any long conversation," since that's what the underlying mechanism can actually deliver.

Addendum: what this session actually shipped

Reviewing #2 raised a GPU-capacity question that turned out to be real and already live, not hypothetical. Measured on this box's Quadro RTX 6000 (24GB): mistral-nemo:12b (classifier) and qwen3-vl:4b (local_vision fallback), both loaded at Ollama's untagged default context (32768), together used 22.1GB — 1.9GB free, one browser tab of GPU use away from contention. Neither model needed anywhere near that: max_input_chars + system_prompt + max_output_tokens bound the classifier well under 4k tokens, and 32768 was never a value anyone chose — it's just what the base model's library Modelfile defaults to.

Fixed by baking a right-sized context into a Modelfile-tagged variant of each model (mistral-nemo-router:12b @ 8192, qwen3-vl-router:4b @ 16384) and pointing classifier.model / verification.model / local_vision.model at the tags instead of the base ones. Verified live: 8.6GB and 6.9GB resident respectively, 15.9GB combined worst case (both hot at once) versus 22.1GB before — free VRAM went from 1.9GB to 8.1GB in that case, and from 10.8GB to 14.6GB in the common case (classifier alone). Confirmed empirically along the way, not assumed: Ollama's OpenAI-compatible endpoint (0.22.0) silently ignores a per-request num_ctx or keep_alive override under every field shape tried (options.num_ctx, top-level num_ctx, context_length, top-level keep_alive) — a 200 comes back and nothing changes. Only the native /api/chat endpoint honors either, which is why the fix has to live in the model tag rather than in a request parameter, and why a shorter keep_alive for the vision fallback specifically was left as a follow-up (it would need that path rewritten onto the native API, including its image format).

This doesn't change the verdict on #2: headroom existing now is not the same as local generation being cheaper than a cloud completion, which is still unmeasured. It does mean the box has real spare VRAM again, which is one less reason to reach for anything drastic to free some.

Overall

None of the three should be built as pitched. #1's real, low-cost half (the timer) is worth doing now. #2's cost mechanism doesn't hold up against this project's own numbers and depends on an open item (#5) that isn't done yet — hold it. #3 is the one worth real design time, once scoped down to what pinch actually touches.