Files
6krrt/plans/magic-brainstorming-review.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

142 lines
8.0 KiB
Markdown

# Review: opencode's "magic" brainstorm (closed-loop optimizer, local drafting, context-aware pinch)
Status: done -- review of shipped work
**What it was reviewing:** three feature proposals opencode (via the
llm-router `auto` model) generated when asked to brainstorm "high-magic"
next steps for this router — a `quality_monitor.py` closed-loop optimizer, a
local-draft/cloud-refine hybrid ("Local-Cloud Hybrid Synergy"), and a
model-aware upgrade to the disabled `context_prune.py` ("pinch") module.
Not a diff — nothing here had been implemented — so this checks the pitch's
claims against the actual current code (`feedback.py`, `context_prune.py`,
`iteration.py`, `dispatcher.py`) rather than against the pitch's own framing.
## Verdict: none as pitched. One is nearly free, one has a broken cost
## mechanism, one is real but narrower than claimed.
### 1. Closed-loop quality monitor — mostly already built; the risky half fights the project's own design
**Claim:** `feedback.py` is "on-demand," so a new `quality_monitor.py`
background service is needed to detect drift via sliding-window failure
rates and auto-dampen a model's score or flag it for re-eval.
**Finding.** `feedback.py:79-114` already does the sliding-window
aggregation this proposes: `unapplied_failures` groups `verifications` rows
by `(model, category)`, and `apply_failures` folds them into `proficiency`
via `add_self_eval`. The only missing piece is a schedule — turning it into
a background service is a `llm-router-feedback.timer` unit identical in
shape to the two already shipped (`deploy/llm-router-poller.timer`,
`deploy/llm-router-seed.timer`), not new logic.
The "detect drift and auto-dampen or flag for re-eval" half is the part
actually being proposed, and it cuts against this project's own history:
three separate harness bugs (CLAUDE.md's "Harness bugs this shook out" — the
shared token budget, kimi's leading-space indentation error, unparseable
judge output) each looked exactly like a quality signal and were only
caught by a human reading per-task detail, never by a threshold.
`add_self_eval`'s running mean is deliberately smoothing for the same
reason. An automated dampening layer on top would fight that design, not
extend it.
**Recommendation.** Ship the timer — near-zero cost, real value. Treat
"auto-dampen" as "flag for human review," not "silently adjust routing."
### 2. Local drafting ("Local-Cloud Hybrid Synergy") — the savings mechanism doesn't check out, and it reopens a gap this project already got burned by once
**Claim:** a local model drafts a skeleton/plan, a cloud model "refines" it,
cutting billed `completion_tokens` — the expensive side of the bill.
**Finding.** The draft becomes part of the cloud call's *prompt*, not its
completion — "refine this draft" still requires the cloud model to generate
a full final answer as completion tokens, unless doing constrained
diff-editing, which was not proposed and is not a small addition. Given
CLAUDE.md's own measurement that a completion token costs **201x** a prompt
token ("Verification: what local compute is actually good for"), a scheme
that's really "local draft + full prompt + full cloud completion" can
easily cost *more*, not less. The pitch's central cost claim was never
checked against the codebase's own numbers.
It also reopens open item #5 ("Local energy is not on the ledger") at a
much larger scale — full local generation instead of just classification —
before the metering that would let anyone tell if it's actually cheaper
exists. This project already spent real effort discovering "local is free"
was wrong once (the hosted classifier beat the local one on speed,
accuracy, *and* attributed energy); this proposal reintroduces the same
unverified assumption in a bigger, less reversible form. It also serializes
two model calls on the interactive path, against `iteration.py`'s explicit
stance that every retry/extra hop is a latency cost.
**Recommendation.** Do not build. If local drafting comes back, item #5
(real local energy metering) needs to exist first, so the cost claim can be
checked instead of assumed.
### 3. Context-aware pinch — the most grounded idea, but narrower than pitched
**Claim:** make `context_prune.py` ("pinch," shipped `enabled: false`)
model-aware — size the prune to the *selected* model's context window, so a
request can trade context precision for routing to a cheaper/smaller model.
**Finding.** `context_prune.py` and `PinchConfig` are fully built and wired
into both dispatch call sites (`dispatcher.py:2107`, `:2334`, confirmed by
direct read). Making it model-aware is real but non-trivial: today pinch
runs *before* model selection with one static `budget_tokens`, and its
output already feeds the tier/cost decision (`dispatcher.py:2099-2123`) —
so "pinch to fit the chosen model" needs either a provisional-selection →
pinch-to-fit → reselect loop, or moving pinch to after a first-pass
candidate pick. It reuses tested code rather than inventing a new
subsystem, which is the strongest thing in its favor.
One overclaim in the pitch: pinch only ever trims *tool-result* messages —
user/assistant/system content is always kept verbatim by explicit design
(`context_prune.py:9`). So "massive cost drop by shifting to a cheaper
small-window model" only applies to long agent sessions with tool-call
history. A single huge pasted document in a user message — a common way to
block a cheap model — gets none of this benefit today, and the pitch didn't
note that limit.
**Recommendation.** If one of the three ships, this is it — but scope it
explicitly to agent/tool-heavy sessions, not "any long conversation," since
that's what the underlying mechanism can actually deliver.
## Addendum: what this session actually shipped
Reviewing #2 raised a GPU-capacity question that turned out to be real and
already live, not hypothetical. Measured on this box's Quadro RTX 6000
(24GB): `mistral-nemo:12b` (classifier) and `qwen3-vl:4b` (`local_vision`
fallback), both loaded at Ollama's untagged default context (32768),
together used 22.1GB — 1.9GB free, one browser tab of GPU use away from
contention. Neither model needed anywhere near that: `max_input_chars` +
`system_prompt` + `max_output_tokens` bound the classifier well under 4k
tokens, and 32768 was never a value anyone chose — it's just what the base
model's library Modelfile defaults to.
Fixed by baking a right-sized context into a Modelfile-tagged variant of
each model (`mistral-nemo-router:12b` @ 8192, `qwen3-vl-router:4b` @ 16384)
and pointing `classifier.model` / `verification.model` / `local_vision.model`
at the tags instead of the base ones. Verified live: 8.6GB and 6.9GB
resident respectively, 15.9GB combined worst case (both hot at once) versus
22.1GB before — free VRAM went from 1.9GB to 8.1GB in that case, and from
10.8GB to 14.6GB in the common case (classifier alone). Confirmed
empirically along the way, not assumed: Ollama's OpenAI-compatible endpoint
(0.22.0) silently ignores a per-request `num_ctx` or `keep_alive` override
under every field shape tried (`options.num_ctx`, top-level `num_ctx`,
`context_length`, top-level `keep_alive`) — a 200 comes back and nothing
changes. Only the native `/api/chat` endpoint honors either, which is why
the fix has to live in the model tag rather than in a request parameter,
and why a shorter `keep_alive` for the vision fallback specifically was
left as a follow-up (it would need that path rewritten onto the native API,
including its image format).
This doesn't change the verdict on #2: headroom existing now is not the
same as local generation being cheaper than a cloud completion, which is
still unmeasured. It does mean the box has real spare VRAM again, which is
one less reason to reach for anything drastic to free some.
## Overall
None of the three should be built as pitched. #1's real, low-cost half (the
timer) is worth doing now. #2's cost mechanism doesn't hold up against this
project's own numbers and depends on an open item (#5) that isn't done yet
— hold it. #3 is the one worth real design time, once scoped down to what
pinch actually touches.