plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
8.0 KiB
Review: opencode's "magic" brainstorm (closed-loop optimizer, local drafting, context-aware pinch)
Status: done -- review of shipped work
What it was reviewing: three feature proposals opencode (via the
llm-router auto model) generated when asked to brainstorm "high-magic"
next steps for this router — a quality_monitor.py closed-loop optimizer, a
local-draft/cloud-refine hybrid ("Local-Cloud Hybrid Synergy"), and a
model-aware upgrade to the disabled context_prune.py ("pinch") module.
Not a diff — nothing here had been implemented — so this checks the pitch's
claims against the actual current code (feedback.py, context_prune.py,
iteration.py, dispatcher.py) rather than against the pitch's own framing.
Verdict: none as pitched. One is nearly free, one has a broken cost
mechanism, one is real but narrower than claimed.
1. Closed-loop quality monitor — mostly already built; the risky half fights the project's own design
Claim: feedback.py is "on-demand," so a new quality_monitor.py
background service is needed to detect drift via sliding-window failure
rates and auto-dampen a model's score or flag it for re-eval.
Finding. feedback.py:79-114 already does the sliding-window
aggregation this proposes: unapplied_failures groups verifications rows
by (model, category), and apply_failures folds them into proficiency
via add_self_eval. The only missing piece is a schedule — turning it into
a background service is a llm-router-feedback.timer unit identical in
shape to the two already shipped (deploy/llm-router-poller.timer,
deploy/llm-router-seed.timer), not new logic.
The "detect drift and auto-dampen or flag for re-eval" half is the part
actually being proposed, and it cuts against this project's own history:
three separate harness bugs (CLAUDE.md's "Harness bugs this shook out" — the
shared token budget, kimi's leading-space indentation error, unparseable
judge output) each looked exactly like a quality signal and were only
caught by a human reading per-task detail, never by a threshold.
add_self_eval's running mean is deliberately smoothing for the same
reason. An automated dampening layer on top would fight that design, not
extend it.
Recommendation. Ship the timer — near-zero cost, real value. Treat "auto-dampen" as "flag for human review," not "silently adjust routing."
2. Local drafting ("Local-Cloud Hybrid Synergy") — the savings mechanism doesn't check out, and it reopens a gap this project already got burned by once
Claim: a local model drafts a skeleton/plan, a cloud model "refines" it,
cutting billed completion_tokens — the expensive side of the bill.
Finding. The draft becomes part of the cloud call's prompt, not its completion — "refine this draft" still requires the cloud model to generate a full final answer as completion tokens, unless doing constrained diff-editing, which was not proposed and is not a small addition. Given CLAUDE.md's own measurement that a completion token costs 201x a prompt token ("Verification: what local compute is actually good for"), a scheme that's really "local draft + full prompt + full cloud completion" can easily cost more, not less. The pitch's central cost claim was never checked against the codebase's own numbers.
It also reopens open item #5 ("Local energy is not on the ledger") at a
much larger scale — full local generation instead of just classification —
before the metering that would let anyone tell if it's actually cheaper
exists. This project already spent real effort discovering "local is free"
was wrong once (the hosted classifier beat the local one on speed,
accuracy, and attributed energy); this proposal reintroduces the same
unverified assumption in a bigger, less reversible form. It also serializes
two model calls on the interactive path, against iteration.py's explicit
stance that every retry/extra hop is a latency cost.
Recommendation. Do not build. If local drafting comes back, item #5 (real local energy metering) needs to exist first, so the cost claim can be checked instead of assumed.
3. Context-aware pinch — the most grounded idea, but narrower than pitched
Claim: make context_prune.py ("pinch," shipped enabled: false)
model-aware — size the prune to the selected model's context window, so a
request can trade context precision for routing to a cheaper/smaller model.
Finding. context_prune.py and PinchConfig are fully built and wired
into both dispatch call sites (dispatcher.py:2107, :2334, confirmed by
direct read). Making it model-aware is real but non-trivial: today pinch
runs before model selection with one static budget_tokens, and its
output already feeds the tier/cost decision (dispatcher.py:2099-2123) —
so "pinch to fit the chosen model" needs either a provisional-selection →
pinch-to-fit → reselect loop, or moving pinch to after a first-pass
candidate pick. It reuses tested code rather than inventing a new
subsystem, which is the strongest thing in its favor.
One overclaim in the pitch: pinch only ever trims tool-result messages —
user/assistant/system content is always kept verbatim by explicit design
(context_prune.py:9). So "massive cost drop by shifting to a cheaper
small-window model" only applies to long agent sessions with tool-call
history. A single huge pasted document in a user message — a common way to
block a cheap model — gets none of this benefit today, and the pitch didn't
note that limit.
Recommendation. If one of the three ships, this is it — but scope it explicitly to agent/tool-heavy sessions, not "any long conversation," since that's what the underlying mechanism can actually deliver.
Addendum: what this session actually shipped
Reviewing #2 raised a GPU-capacity question that turned out to be real and
already live, not hypothetical. Measured on this box's Quadro RTX 6000
(24GB): mistral-nemo:12b (classifier) and qwen3-vl:4b (local_vision
fallback), both loaded at Ollama's untagged default context (32768),
together used 22.1GB — 1.9GB free, one browser tab of GPU use away from
contention. Neither model needed anywhere near that: max_input_chars +
system_prompt + max_output_tokens bound the classifier well under 4k
tokens, and 32768 was never a value anyone chose — it's just what the base
model's library Modelfile defaults to.
Fixed by baking a right-sized context into a Modelfile-tagged variant of
each model (mistral-nemo-router:12b @ 8192, qwen3-vl-router:4b @ 16384)
and pointing classifier.model / verification.model / local_vision.model
at the tags instead of the base ones. Verified live: 8.6GB and 6.9GB
resident respectively, 15.9GB combined worst case (both hot at once) versus
22.1GB before — free VRAM went from 1.9GB to 8.1GB in that case, and from
10.8GB to 14.6GB in the common case (classifier alone). Confirmed
empirically along the way, not assumed: Ollama's OpenAI-compatible endpoint
(0.22.0) silently ignores a per-request num_ctx or keep_alive override
under every field shape tried (options.num_ctx, top-level num_ctx,
context_length, top-level keep_alive) — a 200 comes back and nothing
changes. Only the native /api/chat endpoint honors either, which is why
the fix has to live in the model tag rather than in a request parameter,
and why a shorter keep_alive for the vision fallback specifically was
left as a follow-up (it would need that path rewritten onto the native API,
including its image format).
This doesn't change the verdict on #2: headroom existing now is not the same as local generation being cheaper than a cloud completion, which is still unmeasured. It does mean the box has real spare VRAM again, which is one less reason to reach for anything drastic to free some.
Overall
None of the three should be built as pitched. #1's real, low-cost half (the timer) is worth doing now. #2's cost mechanism doesn't hold up against this project's own numbers and depends on an open item (#5) that isn't done yet — hold it. #3 is the one worth real design time, once scoped down to what pinch actually touches.