plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
142 lines
8.0 KiB
Markdown
142 lines
8.0 KiB
Markdown
# Review: opencode's "magic" brainstorm (closed-loop optimizer, local drafting, context-aware pinch)
|
|
|
|
Status: done -- review of shipped work
|
|
|
|
**What it was reviewing:** three feature proposals opencode (via the
|
|
llm-router `auto` model) generated when asked to brainstorm "high-magic"
|
|
next steps for this router — a `quality_monitor.py` closed-loop optimizer, a
|
|
local-draft/cloud-refine hybrid ("Local-Cloud Hybrid Synergy"), and a
|
|
model-aware upgrade to the disabled `context_prune.py` ("pinch") module.
|
|
Not a diff — nothing here had been implemented — so this checks the pitch's
|
|
claims against the actual current code (`feedback.py`, `context_prune.py`,
|
|
`iteration.py`, `dispatcher.py`) rather than against the pitch's own framing.
|
|
|
|
## Verdict: none as pitched. One is nearly free, one has a broken cost
|
|
## mechanism, one is real but narrower than claimed.
|
|
|
|
### 1. Closed-loop quality monitor — mostly already built; the risky half fights the project's own design
|
|
|
|
**Claim:** `feedback.py` is "on-demand," so a new `quality_monitor.py`
|
|
background service is needed to detect drift via sliding-window failure
|
|
rates and auto-dampen a model's score or flag it for re-eval.
|
|
|
|
**Finding.** `feedback.py:79-114` already does the sliding-window
|
|
aggregation this proposes: `unapplied_failures` groups `verifications` rows
|
|
by `(model, category)`, and `apply_failures` folds them into `proficiency`
|
|
via `add_self_eval`. The only missing piece is a schedule — turning it into
|
|
a background service is a `llm-router-feedback.timer` unit identical in
|
|
shape to the two already shipped (`deploy/llm-router-poller.timer`,
|
|
`deploy/llm-router-seed.timer`), not new logic.
|
|
|
|
The "detect drift and auto-dampen or flag for re-eval" half is the part
|
|
actually being proposed, and it cuts against this project's own history:
|
|
three separate harness bugs (CLAUDE.md's "Harness bugs this shook out" — the
|
|
shared token budget, kimi's leading-space indentation error, unparseable
|
|
judge output) each looked exactly like a quality signal and were only
|
|
caught by a human reading per-task detail, never by a threshold.
|
|
`add_self_eval`'s running mean is deliberately smoothing for the same
|
|
reason. An automated dampening layer on top would fight that design, not
|
|
extend it.
|
|
|
|
**Recommendation.** Ship the timer — near-zero cost, real value. Treat
|
|
"auto-dampen" as "flag for human review," not "silently adjust routing."
|
|
|
|
### 2. Local drafting ("Local-Cloud Hybrid Synergy") — the savings mechanism doesn't check out, and it reopens a gap this project already got burned by once
|
|
|
|
**Claim:** a local model drafts a skeleton/plan, a cloud model "refines" it,
|
|
cutting billed `completion_tokens` — the expensive side of the bill.
|
|
|
|
**Finding.** The draft becomes part of the cloud call's *prompt*, not its
|
|
completion — "refine this draft" still requires the cloud model to generate
|
|
a full final answer as completion tokens, unless doing constrained
|
|
diff-editing, which was not proposed and is not a small addition. Given
|
|
CLAUDE.md's own measurement that a completion token costs **201x** a prompt
|
|
token ("Verification: what local compute is actually good for"), a scheme
|
|
that's really "local draft + full prompt + full cloud completion" can
|
|
easily cost *more*, not less. The pitch's central cost claim was never
|
|
checked against the codebase's own numbers.
|
|
|
|
It also reopens open item #5 ("Local energy is not on the ledger") at a
|
|
much larger scale — full local generation instead of just classification —
|
|
before the metering that would let anyone tell if it's actually cheaper
|
|
exists. This project already spent real effort discovering "local is free"
|
|
was wrong once (the hosted classifier beat the local one on speed,
|
|
accuracy, *and* attributed energy); this proposal reintroduces the same
|
|
unverified assumption in a bigger, less reversible form. It also serializes
|
|
two model calls on the interactive path, against `iteration.py`'s explicit
|
|
stance that every retry/extra hop is a latency cost.
|
|
|
|
**Recommendation.** Do not build. If local drafting comes back, item #5
|
|
(real local energy metering) needs to exist first, so the cost claim can be
|
|
checked instead of assumed.
|
|
|
|
### 3. Context-aware pinch — the most grounded idea, but narrower than pitched
|
|
|
|
**Claim:** make `context_prune.py` ("pinch," shipped `enabled: false`)
|
|
model-aware — size the prune to the *selected* model's context window, so a
|
|
request can trade context precision for routing to a cheaper/smaller model.
|
|
|
|
**Finding.** `context_prune.py` and `PinchConfig` are fully built and wired
|
|
into both dispatch call sites (`dispatcher.py:2107`, `:2334`, confirmed by
|
|
direct read). Making it model-aware is real but non-trivial: today pinch
|
|
runs *before* model selection with one static `budget_tokens`, and its
|
|
output already feeds the tier/cost decision (`dispatcher.py:2099-2123`) —
|
|
so "pinch to fit the chosen model" needs either a provisional-selection →
|
|
pinch-to-fit → reselect loop, or moving pinch to after a first-pass
|
|
candidate pick. It reuses tested code rather than inventing a new
|
|
subsystem, which is the strongest thing in its favor.
|
|
|
|
One overclaim in the pitch: pinch only ever trims *tool-result* messages —
|
|
user/assistant/system content is always kept verbatim by explicit design
|
|
(`context_prune.py:9`). So "massive cost drop by shifting to a cheaper
|
|
small-window model" only applies to long agent sessions with tool-call
|
|
history. A single huge pasted document in a user message — a common way to
|
|
block a cheap model — gets none of this benefit today, and the pitch didn't
|
|
note that limit.
|
|
|
|
**Recommendation.** If one of the three ships, this is it — but scope it
|
|
explicitly to agent/tool-heavy sessions, not "any long conversation," since
|
|
that's what the underlying mechanism can actually deliver.
|
|
|
|
## Addendum: what this session actually shipped
|
|
|
|
Reviewing #2 raised a GPU-capacity question that turned out to be real and
|
|
already live, not hypothetical. Measured on this box's Quadro RTX 6000
|
|
(24GB): `mistral-nemo:12b` (classifier) and `qwen3-vl:4b` (`local_vision`
|
|
fallback), both loaded at Ollama's untagged default context (32768),
|
|
together used 22.1GB — 1.9GB free, one browser tab of GPU use away from
|
|
contention. Neither model needed anywhere near that: `max_input_chars` +
|
|
`system_prompt` + `max_output_tokens` bound the classifier well under 4k
|
|
tokens, and 32768 was never a value anyone chose — it's just what the base
|
|
model's library Modelfile defaults to.
|
|
|
|
Fixed by baking a right-sized context into a Modelfile-tagged variant of
|
|
each model (`mistral-nemo-router:12b` @ 8192, `qwen3-vl-router:4b` @ 16384)
|
|
and pointing `classifier.model` / `verification.model` / `local_vision.model`
|
|
at the tags instead of the base ones. Verified live: 8.6GB and 6.9GB
|
|
resident respectively, 15.9GB combined worst case (both hot at once) versus
|
|
22.1GB before — free VRAM went from 1.9GB to 8.1GB in that case, and from
|
|
10.8GB to 14.6GB in the common case (classifier alone). Confirmed
|
|
empirically along the way, not assumed: Ollama's OpenAI-compatible endpoint
|
|
(0.22.0) silently ignores a per-request `num_ctx` or `keep_alive` override
|
|
under every field shape tried (`options.num_ctx`, top-level `num_ctx`,
|
|
`context_length`, top-level `keep_alive`) — a 200 comes back and nothing
|
|
changes. Only the native `/api/chat` endpoint honors either, which is why
|
|
the fix has to live in the model tag rather than in a request parameter,
|
|
and why a shorter `keep_alive` for the vision fallback specifically was
|
|
left as a follow-up (it would need that path rewritten onto the native API,
|
|
including its image format).
|
|
|
|
This doesn't change the verdict on #2: headroom existing now is not the
|
|
same as local generation being cheaper than a cloud completion, which is
|
|
still unmeasured. It does mean the box has real spare VRAM again, which is
|
|
one less reason to reach for anything drastic to free some.
|
|
|
|
## Overall
|
|
|
|
None of the three should be built as pitched. #1's real, low-cost half (the
|
|
timer) is worth doing now. #2's cost mechanism doesn't hold up against this
|
|
project's own numbers and depends on an open item (#5) that isn't done yet
|
|
— hold it. #3 is the one worth real design time, once scoped down to what
|
|
pinch actually touches.
|