plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
11 KiB
Documentation sweep and context diet
Status: done -- docs/ reference set
Status: FINAL — decision-complete, ready for planning.
Every number in this plan was measured on 2026-09-02 against the working tree
at feat/pinch-instrumentation-and-token-accounting (d7f051b). Re-measure
before trusting any of them; the point of the plan is that they drift.
Why now
Three unrelated pressures landed on the same files at once:
CLAUDE.mdis loaded into every session. It is 72,549 bytes, roughly 24,000 tokens, paid on every single session in this repo before a word of work happens. It has grown by accretion — each correction appended, none retired.- A core dispatch-path feature is undocumented.
pinchprunes context before any paid token is sent upstream. It has a config block, a hard position in the dispatch path, and now a metrics surface and an admin card. It is mentioned inCLAUDE.mdexactly once (line 1156, an incidental parenthetical about local-vision context sizing) and zero times across the eleven files indocs/. - Load-bearing numbers have gone stale, and they are the kind a reader reasonably trusts.
None of these is urgent alone. Together they mean the file that is supposed to be "the one to trust on what is currently true" is neither current nor cheap.
The organizing principle
CLAUDE.md already states its own job: README.md is the front door, docs/
is the module-by-module reference, design/ holds architecture and rationale,
and CLAUDE.md is "the working state + immediate next steps, and is the one
to trust on what is currently true."
It has drifted from that. Much of its bulk is evidence for conclusions — measurement tables, split-half checks, superseded reasoning — rather than the current state itself.
So the rule for this sweep:
CLAUDE.mdkeeps what changes a decision inside a session.docs/keeps the evidence that established it.
Relocate, never delete. This project deliberately records why, including conclusions it later overturned — the cost-axis inversion, the tier-on-price mistake, the four harness bugs that scored the rig rather than the model. That history is the most valuable content here and several sections exist purely to stop a future reader re-making the same substitution. Every extraction leaves a 2–4 line summary plus a link. A fact that exists only in one place after this sweep has been lost, and that is a failure, not a saving.
Part A — Correct the stale facts
Measured against the live router.db and a real collection run:
CLAUDE.md says |
Actually | Where |
|---|---|---|
| "796 tests across 40 files" | 860 tests across 44 files | line 291 |
| "19 models" | 20 | line 205 |
| "13 of 19 rows are routable" | 14 of 20 (14 public, 4 preview, 2 canary) | line 435 |
The catalog moved because glm-5.3 was reactivated. Note the second and third
are the same drift surfacing twice, which is the argument for deriving them
rather than restating them: prefer "the poller reports N; access_level
excludes preview and canary" over a hard-coded pair of integers that must be
edited in two places and will not be.
Historical measurements that carry a date — the 65-sample sweep, the $8.00/kWh table, the attribution split-half figures — are observations, not claims about now, and stay as they are. They are already framed that way. Do not "correct" them.
Audit What's NOT built yet (3,368 bytes) in the same pass: items 4 and 5 are
already marked built and describe shipped behaviour, so they are no longer
"not built" and belong in What's built and working or in docs/. Confirm the
status of items 1–3 rather than assuming.
Part B — Document pinch
This is the largest genuine gap and the highest-value half of the sweep.
Add docs/pinch.md, matching the shape of the existing reference docs
(routing.md, verification.md, data-model.md). It must cover:
- What pinch is and where it sits in the dispatch path — before the measured-context decision on the routed path, so window/tier/cost selection sees the size that will actually ship; separately on the passthrough path.
- That it trims only tool results, never user/assistant/system messages, and never the classifier input (which clamps to head+tail independently).
- The config block:
enabled,budget_tokens(50000),keep_last_turns,max_summarize_chars,relevance. - The token accounting, including the tool-definition overhead. opencode
sends ~32k prompt tokens of tool definitions on a trivial request, and until
d7f051bpinch excluded all of it from the budget comparison. Say why_relevance_order_fortakesextra_fixed_tokenstoo: without it, relevance ordering silently declines to run on exactly the requests that newly need pruning — no error, no log, just uniform instead of least-relevant-first. - The instrumentation:
route_decisions.pinch_original_tokens/pinch_final_tokens,metrics.pinch_summary, thepinchblock on/metricsand/admin/api/snapshot, the TUI rows and the admin card. - How dollars-saved is priced, and why it must stay that way: the blended
rate from
routing.estimated_cost(src/routing.py:363-374), readingcfg.objective.assumed_cache_rate. Record the error that was caught in review — pricing at the cached rate alone double-applies the discount and drops the fraction billed at full price, producing a dashboard figure that disagrees with the cost model the router ranks on. This is exactly the class of mistake this repo documents rather than quietly fixes. - The passthrough ordering constraint. The prune is hoisted above the
persist so stats are recorded rather than NULL, but the persist stays above
_check_pinned_capabilities, which raises 422 — so a rejected pin still writes the decision row that explains it. Anyone "tidying" that order will silently delete rejection records.
Then add a short pinch entry to CLAUDE.md's What's built and working
list — a few lines and a link to docs/pinch.md, in the style of the existing
bullets. Not a full section; that is what Part C is trying to undo.
Update docs/data-model.md for the two new columns and docs/api.md for the
pinch block on /metrics. docs/admin-portal.md gains the card.
Part C — The context diet
Three sections are 51.1% of the file:
| section | bytes | share |
|---|---|---|
What's built and working |
18,440 | 25.8% |
Proficiency: category now changes routing |
10,794 | 15.1% |
The classifier is the latency floor |
7,307 | 10.2% |
| — | 36,541 | 51.1% |
Target: CLAUDE.md ≤ 40,000 bytes, from 72,549. That is ~11,000 tokens
saved per session, and it is reachable by relocating these three plus the
already-identified surplus, without touching anything else.
Precedent: docs/incidents.md was extracted this way and it worked — the
symptom → one-line-check table is more useful in its own file than it was
buried mid-CLAUDE.md, and CLAUDE.md kept a four-line pointer.
Proposed destinations:
What's built and working→ the per-module detail moves todocs/(most modules already have a home;architecture.mdcan carry the rest).CLAUDE.mdkeeps a compact inventory: module name, one line, link. The sub-sections that are really rationale — the circuit-breaker/eval-harness isolation argument, the request-side capability gates and the fail-closed asymmetry — move todocs/routing.mdanddocs/architecture.md.Proficiency: category now changes routing→docs/evaluation.mdanddocs/routing.md. The measurement tables, thedocs_writingsampling story, the harness-bug list and the per-category spread table are all evidence.CLAUDE.mdkeeps the conclusions that change behaviour: category affects routing, one winner can be legitimate dominance, check for that before reaching forquality_tolerance, and the lever isquality_tolerance/max_energy_per_request.The classifier is the latency floor→docs/local-models.md. Theqwen3.5vsmistral-nemocomparison table, the num_ctx/Modelfile investigation and the cloud-classifier measurement are all reference material.CLAUDE.mdkeeps: the classifier is the latency floor, the four settings that keep it usable and why each exists, and that "local" means your hardware, not this machine.
Non-negotiable retentions. These stay in CLAUDE.md verbatim regardless
of length, because they are the rules that bind a session and a reader who
misses them causes an incident:
- 8080 is production; throwaway instances bind 8081; never signal a process matched by name or port.
- Never point
config/config.yamlat test fixtures. - The service holds a billable key and has no auth; loopback is the only protection. Same for Ollama.
- Config is strict — an unknown key is an error, and every knob belongs in
config.yaml, not only in a Pydantic default. - The
docs/incidents.mdpointer and its recovery one-liner.
Part D — Scope boundary with the unmerged proficiency branch
feat/proficiency-exposure-bias-and-exploration (5 commits) adds
src/exploration.py, the empirical-Bayes blending and the three-way source
taxonomy. None of it is documented, and none of it is on main.
Decision: this sweep documents only what is on main plus the pinch
branch. Writing docs for unmerged behaviour would put the repo in the state
this plan exists to fix, one branch earlier. Documenting exploration is a task
for that branch's own merge, and this plan should say so explicitly so it is
not silently dropped.
Verified while scoping: docs/routing.md line 35 uses the phrase "empirical
question" in an unrelated sentence. It is not documentation of empirical
Bayes, and docs/ is not currently ahead of code.
Non-goals
- No rewriting of
design/local-llm-model-router.md. It is explicitly the architecture-and-rationale document including unbuilt parts, and it is allowed to describe things that do not exist. - No changes to
AGENTS.mdbeyond factual corrections. Its conventions came out of incident #1 and are load-bearing as written. - No new measurements. Where a number is unknown, the sweep reports it as unknown rather than generating one.
- No prose polish for its own sake. Every edit is a correction, a relocation, or a deletion of something now stated elsewhere.
Success criteria
CLAUDE.md≤ 40,000 bytes (from 72,549), with every relocated section leaving a summary and a working link.docs/pinch.mdexists and covers accounting, instrumentation, the blended pricing rule, and the passthrough ordering constraint.- The three stale figures are corrected, and the model counts are derived rather than restated where practical.
What's NOT built yetlists only things that are genuinely not built.- No fact lost. For each extracted section, the reviewer can name where each load-bearing claim now lives. This is the criterion most likely to be quietly failed, so check it explicitly rather than trusting the diff.
- Every internal link resolves — relative paths, correct filenames. A sweep that halves the file and breaks the pointers has made things worse.
- The non-negotiable retentions above are still present verbatim.
- No code changes. This plan touches documentation only; if it turns out a doc is wrong because the code is wrong, record it and stop rather than fixing it here.