plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
225 lines
11 KiB
Markdown
225 lines
11 KiB
Markdown
# Documentation sweep and context diet
|
||
|
||
Status: done -- docs/ reference set
|
||
|
||
**Status: FINAL — decision-complete, ready for planning.**
|
||
|
||
Every number in this plan was measured on 2026-09-02 against the working tree
|
||
at `feat/pinch-instrumentation-and-token-accounting` (`d7f051b`). Re-measure
|
||
before trusting any of them; the point of the plan is that they drift.
|
||
|
||
## Why now
|
||
|
||
Three unrelated pressures landed on the same files at once:
|
||
|
||
1. **`CLAUDE.md` is loaded into every session.** It is 72,549 bytes, roughly
|
||
24,000 tokens, paid on every single session in this repo before a word of
|
||
work happens. It has grown by accretion — each correction appended, none
|
||
retired.
|
||
2. **A core dispatch-path feature is undocumented.** `pinch` prunes context
|
||
before any paid token is sent upstream. It has a config block, a hard
|
||
position in the dispatch path, and now a metrics surface and an admin
|
||
card. It is mentioned in `CLAUDE.md` exactly **once** (line 1156, an
|
||
incidental parenthetical about local-vision context sizing) and **zero**
|
||
times across the eleven files in `docs/`.
|
||
3. **Load-bearing numbers have gone stale**, and they are the kind a reader
|
||
reasonably trusts.
|
||
|
||
None of these is urgent alone. Together they mean the file that is supposed
|
||
to be "the one to trust on what is currently true" is neither current nor
|
||
cheap.
|
||
|
||
## The organizing principle
|
||
|
||
`CLAUDE.md` already states its own job: `README.md` is the front door, `docs/`
|
||
is the module-by-module reference, `design/` holds architecture and rationale,
|
||
and `CLAUDE.md` is "the working state + immediate next steps, and is the one
|
||
to trust on what is currently true."
|
||
|
||
It has drifted from that. Much of its bulk is *evidence for* conclusions —
|
||
measurement tables, split-half checks, superseded reasoning — rather than the
|
||
current state itself.
|
||
|
||
So the rule for this sweep:
|
||
|
||
> **`CLAUDE.md` keeps what changes a decision inside a session. `docs/` keeps
|
||
> the evidence that established it.**
|
||
|
||
**Relocate, never delete.** This project deliberately records *why*, including
|
||
conclusions it later overturned — the cost-axis inversion, the tier-on-price
|
||
mistake, the four harness bugs that scored the rig rather than the model. That
|
||
history is the most valuable content here and several sections exist purely to
|
||
stop a future reader re-making the same substitution. Every extraction leaves
|
||
a 2–4 line summary plus a link. A fact that exists only in one place after
|
||
this sweep has been lost, and that is a failure, not a saving.
|
||
|
||
## Part A — Correct the stale facts
|
||
|
||
Measured against the live `router.db` and a real collection run:
|
||
|
||
| `CLAUDE.md` says | Actually | Where |
|
||
|---|---|---|
|
||
| "796 tests across 40 files" | **860 tests across 44 files** | line 291 |
|
||
| "19 models" | **20** | line 205 |
|
||
| "13 of 19 rows are routable" | **14 of 20** (14 public, 4 preview, 2 canary) | line 435 |
|
||
|
||
The catalog moved because `glm-5.3` was reactivated. Note the second and third
|
||
are the same drift surfacing twice, which is the argument for deriving them
|
||
rather than restating them: prefer "the poller reports N; `access_level`
|
||
excludes preview and canary" over a hard-coded pair of integers that must be
|
||
edited in two places and will not be.
|
||
|
||
Historical measurements that carry a date — the 65-sample sweep, the $8.00/kWh
|
||
table, the attribution split-half figures — are **observations, not claims
|
||
about now**, and stay as they are. They are already framed that way. Do not
|
||
"correct" them.
|
||
|
||
Audit `What's NOT built yet` (3,368 bytes) in the same pass: items 4 and 5 are
|
||
already marked built and describe shipped behaviour, so they are no longer
|
||
"not built" and belong in `What's built and working` or in `docs/`. Confirm the
|
||
status of items 1–3 rather than assuming.
|
||
|
||
## Part B — Document pinch
|
||
|
||
This is the largest genuine gap and the highest-value half of the sweep.
|
||
|
||
**Add `docs/pinch.md`**, matching the shape of the existing reference docs
|
||
(`routing.md`, `verification.md`, `data-model.md`). It must cover:
|
||
|
||
- What pinch is and where it sits in the dispatch path — before the
|
||
measured-context decision on the routed path, so window/tier/cost selection
|
||
sees the size that will actually ship; separately on the passthrough path.
|
||
- That it trims **only tool results**, never user/assistant/system messages,
|
||
and never the classifier input (which clamps to head+tail independently).
|
||
- The config block: `enabled`, `budget_tokens` (50000), `keep_last_turns`,
|
||
`max_summarize_chars`, `relevance`.
|
||
- **The token accounting**, including the tool-definition overhead. opencode
|
||
sends ~32k prompt tokens of tool definitions on a trivial request, and until
|
||
`d7f051b` pinch excluded all of it from the budget comparison. Say why
|
||
`_relevance_order_for` takes `extra_fixed_tokens` too: without it, relevance
|
||
ordering silently declines to run on exactly the requests that newly need
|
||
pruning — no error, no log, just uniform instead of least-relevant-first.
|
||
- **The instrumentation**: `route_decisions.pinch_original_tokens` /
|
||
`pinch_final_tokens`, `metrics.pinch_summary`, the `pinch` block on
|
||
`/metrics` and `/admin/api/snapshot`, the TUI rows and the admin card.
|
||
- **How dollars-saved is priced**, and why it must stay that way: the blended
|
||
rate from `routing.estimated_cost` (`src/routing.py:363-374`), reading
|
||
`cfg.objective.assumed_cache_rate`. Record the error that was caught in
|
||
review — pricing at the cached rate alone double-applies the discount *and*
|
||
drops the fraction billed at full price, producing a dashboard figure that
|
||
disagrees with the cost model the router ranks on. This is exactly the class
|
||
of mistake this repo documents rather than quietly fixes.
|
||
- **The passthrough ordering constraint.** The prune is hoisted above the
|
||
persist so stats are recorded rather than NULL, but the persist stays above
|
||
`_check_pinned_capabilities`, which raises 422 — so a rejected pin still
|
||
writes the decision row that explains it. Anyone "tidying" that order will
|
||
silently delete rejection records.
|
||
|
||
Then add a **short** pinch entry to `CLAUDE.md`'s `What's built and working`
|
||
list — a few lines and a link to `docs/pinch.md`, in the style of the existing
|
||
bullets. Not a full section; that is what Part C is trying to undo.
|
||
|
||
Update `docs/data-model.md` for the two new columns and `docs/api.md` for the
|
||
`pinch` block on `/metrics`. `docs/admin-portal.md` gains the card.
|
||
|
||
## Part C — The context diet
|
||
|
||
Three sections are **51.1%** of the file:
|
||
|
||
| section | bytes | share |
|
||
|---|---|---|
|
||
| `What's built and working` | 18,440 | 25.8% |
|
||
| `Proficiency: category now changes routing` | 10,794 | 15.1% |
|
||
| `The classifier is the latency floor` | 7,307 | 10.2% |
|
||
| — | **36,541** | **51.1%** |
|
||
|
||
Target: **`CLAUDE.md` ≤ 40,000 bytes**, from 72,549. That is ~11,000 tokens
|
||
saved per session, and it is reachable by relocating these three plus the
|
||
already-identified surplus, without touching anything else.
|
||
|
||
Precedent: `docs/incidents.md` was extracted this way and it worked — the
|
||
symptom → one-line-check table is more useful in its own file than it was
|
||
buried mid-`CLAUDE.md`, and `CLAUDE.md` kept a four-line pointer.
|
||
|
||
Proposed destinations:
|
||
|
||
- **`What's built and working`** → the per-module detail moves to `docs/` (most
|
||
modules already have a home; `architecture.md` can carry the rest).
|
||
`CLAUDE.md` keeps a compact inventory: module name, one line, link. The
|
||
sub-sections that are really *rationale* — the circuit-breaker/eval-harness
|
||
isolation argument, the request-side capability gates and the fail-closed
|
||
asymmetry — move to `docs/routing.md` and `docs/architecture.md`.
|
||
- **`Proficiency: category now changes routing`** → `docs/evaluation.md` and
|
||
`docs/routing.md`. The measurement tables, the `docs_writing` sampling
|
||
story, the harness-bug list and the per-category spread table are all
|
||
evidence. `CLAUDE.md` keeps the conclusions that change behaviour: category
|
||
affects routing, one winner can be legitimate dominance, check for that
|
||
before reaching for `quality_tolerance`, and the lever is
|
||
`quality_tolerance` / `max_energy_per_request`.
|
||
- **`The classifier is the latency floor`** → `docs/local-models.md`. The
|
||
`qwen3.5` vs `mistral-nemo` comparison table, the num_ctx/Modelfile
|
||
investigation and the cloud-classifier measurement are all reference
|
||
material. `CLAUDE.md` keeps: the classifier is the latency floor, the four
|
||
settings that keep it usable and why each exists, and that "local" means
|
||
your hardware, not this machine.
|
||
|
||
**Non-negotiable retentions.** These stay in `CLAUDE.md` verbatim regardless
|
||
of length, because they are the rules that bind a session and a reader who
|
||
misses them causes an incident:
|
||
|
||
- 8080 is production; throwaway instances bind 8081; never signal a process
|
||
matched by name or port.
|
||
- Never point `config/config.yaml` at test fixtures.
|
||
- The service holds a billable key and has no auth; loopback is the only
|
||
protection. Same for Ollama.
|
||
- Config is strict — an unknown key is an error, and every knob belongs in
|
||
`config.yaml`, not only in a Pydantic default.
|
||
- The `docs/incidents.md` pointer and its recovery one-liner.
|
||
|
||
## Part D — Scope boundary with the unmerged proficiency branch
|
||
|
||
`feat/proficiency-exposure-bias-and-exploration` (5 commits) adds
|
||
`src/exploration.py`, the empirical-Bayes blending and the three-way source
|
||
taxonomy. **None of it is documented, and none of it is on `main`.**
|
||
|
||
**Decision: this sweep documents only what is on `main` plus the pinch
|
||
branch.** Writing docs for unmerged behaviour would put the repo in the state
|
||
this plan exists to fix, one branch earlier. Documenting exploration is a task
|
||
for that branch's own merge, and this plan should say so explicitly so it is
|
||
not silently dropped.
|
||
|
||
Verified while scoping: `docs/routing.md` line 35 uses the phrase "empirical
|
||
question" in an unrelated sentence. It is not documentation of empirical
|
||
Bayes, and `docs/` is not currently ahead of code.
|
||
|
||
## Non-goals
|
||
|
||
- No rewriting of `design/local-llm-model-router.md`. It is explicitly the
|
||
architecture-and-rationale document including unbuilt parts, and it is
|
||
allowed to describe things that do not exist.
|
||
- No changes to `AGENTS.md` beyond factual corrections. Its conventions came
|
||
out of incident #1 and are load-bearing as written.
|
||
- No new measurements. Where a number is unknown, the sweep reports it as
|
||
unknown rather than generating one.
|
||
- No prose polish for its own sake. Every edit is a correction, a relocation,
|
||
or a deletion of something now stated elsewhere.
|
||
|
||
## Success criteria
|
||
|
||
- `CLAUDE.md` ≤ 40,000 bytes (from 72,549), with every relocated section
|
||
leaving a summary and a working link.
|
||
- `docs/pinch.md` exists and covers accounting, instrumentation, the blended
|
||
pricing rule, and the passthrough ordering constraint.
|
||
- The three stale figures are corrected, and the model counts are derived
|
||
rather than restated where practical.
|
||
- `What's NOT built yet` lists only things that are genuinely not built.
|
||
- **No fact lost.** For each extracted section, the reviewer can name where
|
||
each load-bearing claim now lives. This is the criterion most likely to be
|
||
quietly failed, so check it explicitly rather than trusting the diff.
|
||
- Every internal link resolves — relative paths, correct filenames. A sweep
|
||
that halves the file and breaks the pointers has made things worse.
|
||
- The non-negotiable retentions above are still present verbatim.
|
||
- No code changes. This plan touches documentation only; if it turns out a doc
|
||
is wrong because the *code* is wrong, record it and stop rather than fixing
|
||
it here.
|