# Documentation sweep and context diet Status: done -- docs/ reference set **Status: FINAL — decision-complete, ready for planning.** Every number in this plan was measured on 2026-09-02 against the working tree at `feat/pinch-instrumentation-and-token-accounting` (`d7f051b`). Re-measure before trusting any of them; the point of the plan is that they drift. ## Why now Three unrelated pressures landed on the same files at once: 1. **`CLAUDE.md` is loaded into every session.** It is 72,549 bytes, roughly 24,000 tokens, paid on every single session in this repo before a word of work happens. It has grown by accretion — each correction appended, none retired. 2. **A core dispatch-path feature is undocumented.** `pinch` prunes context before any paid token is sent upstream. It has a config block, a hard position in the dispatch path, and now a metrics surface and an admin card. It is mentioned in `CLAUDE.md` exactly **once** (line 1156, an incidental parenthetical about local-vision context sizing) and **zero** times across the eleven files in `docs/`. 3. **Load-bearing numbers have gone stale**, and they are the kind a reader reasonably trusts. None of these is urgent alone. Together they mean the file that is supposed to be "the one to trust on what is currently true" is neither current nor cheap. ## The organizing principle `CLAUDE.md` already states its own job: `README.md` is the front door, `docs/` is the module-by-module reference, `design/` holds architecture and rationale, and `CLAUDE.md` is "the working state + immediate next steps, and is the one to trust on what is currently true." It has drifted from that. Much of its bulk is *evidence for* conclusions — measurement tables, split-half checks, superseded reasoning — rather than the current state itself. So the rule for this sweep: > **`CLAUDE.md` keeps what changes a decision inside a session. `docs/` keeps > the evidence that established it.** **Relocate, never delete.** This project deliberately records *why*, including conclusions it later overturned — the cost-axis inversion, the tier-on-price mistake, the four harness bugs that scored the rig rather than the model. That history is the most valuable content here and several sections exist purely to stop a future reader re-making the same substitution. Every extraction leaves a 2–4 line summary plus a link. A fact that exists only in one place after this sweep has been lost, and that is a failure, not a saving. ## Part A — Correct the stale facts Measured against the live `router.db` and a real collection run: | `CLAUDE.md` says | Actually | Where | |---|---|---| | "796 tests across 40 files" | **860 tests across 44 files** | line 291 | | "19 models" | **20** | line 205 | | "13 of 19 rows are routable" | **14 of 20** (14 public, 4 preview, 2 canary) | line 435 | The catalog moved because `glm-5.3` was reactivated. Note the second and third are the same drift surfacing twice, which is the argument for deriving them rather than restating them: prefer "the poller reports N; `access_level` excludes preview and canary" over a hard-coded pair of integers that must be edited in two places and will not be. Historical measurements that carry a date — the 65-sample sweep, the $8.00/kWh table, the attribution split-half figures — are **observations, not claims about now**, and stay as they are. They are already framed that way. Do not "correct" them. Audit `What's NOT built yet` (3,368 bytes) in the same pass: items 4 and 5 are already marked built and describe shipped behaviour, so they are no longer "not built" and belong in `What's built and working` or in `docs/`. Confirm the status of items 1–3 rather than assuming. ## Part B — Document pinch This is the largest genuine gap and the highest-value half of the sweep. **Add `docs/pinch.md`**, matching the shape of the existing reference docs (`routing.md`, `verification.md`, `data-model.md`). It must cover: - What pinch is and where it sits in the dispatch path — before the measured-context decision on the routed path, so window/tier/cost selection sees the size that will actually ship; separately on the passthrough path. - That it trims **only tool results**, never user/assistant/system messages, and never the classifier input (which clamps to head+tail independently). - The config block: `enabled`, `budget_tokens` (50000), `keep_last_turns`, `max_summarize_chars`, `relevance`. - **The token accounting**, including the tool-definition overhead. opencode sends ~32k prompt tokens of tool definitions on a trivial request, and until `d7f051b` pinch excluded all of it from the budget comparison. Say why `_relevance_order_for` takes `extra_fixed_tokens` too: without it, relevance ordering silently declines to run on exactly the requests that newly need pruning — no error, no log, just uniform instead of least-relevant-first. - **The instrumentation**: `route_decisions.pinch_original_tokens` / `pinch_final_tokens`, `metrics.pinch_summary`, the `pinch` block on `/metrics` and `/admin/api/snapshot`, the TUI rows and the admin card. - **How dollars-saved is priced**, and why it must stay that way: the blended rate from `routing.estimated_cost` (`src/routing.py:363-374`), reading `cfg.objective.assumed_cache_rate`. Record the error that was caught in review — pricing at the cached rate alone double-applies the discount *and* drops the fraction billed at full price, producing a dashboard figure that disagrees with the cost model the router ranks on. This is exactly the class of mistake this repo documents rather than quietly fixes. - **The passthrough ordering constraint.** The prune is hoisted above the persist so stats are recorded rather than NULL, but the persist stays above `_check_pinned_capabilities`, which raises 422 — so a rejected pin still writes the decision row that explains it. Anyone "tidying" that order will silently delete rejection records. Then add a **short** pinch entry to `CLAUDE.md`'s `What's built and working` list — a few lines and a link to `docs/pinch.md`, in the style of the existing bullets. Not a full section; that is what Part C is trying to undo. Update `docs/data-model.md` for the two new columns and `docs/api.md` for the `pinch` block on `/metrics`. `docs/admin-portal.md` gains the card. ## Part C — The context diet Three sections are **51.1%** of the file: | section | bytes | share | |---|---|---| | `What's built and working` | 18,440 | 25.8% | | `Proficiency: category now changes routing` | 10,794 | 15.1% | | `The classifier is the latency floor` | 7,307 | 10.2% | | — | **36,541** | **51.1%** | Target: **`CLAUDE.md` ≤ 40,000 bytes**, from 72,549. That is ~11,000 tokens saved per session, and it is reachable by relocating these three plus the already-identified surplus, without touching anything else. Precedent: `docs/incidents.md` was extracted this way and it worked — the symptom → one-line-check table is more useful in its own file than it was buried mid-`CLAUDE.md`, and `CLAUDE.md` kept a four-line pointer. Proposed destinations: - **`What's built and working`** → the per-module detail moves to `docs/` (most modules already have a home; `architecture.md` can carry the rest). `CLAUDE.md` keeps a compact inventory: module name, one line, link. The sub-sections that are really *rationale* — the circuit-breaker/eval-harness isolation argument, the request-side capability gates and the fail-closed asymmetry — move to `docs/routing.md` and `docs/architecture.md`. - **`Proficiency: category now changes routing`** → `docs/evaluation.md` and `docs/routing.md`. The measurement tables, the `docs_writing` sampling story, the harness-bug list and the per-category spread table are all evidence. `CLAUDE.md` keeps the conclusions that change behaviour: category affects routing, one winner can be legitimate dominance, check for that before reaching for `quality_tolerance`, and the lever is `quality_tolerance` / `max_energy_per_request`. - **`The classifier is the latency floor`** → `docs/local-models.md`. The `qwen3.5` vs `mistral-nemo` comparison table, the num_ctx/Modelfile investigation and the cloud-classifier measurement are all reference material. `CLAUDE.md` keeps: the classifier is the latency floor, the four settings that keep it usable and why each exists, and that "local" means your hardware, not this machine. **Non-negotiable retentions.** These stay in `CLAUDE.md` verbatim regardless of length, because they are the rules that bind a session and a reader who misses them causes an incident: - 8080 is production; throwaway instances bind 8081; never signal a process matched by name or port. - Never point `config/config.yaml` at test fixtures. - The service holds a billable key and has no auth; loopback is the only protection. Same for Ollama. - Config is strict — an unknown key is an error, and every knob belongs in `config.yaml`, not only in a Pydantic default. - The `docs/incidents.md` pointer and its recovery one-liner. ## Part D — Scope boundary with the unmerged proficiency branch `feat/proficiency-exposure-bias-and-exploration` (5 commits) adds `src/exploration.py`, the empirical-Bayes blending and the three-way source taxonomy. **None of it is documented, and none of it is on `main`.** **Decision: this sweep documents only what is on `main` plus the pinch branch.** Writing docs for unmerged behaviour would put the repo in the state this plan exists to fix, one branch earlier. Documenting exploration is a task for that branch's own merge, and this plan should say so explicitly so it is not silently dropped. Verified while scoping: `docs/routing.md` line 35 uses the phrase "empirical question" in an unrelated sentence. It is not documentation of empirical Bayes, and `docs/` is not currently ahead of code. ## Non-goals - No rewriting of `design/local-llm-model-router.md`. It is explicitly the architecture-and-rationale document including unbuilt parts, and it is allowed to describe things that do not exist. - No changes to `AGENTS.md` beyond factual corrections. Its conventions came out of incident #1 and are load-bearing as written. - No new measurements. Where a number is unknown, the sweep reports it as unknown rather than generating one. - No prose polish for its own sake. Every edit is a correction, a relocation, or a deletion of something now stated elsewhere. ## Success criteria - `CLAUDE.md` ≤ 40,000 bytes (from 72,549), with every relocated section leaving a summary and a working link. - `docs/pinch.md` exists and covers accounting, instrumentation, the blended pricing rule, and the passthrough ordering constraint. - The three stale figures are corrected, and the model counts are derived rather than restated where practical. - `What's NOT built yet` lists only things that are genuinely not built. - **No fact lost.** For each extracted section, the reviewer can name where each load-bearing claim now lives. This is the criterion most likely to be quietly failed, so check it explicitly rather than trusting the diff. - Every internal link resolves — relative paths, correct filenames. A sweep that halves the file and breaks the pointers has made things worse. - The non-negotiable retentions above are still present verbatim. - No code changes. This plan touches documentation only; if it turns out a doc is wrong because the *code* is wrong, record it and stop rather than fixing it here.