plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
8.9 KiB
Named routing profiles
Status: done -- BUILTIN_PROFILES plus overlay profiles
Status: FINAL — decision-complete. §3 was answered by the user on 2026-09-03
(filter only). Written 2026-09-03 against main at c2893b5.
What this is
Today the router has exactly one behaviour: classify, hard-filter, rank
quality-first with cost as a tiebreak. The user wants named alternatives —
locality, onlycheaps, bigboybritches — alongside a default that stays
the feature-rich automatic router it is now.
The mechanism already exists and is half-built. auto:batch is a profile:
a named variant of auto that changes one hard-filter parameter
(latency_tolerance), documented in README and shipped for months. This plan
generalises that one hard-coded special case into a named, configurable set.
It is deliberately NOT a new ranking path.
1. Where it slots in, precisely
Three call sites, all already doing most of the work:
src/dispatcher.py:2937—wants_routing = bare in (ROUTER_MODEL, ROUTER_MODEL_BATCH). An exact membership test against two constants. This becomes a parse: splitauto:<profile>, look the profile up, reject an unknown one.src/dispatcher.py:2955—latency = BATCH if requested == ROUTER_MODEL_BATCH else INTERACTIVE. This is the existing profile, expressed as a ternary. It must become one case of the general mechanism, not survive alongside it.src/routing.py:313select_candidates— already takesexclude_models: set[str],task_category,latency_tolerance,allowed_access_levels,min_tool_proficiency. A profile supplies or overrides these. All the filtering machinery is already there.
routing.py is a pure module with injected dependencies and
rejection_reason is the single copy of the rules (is_eligible is a thin
wrapper over it). Keep that: a profile must not introduce a second place where
eligibility is decided.
2. The one genuinely missing primitive
select_candidates has a denylist (exclude_models) and no allowlist.
Every profile the user named is naturally an allowlist:
locality→ onlyprovider = 'ollama-local'rowsonlycheaps→ only rows under a cost barbigboybritches→ only frontier rows
Add one parameter, restrict_to: set[str] | None = None, threaded through
rejection_reason → is_eligible → select_candidates, with None meaning
"unrestricted" (NOT an empty set, which must mean "nothing eligible"). Give it
its own rejection_reason string so a 422 says which profile emptied the
candidate set — the catalog-staleness incident is the precedent: an empty
candidate set with an opaque reason cost ~19 hours to diagnose.
Prefer expressing profiles as predicates over catalog columns (provider, cost, tier) rather than hard-coded model-id lists, so a profile does not go stale the moment the poller adds a row. A literal id list stays available for the cases where that is genuinely what the user means.
3. THE OPEN DECISION — filter only, or objective overrides too?
DECIDED 2026-09-03 by the user: FILTER ONLY. A profile restricts the
candidate set and nothing else. Ranking stays quality-first with cost as the
tiebreak, one rule for the whole system. Do NOT implement objective overrides,
and do NOT add a quality_tolerance field to a profile definition — if a
future preference-shaped profile needs one, that is a separate plan with its
own justification.
The reasoning is preserved below because it explains WHY filter-only is sufficient, which is not obvious.
A profile is unambiguously a candidate-set filter. The question is whether it
may ALSO override objective.* (quality_tolerance,
max_energy_per_request).
It matters concretely, and we have the measurement:
plans/local-dispatch-inert-and-test-coupling.md established that the local
model scores 0.767 on file_summarization against cloud's 0.95 — a 0.183 gap
against a quality_tolerance of 0.1 — so local never routes under
quality-first ranking.
- A pure-filter
localitystill works: restrict the set to local rows and the best local model wins by default, because there is no cloud row left to lose to. This covers the user's stated case. - A mixed profile that prefers local without excluding cloud does NOT work as a pure filter. It reproduces the dormancy exactly.
So: pure filters satisfy the three named profiles. Objective overrides are only needed for preference-shaped profiles nobody has asked for yet.
Recommendation: ship filter-only. It is the smaller change, it satisfies
every named use case, and it keeps one ranking rule in the system. Add
overrides later if a real preference-shaped profile appears. If the user wants
overrides now, they must be scoped strictly to the profile and never mutate
global config — and quality_tolerance in particular describes measurement
noise on 2-3 samples, so overriding it means asserting a preference through a
knob that means something else.
4. Naming and selection
Profiles are selected the way auto:batch already is: the model field.
auto -> default profile
auto:batch -> MUST keep working, byte-identical behaviour
auto:locality -> named profile
<real model id> -> passthrough, unchanged
auto:batch is load-bearing: it is in README, docs/, and possibly user
clients. Reimplement it as a profile whose definition sets
latency_tolerance: batch, and pin its behaviour with a test that would fail
if the general mechanism changed it. Do not leave the ternary in place beside
the new path — two mechanisms for one behaviour is how they drift.
An unknown profile (auto:nonsense) must fail loudly with a 422 naming the
valid profiles. It must NOT silently fall back to default: a user who
mistypes auto:locallity would otherwise get cloud routing and a surprising
bill, with nothing indicating why.
Config lives under a new top-level profiles: key. Config is strict
(extra="forbid"), so every field needs declaring; an unknown key is an error
by design.
5. Observability
route_decisions must record which profile served each decision, or the
proficiency loop cannot tell a bigboybritches win from a default one and
will read profile-forced selections as evidence of model quality. This is the
same reasoning that made the local fallback record
kind='local_dispatch_fallback' rather than chat.
Add a profile column (code-side migration in ensure_route_decisions, as
every prior column addition did — config/schema.sql is
CREATE TABLE IF NOT EXISTS and does not alter a live DB). Surface it in the
admin decisions table and its filters, which already filter by kind, category
and tier.
6. Interactions to get right
eligible_categories(local dispatch) is a per-MODEL filter; a profile is a per-REQUEST filter. They compose as AND. Alocalityprofile must not bypasseligible_categories— that gate is what keeps a summarization-grade model away from code, andnemotron-mini:4bdetecting 0/4 bugs is why.- The local-dispatch fallback (
kind='local_dispatch_fallback') is a degraded path, not a profile. Leave it alone. Underlocalitythe local model is the primary choice, so the fallback should simply never fire. - Circuit breaker and admin overrides already feed
exclude_models. Profiles must AND with them, never replace them: a profile must not resurrect a model an operator deprecated or the breaker has marked down. exploration(epsilon-greedy) picks from eligible candidates. Confirm it reads the profile-filtered set, or exploration will select models the profile excluded.
Non-goals
- No new ranking algorithm. Quality-first with cost as tiebreak stays.
- No change to
objective.quality_tolerance's global value. - No per-category preference logic (that was §1(a) of the previous plan and is explicitly out of scope here).
- No changes to the local-dispatch fallback path.
- Do not touch the classifier: profiles select candidates, not categories.
Success criteria
autoandauto:batchbehave exactly as today, pinned by tests that would catch a regression in the general mechanism.- At least the three named profiles ship, defined as predicates over catalog columns where possible.
- An unknown profile returns 422 naming the valid ones; it never silently
degrades to
default. - A profile that empties the candidate set returns a rejection reason naming the profile.
route_decisions.profileis populated and visible in the admin decisions table.- Profiles AND with circuit-breaker exclusions, admin overrides, and
eligible_categories— proven by a test for each, not by inspection. rejection_reasonremains the single copy of the eligibility rules.- Full suite green with
local_energy.enabledboth true and false (the Plan 5 invariant must not regress). - The user's
config/config.yamltariff lines remain uncommitted and verbatim.