Files
6krrt/docs/routing.md
2026-09-20 20:28:22 -04:00

643 lines
37 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
> Deep dive into routing internals. Back: [README](../README.md).
## Quality-First Selection
Seven hard filters are applied **before** scoring (not weighted — outright disqualification):
1. `effective_context_window ≥ required_context_tokens`
2. `tier ≥ required_tier` (from classifier)
3. Serving class compatible with request's `latency_tolerance`; `access_level` reachable
4. `tool_use_agentic` proficiency ≥ `routing.min_tool_proficiency`, **only when
the request carries a `tools` array** (ships disabled — `null`)
5. `supports_vision = 1` when the request carries `image_url` parts; `NULL`
fails closed (an unknown flag means the capability cannot be confirmed)
6. `supports_json_mode = 1` when `response_format.type` is `json_object` or
`json_schema`; `NULL` also fails closed
7. `eligible_categories` restricts only when set: a row with a non-empty
category list is admitted only when the task's category is in that list;
`NULL`/absent means unrestricted (all cloud rows today)
Filters 4–6 are read from the request body rather than inferred. A `tools`
array states whether tool definitions are on the table, `image_url` parts state
whether vision is needed, and `response_format` states whether JSON mode is
needed. Local classifiers identified unambiguous tool-use prompts only 1-2
times in 6, so inferring capability requirements does not work; the request
states them exactly and for free.
The tool gate is a **quality measurement** filter: a model with no measured
`tool_use_agentic` score is admitted, because "unproven" is not "proven bad".
Only a measured score below `routing.min_tool_proficiency` is dropped. The
vision and JSON-mode gates are **capability flag** filters. There every
routable catalog row has `supports_tools = 1`, so a flag gate would be inert;
vision and JSON mode are not universal, and an absent or `NULL` flag means
"cannot confirm the capability", so the model is dropped. A wrong guess on a
capability flag is a guaranteed provider 400.
The tool filter ships **off** (`null`), because agent clients send `tools` on
nearly every request, so enabling it excludes the cheapest model from ordinary
agent traffic. Whether that trade is worth it is an empirical question best
settled with `POST /outcome` data rather than a 3-task benchmark score. Set it
to `0.5` to turn it on. The vision and JSON-mode gates ship **on**, because a
wrong guess is a guaranteed 400.
### Local Vision Fallback
Cloud vision is not universal in the catalog, and the most economical coding
rows do not declare `supports_vision`. Routing an image request through them
would earn a provider-side 400, so when `routing.require_vision` is true only
rows with `supports_vision = 1` survive the hard filters. If **no** cloud
candidate survives, the router can fall back to a local vision model instead
of returning 422.
`local_vision:` in `config/config.yaml` controls this path:
| Key | Default | Purpose |
|---|---|---|
| `enabled` | `true` | Whether the fallback runs at all |
| `base_url` | `http://localhost:11434/v1` | OpenAI-compatible Ollama endpoint |
| `api_key_env` | `null` | Env var holding an API key, if the endpoint needs one |
| `model` | `qwen3-vl:4b` | Vision model on that Ollama |
| `timeout_seconds` | `60` | Request timeout |
| `max_images` | `4` | Refuse requests with more image parts |
| `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this |
The fallback is **enabled by default** both in `config/config.yaml` and in
`LocalVisionConfig`, so omitting the section still turns it on. Disable it
explicitly (`enabled: false`) on a host with no local Ollama or one that has
not pulled the vision model.
When enabled and no cloud candidate is selected, `_run_local_vision` sends the
original message list — with `image_url` parts intact — to the configured
Ollama model. The local answer then **replaces** the cloud completion:
`_local_vision_response` returns a normal OpenAI-shaped response, including a
stream-wrapped version when the client asked for `stream: true`. It does *not*
inject a caption into a cloud call, because the streaming proxy cannot rewrite
bytes mid-stream.
Security and budget guards:
- **Only inline `data:` URIs are accepted.** A remote `http(s)` image URL is
declined, because pointing a local model at an arbitrary URL would let an
unauthenticated caller make the router fetch internal resources (SSRF).
- Image count and total payload size are bounded by `max_images` and
`max_image_bytes` before the local call is made.
- If the local call fails for any reason, the request falls through to the
ordinary `422 No model satisfies the hard filters` rather than returning a
silent empty response.
Pull the model on whichever Ollama the fallback points at:
```bash
ollama pull qwen3-vl:4b
```
Config is strict (`extra="forbid"`): a misspelled or misplaced key fails at
load instead of being silently ignored.
### Local dispatch branch
Tasks in a local model's `eligible_categories` (for example `file_summarization`
and `diff_checking` for the shipped `qwen2.5-coder-router:14b`) survive the hard
filters the same way cloud rows do. Once selected, the request is sent to the
configured Ollama endpoint (`provider='ollama-local'`) instead of to
`dispatch_providers`. A local 502 trips the circuit breaker, so the next request
sees the local row as excluded and reroutes to a cloud candidate.
Local dispatch is **quality-gated and dormant under the default profile**.
The local row competes in the same ranking as every cloud row and must score
within `objective.quality_tolerance` of the category's cloud leader before its
price advantage is even consulted. Measured on 2026-09-02/03,
`qwen2.5-coder-router:14b` scores **0.767** on `file_summarization` against the
cloud leader's **0.95** — a 0.183 gap against a 0.10 tolerance — so under the
default profile the local row is dormant **by design**, today, until a local
model scores within tolerance. This is not a bug: quality is the objective,
and the local model has not yet earned the tiebreak (its 9.1x cost advantage
applies mainly to prompt tokens, so it wins most where large-context
summarization dominates — and it is assigned to summarization precisely
because it is not good enough at code).
Do NOT widen `objective.quality_tolerance` to fix this: the tolerance
describes measurement noise (bench sample sizes are small enough that a
smaller gap is genuinely indistinguishable), not preference. The project's
direction is named profiles (`auto:batch` is already one — a named widening
of the candidate set), and a `locality` profile is the right future home for
a per-category local preference, not a global tolerance change.
When every cloud candidate is exhausted, or the upstream returns an
account-level 4xx (any 4xx except 400/404/422 — those mean the request
itself is malformed and would fail locally too), an eligible routed
request degrades to the local dispatch model instead of surfacing the
cloud error. The trigger is deliberately conservative because the
upstream's exact credit-exhaustion signature is not yet observed
(full-body upstream-failure logging is in place to capture it); the
remaining cloud rows share one account, so no time is wasted walking
them after an account-level refusal. The fallback respects
`eligible_categories` even here: for everything else, the cloud error
fails loudly rather than handing a code task to a summarization-grade
model. The degraded answer is recorded as `kind='local_dispatch_fallback'`
in `route_decisions` so the proficiency loop never mistakes it for local
winning on merit, and if the local call itself fails the original cloud
error surfaces (the same fail-through contract as the local vision
fallback). There is no separate config knob — the fallback is armed
exactly by a `local_dispatch_models` entry listing the category in its
`eligible_categories`.
```
1. drop candidates whose measured energy exceeds objective.max_energy_per_request
2. rank by expected pass rate for the task's category (proficiency.blended_score)
3. treat differences smaller than objective.quality_tolerance as equal
4. among equals, pick the cheapest
5. on a small share of eligible requests, explore the least-evidenced candidate
```
| Setting | Value | Notes |
|---|---|---|
| `quality_tolerance` | 0.10 | % of real success rate, not an abstract quality unit: a candidate must beat the best by more than this to outrank a cheaper alternative |
| `assumed_cache_rate` | 0.917 | Share of prompt tokens served from the provider's prefix cache. Agent clients resend the conversation each turn, so most of it hits. Measured token-weighted over 40.7M tokens; measure your own against the provider's session view |
| `assumed_completion_tokens` | 500 | Completion length assumed when pricing a candidate |
| `max_energy_per_request` | null | Per-request kWh ceiling — a wall, not a bill. Null disables it |
| `plan_kwh_per_period` | 6.25 | **Set this to your own plan's quota.** Reported in `/health` as burn against the allowance; it does not gate anything |
This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4).
Measurement retired it: turning the cost weight from 0.4 to **zero** changed
the winner in only 2 of 6 categories, so the blend was never steering on
quality — while 60% of every decision adjudicated fractions of a cent.
### What the proficiency score means now
`proficiency.blended_score` is an **expected pass rate on your traffic**, not a
raw benchmark quality level. The benchmark (leaderboard + self-eval) provides a
prior. Client-reported outcomes from `POST /outcome` provide the per-model
signal. The two are combined through empirical-Bayes shrinkage:
- A category with no outcome traffic keeps the benchmark score verbatim.
- A trafficked category row with no own outcomes inherits a peer prior.
- A row with its own outcomes blends the observed rate with that prior;
`outcome_prior_strength` (default 20 pseudo-observations) pulls thin data
toward the category mean so one lucky sample cannot dominate.
So `quality_tolerance` asks "is this candidate's expected real success rate at
least 10% better than the cheaper one?" If not, cost breaks the tie. The score
is calibrated against the same client reports that `feedback.py` folds in, so
routing learns from traffic rather than from the benchmark alone.
**Cost is priced per request, from catalog token prices scaled to the
request's shape** — prompt size, `assumed_completion_tokens`,
`assumed_cache_rate`. It is deliberately *not* a benchmark average.
Neuralwatt bills flat per-kWh rather than per-token, so list price is not what
gets charged. Scoring used measured billed cost from a fixed 400-token
reference sweep for exactly that reason — and that was wrong for real traffic,
because the ranking depends on the workload's *shape*, not just the model. On
a 400-token prompt one model looked 3.2× cheaper than another; on a realistic
70k-token prompt the same pair inverted and the second was 5.0× cheaper. The
provider's attribution ratio moves with prompt size, so a fixed-shape
benchmark cannot rank models for a workload of a different shape.
List price is still not what is billed, but billing is capped at a multiple of
it, so it tracks the real ordering and bounds it — and it is free, needs no
sweep, and refreshes whenever the poller runs.
**Why cost ≠ eco:** Cost tracks energy (kWh), but carbon is energy × grid
intensity. Grid intensity spanned ~49 gCO2/kWh (`FI`) to ~442
(`US-MIDA-PJM`) when measured — and it moves with time of day. The models
disagree: `glm-5.2-fast` is
2nd cheapest but 6th cleanest; `kimi-k3-flex` draws 3.7× less energy than
`kimi-k2.7-code` while emitting 3.6× more carbon. Collapsing them picks a
side.
### Exploration
Routing runs an epsilon-greedy exploration pass after ranking. If enabled
(`exploration.enabled: true`, default), a random share `ε = 0.03` of eligible
requests replaces the winner with the hard-filter-eligible candidate that has
the fewest `outcome_samples`, tie-broken by lowest cost. The alternative is
skipped if its cost exceeds `max_cost_ratio ×` the winner's cost (default 4×).
Only tier 1 and tier 2 requests explore (`max_tier: 2`); tier 3 stays
exploitative so frontier work gets the best expected rate.
`was_exploration` is persisted to `route_decisions.exploration` so metrics can
tell real preference from forced discovery. This is the mechanism that breaks
the exposure-bias loop: without it, the router would send most requests to the
models already richest in outcome samples and the least-sampled rows would
never catch up.
### Incumbency and cache pricing
The ranking above treats every candidate as if its prompt were fresh. On a real
session it is not: the provider caches the conversation prefix, so the model
that served the last turn re-serves it with most of its prompt tokens billed at
the cached rate, while a different model pays full price for the entire prefix
on its first turn. `assumed_cache_rate` already prices that discount uniformly
for every row. Incumbency pricing makes it per-row: the incumbent keeps its
measured cache rate, challengers are priced toward cold, and a challenger now
has to beat the incumbent by more than the cache it is about to discard, which
on a long prompt is most of the prompt. This is the first place the
router's own past decision feeds back into its cost model, and it is
deliberately narrow: it changes only the cost key, never the sort.
The incumbent is **per conversation** when the client sends
`X-Router-Conversation`: `session_key` is then the namespaced conversation id,
so each conversation tracks its own incumbent. Without that header the
incumbent is per system-prompt fingerprint, and the fingerprint's hashing
merges concurrent same-prompt agents into one incumbent so they share the
cache discount.
Four knobs under `objective:`, all shipping **off**
(`incumbent_cache_pricing: false`, dial `null`; enable only after the Wave 1
post-restart baseline day, the Wave 2 gate in `plans/token-waste-waves.md`):
| Knob | Default | Purpose |
|---|---|---|
| `incumbent_cache_pricing` | `false` | The gate; off means byte-identical ranking (pinned by test) |
| `incumbent_challenger_cache_rate` | `null` | The challenger dial: `null` follows `assumed_cache_rate`; `0.0` is fully cold; between is a partial penalty |
| `incumbent_rate_refresh_seconds` | `300` | TTL on the measured per-(provider, model) rate table |
| `incumbent_rate_min_observations` | `25` | Observations before a measured rate is trusted **for pricing**; independent of the warning floor |
**What the incumbent is.** The model that served the session's last chat turn,
read back from `route_decisions` by `_session_incumbent_lookup` in
`dispatcher.py`: the most recent row for the session key with a selected model.
The query is an allowlist (`kind IN ('chat')`, never a `!=` denylist), so a
one-request `model` pin (passthrough), a `local_dispatch_fallback` row, or a
`route` probe can never set the incumbent, and neither can any future `kind`
until it argues its way in. The reason is that incumbency is a statement about
*preference*: this session, with its own history, chose that model
repeatedly on merit. A single pinned request says nothing about the session
and a degraded fallback is the opposite of a preference; letting either become
sticky would convert an escape hatch into a rut.
**The rate ladder.** When the feature is on and an incumbent is present, the
incumbent is priced at its measured per-(provider, model) cache rate whenever
that series has at least `incumbent_rate_min_observations` observations
(default 25); below that trust threshold it falls back to `assumed_cache_rate`.
The rates come from `metrics.cache_rate_series` through
`_measured_cache_rates()` in `dispatcher.py`, cached for
`incumbent_rate_refresh_seconds` (the TTL runs on `time.monotonic()` from the
first call after process start, so a brief post-restart cold period is
expected). The pricing floor is deliberately a different knob from the
alerting floor `cache_rate_warn_min_observations`: pricing would rather fall
back to the assumed rate on thin data than steer money on a thin measurement,
alerting wants the opposite trade, and coupling the two would couple operator
intents that should stay separate. Any failure in the measured-rate path fails
open to an empty table and every row prices at `assumed_cache_rate`;
incumbency pricing must never block routing.
**The challenger dial.** `objective.incumbent_challenger_cache_rate` is the
rate every non-incumbent row is priced at. Neutral is decided by **semantic
equality**, not identity: `null` follows `assumed_cache_rate`, and any dial
equal to `assumed_cache_rate` is the same neutral, so at a neutral setting
every row *including the incumbent* prices at `assumed_cache_rate` and the
ranking is byte-identical to today's, even with the flag on (config resolves
`null` to a concrete float at load; downstream code never sees both
representations). `0.0` prices challengers as fully cold prompts, the maximum
incumbent advantage; any value between is a partial cache penalty. The whole
feature tunes from off to full by moving this one config value, never by a
revert: if the switch gap does not narrow, the move is toward cold, then
re-measure.
**The clamp, and why it is load-bearing.** The challenger rate is
`min(dial, incumbent_rate)`, and the invariant it buys is: **the incumbent's
cache rate is always >= every challenger's cache rate.** Without it, the dial
is denominated in an absolute cache rate, so any dial above the incumbent's
measured rate prices the challenger as having a *better* cache than the
incumbent: the feature inverts and penalises the incumbent for being the
incumbent. Because several routed models measure below the assumed rate (a
model measured at ~0.73 against an assumed 0.917, for instance), that inverted
zone is not a corner case: for such a model it spans essentially the whole walk
from neutral down toward full penalty, which is exactly the range an operator
is told to tune through, and from inside it the counterfactual report looks
like "the penalty does not pay here" when in fact the penalty is running
backwards. Do not remove the clamp as redundant. An alternative dial
denominated as a relative penalty (`incumbent_rate * (1 - penalty)`) cannot
invert by construction and was deliberately deferred rather than adopted (see
the wave's second review, `wave2-review-2.md`); until that redesign exists,
the clamp is what makes the absolute form safe, and a property test pins it:
for every dial and every measured rate, the incumbent's priced rate is >=
every challenger's.
**How it composes.** Cache pricing happens inside `estimated_cost`: the
ranking loop simply passes each row a different `cache_rate` argument, so
every consumer of cost, from the band tiebreak to `cost_score` to the
flex-twin swap (a distinct serving endpoint, so it prices as a challenger at
the same dial), sees one coherent number per row. `provider_cost_multipliers`
composes multiplicatively on top of the cache-adjusted estimate, exactly as
before: it inflates the comparison cost inside the quality-band tiebreak and
can never override a genuine quality gap, because the quality band is
computed before the cost key and the incumbent gets no quality advantage. A
model that is genuinely better still wins outright.
**Eviction.** The incumbent is a pricing fact, not a reservation.
`_resolve_incumbent` re-runs the incumbent's catalog row through
`rejection_reason` with this request's actual filters and checks the profile's
`restrict_to`: if a hard filter would drop it (the session outgrew its
context window, a profile switch restricts the candidate set, a circuit is
open), it simply loses incumbency for that turn, no error, and the turn prices
exactly as it did before the feature: every row at `assumed_cache_rate`.
**Exploration becomes session-scoped.** The coin described above now flips
only when no incumbent is present: `route()` adds `incumbent is None` to the
exploration condition, and `exploration.py` itself is unchanged (injected RNG,
no mutable state, per its module contract). The coin flips at session start
only, so one exploratory session costs one cold prompt instead of one per
turn; incumbency pricing then pins the explored model for the session's
remaining turns, which is tiebreak protection without new state. Re-exploring
per turn with an incumbent present would dump the cache every few turns, the
exact leak the feature exists to stop (`plans/token-waste-waves.md` item 2.3).
**Expiry checks.** Both premises carry their own. The `assumed_cache_rate`
fallback is checked by `cache_rate_warnings` in `/metrics`, which compares the
measured aggregate and per-(provider, model) rates against the assumed
constant and warns beyond `cache_rate_warn_margin`. A deployment whose real
rate has diverged is mispriced, not broken, and the fix is to re-measure. The
dial's own premise, "the penalty is paying," is evaluated offline rather than
assumed: `baseline_report.py --incumbent-challenger <rate>` replays recent
`route_decisions` against a cold-challenger ranking and reports what the
penalty changed and what other dial settings would have changed, and the
operator re-measures the same-model/switched cache-rate gap and billed µ$ per
prompt token each re-measurement window after enabling. When the penalty
stops paying, the named move is toward neutral on the dial. No hardcoded
expiry duration: the check is the measurement.
**Shipping state.** Off, as shipped: behavior is byte-identical to
pre-feature traffic (pinned by test). When on, every switch decision is
explainable post-hoc: the rank debug log carries the incumbent identity, its
rate and source (measured or assumed), and the dial; the incumbent row itself
carries an `incumbent_pricing` stamp in the ranked output.
### What routing actually returns, and why it moves
Sweeping `/route` across every category and tier is the fastest way to see
whether your data is doing anything. On the deployment this was written
against, 9 categories × 3 tiers currently yields **5 distinct winners**
(`qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `kimi-k3`,
`kimi-k3-fast`).
That number is a diagnostic, not a target, and it is worth knowing what each
outcome means:
- **One winner everywhere** is a legitimate answer, not a misconfiguration.
It happened here: with cost, eco and proficiency all populated, one model
was Pareto-dominant — cheapest *and* cleanest in the routable set while
scoring within `quality_tolerance` of the best. No defensible weighting
picks anything else. If you see this, check whether the leader really is
dominant before reaching for the config.
- **Winners that change with context size** are the hard filters working.
A model is dropped once `required_context_tokens` exceeds its window, so a
long session can change model mid-conversation. Past the largest window,
`/v1/chat/completions` returns 422 naming the constraint rather than
silently truncating.
- **Winners that change by category** mean proficiency is live. That is the
only category-dependent term, so until the `proficiency` table has data,
`task_category` cannot change a decision at all — the classifier computes
it and the router pays for it for nothing. The score is now
outcome-calibrated, so category-level client reports are the fastest way to
shift these choices.
The spread here widened for two reasons worth copying: cost became a
per-request estimate rather than a benchmark average, and tier stopped being
inferred from price alone. Both had been quietly excluding a cheap
large-context model from every request above tier 1. Outcome samples are now
arriving too; `deepseek-v4-flash` gained substantial samples in
`coding_general`, `coding_refactor`, and `general_chat` from post-stream
`POST /outcome` reports.
If you want a different balance, the levers are `objective.quality_tolerance`
(how big a quality gap must be before it outranks cost) and
`objective.max_energy_per_request` (a hard ceiling). There is no weight to
tune — quality is the objective and cost is the tiebreak, which replaced an
earlier weighted blend.
### Circuit Breaker — `circuit_breaker.py`
A seventh, dynamic filter sits alongside the six hard filters above: a model
that 5xx'd recently is passively excluded from the candidate set. On by
default (`circuit_breaker.enabled: true`).
It records nothing until a model actually fails — a 5xx from the upstream
call marks `(model_id, provider)` down for `initial_cooldown_seconds` (30s
default), doubling on each further failure (`backoff_multiplier`, capped at
`max_cooldown_seconds`, 600s) and clearing on the next success. Recovery is
**passive by design** — no background poller, no health-check loop. The next
real request that would have picked the down model becomes its own recovery
probe once the cooldown has passed.
Two call sites: candidate selection excludes down models outright, and the
dispatch retry loop (see the retry budget in [verification.md](verification.md))
records the failure/success on every upstream call and fails over to the
next-ranked candidate on a 5xx — but only for `auto`-routed requests, since a
pinned request has no alternative to fail over to.
`eval_proficiency.py` deliberately calls providers directly, bypassing both
the circuit breaker and the retry loop — a transient eval-harness failure
against one model's edges shouldn't be able to trip the breaker against real
production traffic.
### The classifier's own circuit breaker
Separate machinery, same idea, different table: `circuit_breaker.py` guards
*upstream models*, while `_last_classifier_failure` in `dispatcher.py` guards
the **local classifier** — the one blocking LLM call on the request path.
For its first release that timestamp was written by `_record_failure()` and
read by nothing: `_classify_cascade` gated only its *cloud* step, and on a
different timestamp (`_last_account_refusal`). So the router re-dialled a
known-dead local classifier on every request. A *stopped* Ollama refuses the
connection immediately and costs little; a **hung** one — or a VPN-bound one
that black-holes — costs the full `classifier.timeout_seconds`, 120s on this
deployment, per request for as long as the outage lasts.
`_classifier_backoff_active()` is the read, consulted in `classify()`
**before** `_classifier_client()` is constructed, because constructing it and
waiting out the timeout is the expensive part.
One detail is load-bearing and easy to undo by accident: recording the failure
lives in `_classify_via_local_llm`'s own exception handlers, **not** in
`_classify_cascade`. The cascade is walked for reasons other than a fresh
failure — a skipped attempt inside an already-open window, or gaming mode —
and if those re-stamped the clock then every request during an outage would
push the deadline forward and the local classifier would never be re-probed
while traffic flowed. That is a permanent outage wearing a circuit breaker's
clothes. `tests/test_classifier_backoff.py` pins it directly, and asserts on
whether the client was *constructed* rather than on the value returned — a
test that only checked the result would pass even if the router had waited
out the timeout first.
A skip and a real failure both reach the classifier's fallback cascade, but
they are logically distinct and `dispatcher.py` keeps them that way: a skip
raises `_ClassifierSkipped` (logged once, specifically, as
`classify_local_skipped`) while a real failure raises the underlying
exception (logged as `fallback`). Conflating the two would tell an operator
reading raw logs that the local classifier is failing when gaming mode simply
turned it off.
### Which classifier is PRIMARY — `classifier.mode`
Separate machinery from the circuit breaker above: that guards *upstream
dispatch models* once a category has already been classified. This guards
*which implementation does the classifying in the first place*.
`classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks
the primary attempt. It is a peer to the classifier's own fallback cascade
(local → stale session cache → session history → optional `cloud_fallback`
→ static guess), not a replacement for it — whichever mode is primary, a
failure in `_classify_via_configured_mode` still funnels into the exact
same, unmodified cascade in `dispatcher.py`.
Two things are load-bearing here and easy to get backwards in a future edit:
- **`cloud_llm` success records `source="classifier"`, not
`"classifier_cloud"`.** The latter string means specifically "the
cascade's post-failure backup step fired," and both the `/metrics`
degradation-share warning and the outcome-attribution set read it as a
*degraded* signal. An intentionally configured cloud primary succeeding
is the opposite of degraded, so reusing `"classifier_cloud"` for it would
make healthy, on-purpose traffic look like an ongoing local outage.
- **`cloud_primary_auto` resolves live, cached, not per-request.**
`routing.cheapest_classifier_candidate` reuses `select_candidates` +
`estimated_cost` — the same functions real dispatch ranking uses — priced
for the classifier's own short-prompt/short-completion call shape rather
than the task's. `dispatcher._resolve_auto_classifier` caches the result
for `classifier.cooldown_seconds` so this does not add a DB scan to every
request's latency floor; the admin portal's GET endpoint calls the
resolver fresh on every load instead, since a human loading a settings
page is not on that latency floor and a stale "live" reading would be
exactly the kind of bug the profiles page's zero-admit badge was.
`local_encoder` mode only ever produces `task_category` — `task_tier` falls
back to `classifier.fallback_tier`, a documented limitation rather than a
second heuristic. See [local-models.md](local-models.md) for why zero-shot
rather than fine-tuned (this router never stores raw task text).
### Checking whether scoring earns its complexity
`baseline_report.py` automates the "check whether the leader really is
dominant" step above: it replays recent `route_decisions` against two trivial
counterfactuals (always cheapest, always highest proficiency) using the
*current* catalog and proficiency table, and reports how often real scoring
picked something a trivial baseline wouldn't have.
```bash
PYTHONPATH=src python -m baseline_report --since 2026-08-01
PYTHONPATH=src python -m baseline_report --since 2026-08-01 --category coding_refactor --csv
```
A high dominance share paired with a near-zero proficiency delta against
`always_cheapest` means the quality-first ranking isn't earning its complexity
for that slice of traffic. Read-only — adds no schema, spends no quota.
### Tiering
Tier on `reasoning_default_enabled` (from `metadata.reasoning.default_enabled`,
falling back to `capabilities.reasoning`), **not** `supports_reasoning`.
`capabilities.reasoning` only means "the endpoint accepts a reasoning
param" — it is true for 17 of 19 rows, and tiering on it put 17 models in
tier 3 and left tier 1 empty. A `-fast` row does not inherit its sibling's
tier 3. Cost is checked before the reasoning rule so $0.28/1M models can
reach tier 1.
**Cheapness is not a capability ceiling** (`tier1_context_max`, default
512000). Tier is a *floor* — `routing.py` drops any row with
`tier < required_tier` — so tier 1 means "simple work only", not "cheap".
Deciding that on price alone put `deepseek-v4-flash` in tier 1 for no reason
but its $0.28/1M completion price, which excluded it **outright** from every
tier-2 request. It has a 1M advertised window and scores 1.00 on all three
coding categories. That was the same substitution the cost axis already had
to unlearn: price is a market signal, not a capability measurement.
Tier 1 now requires the model to be small **and** cheap. The gate reads the
**advertised** `context_window`, whose catalog values are the clean market
classes — 131056 / 199984 / 262128 / 1048560 — rather than
`effective_context_window`, which varies within a class. 512000 sits in the
empty band between the 256K and 1M classes with a 2x margin either side, so it
is not fitted to any one model. The gate only ever demotes; a huge window
never promotes an expensive model into tier 1, and a *missing* window does not
block it, since absent evidence should not decide anything.
Distribution moved **4 / 6 / 9 -> 1 / 9 / 9** — only the three deepseek rows
changed. With cost priced per-request, `deepseek-v4-flash` now wins
`coding_general` at every context size (16.5x cheaper than `kimi-k3` at 200k)
and is still correctly absent from `tool_use_agentic`, where its measured 0.33
drops it out of the quality band. That is the eval data earning it the slot
rather than a thumb on the scale — no `model_tiers` override was needed.
### Why tools and reasoning stay on their existing signals
Tools stay on the measured `tool_use_agentic` proficiency gate, not a
`supports_tools` flag gate. Every routable catalog row already has
`supports_tools = 1`, so a flag gate would be inert. The real signal is the
measured proficiency, because the observed failure is a model over-reaching for
tools on a non-agentic prompt. `routing.min_tool_proficiency` captures that
measurement and only applies when the request carries a `tools` array.
Reasoning stays on tiering (`reasoning_default_enabled`), not on a new flag gate.
`supports_reasoning` only means the endpoint accepts a reasoning parameter, and
that is true for 17 of 19 rows — nearly the whole catalog. `has_reasoning_request`
is detected purely for observation. Making it a gate would add no useful
filtering, because the decision of whether a request needs reasoning is already
encoded in the requested tier.
The fail-closed asymmetry, stated plainly: **capability flags fail closed on
unknown; quality measurements admit on absent evidence.** A missing
`supports_vision` or `supports_json_mode` flag means "cannot confirm", so the
model is dropped. A missing `tool_use_agentic` score or energy measurement means
"unproven, not bad", so the model is admitted. The first wrong guess is a
guaranteed 400; the second is just an empty data point that the neutral default
handles.
### The classifier's labels are not the proficiency scoring axis
`classifier.candidate_categories` is the set of labels a classifier may
return. `proficiency.categories` is the axis models are scored on. They were
one list, and the two jobs are not the same job: the axis answers "what is
this model good at", the candidate set answers "what should a classifier be
asked to distinguish". Config load validates the candidate set as a **subset**
of the axis — a label outside it joins against nothing in the `proficiency`
table, so routing would rank on NULLs, which is the same phantom-join failure
`routing.tool_use_category` is validated against, arriving from the other
side. Omitting the key means "every category", i.e. the pre-split behaviour.
Both classifier backends read the candidate set: `dispatcher.classify` builds
its allowed-values prompt from it, and `local_encoder.classify_zero_shot`
receives it as its candidate labels. `_classify_once` also validates the
returned label against it, so the exclusion is binding rather than advisory —
a local model that ignores the prompt's list is corrected, exactly as an
invented label always was. The scoring-axis consumers were deliberately left
on the full list: `leaderboard.py` (priors are per scoring category),
`eval_proficiency.py` (the benchmark still evaluates every category),
`admin.py`'s profile-coverage count and `/metrics`' category list (both
describe the scoring axis), and `config.py`'s `tool_use_category` /
`eligible_categories` validators.
**`tool_use_agentic` is the excluded category, and why is the point.** It
describes what a turn mechanically *does* rather than what it is *for*, and
every agent turn does it. Measured 2026-09-14 on live traffic with the session
cache off so every turn classified for real: **31 of 31 consecutive turns**
classified `tool_use_agentic`, all carrying `tools`, all routed to
`z-ai/glm-5.3-flash` — 14.8s time-to-first-token, p95 40s, the slowest model in
the catalog. Per-turn classification became accurate and routing got worse.
Classifying *for* tool use also duplicates a signal the router already has
exactly: the request carries a `tools` array, reading it is free, and
`routing.min_tool_proficiency` is the mechanism that acts on it.
**What the exclusion costs, written down so the next person does not have to
rediscover it:** no new `POST /outcome` report can attribute to
`tool_use_agentic`, because nothing classifies as it any more. Those scores
**freeze** at their current values (`qwen3.6-35b` at 0.902 over 154 samples).
The tool filter keeps reading them and the category keeps ranking; it simply
stops accumulating. That is accepted for now, not fixed.
**Considered and rejected:** attributing outcomes to `tool_use_agentic`
whenever the request carried a `tools` array. opencode sends `tools` on
essentially every request, so the score would converge on each model's overall
pass rate and stop discriminating the one thing the category exists to
detect — a model that *over-reaches* for tools on work that did not need them
(`deepseek-v4-flash` calling two tools to subtract 1:20pm from 3pm). A frozen
honest score beats a live meaningless one.
The readout is on the admin portal's Classifier card: the candidate list, and
the excluded categories with a note that they stop accumulating outcomes.
Read-only on purpose — it is a structural list in the same class as
`proficiency.categories` and `models.eligible_categories`, neither of which
has an editing surface, and what an operator needs from the portal is the
answer to "why does nothing ever classify as X".