cfg.proficiency.categories was one list doing two jobs: the candidate labels a classifier is asked to choose between, and the axis models are scored on. The jobs are not the same job, and the mismatch was not hypothetical: measured 2026-09-14 with the session cache off so every turn classified for real, 31 of 31 consecutive live turns classified tool_use_agentic and every one routed to the slowest model in the catalog (14.8s TTFT, p95 40s). The label names what a turn mechanically DOES, and every agent turn does it, so per-turn classification carried no other signal at all. What the category was standing in for is already read exactly and for free: the request's own tools array, which routing.min_tool_proficiency acts on. So classifier.candidate_categories is the label set, and proficiency.categories stays the scoring axis. Config load validates the candidate set as a SUBSET of the axis -- the same phantom-join guard routing.tool_use_category has, arriving from the other side -- and omitting the key means "every category", the pre-split behaviour. tool_use_agentic is excluded from the labels and STAYS on the axis: routing.tool_use_category still validates against the full list and the tool filter keeps its data. The split narrows the classifier; it does not touch the filter. The exclusion is binding rather than advisory. classify() builds the allowed-values prompt from the new list, local_encoder.classify_zero_shot receives it as its candidates, and _classify_once corrects a returned-but-excluded label exactly as it always corrected an invented one -- without that, the split would be a suggestion a non-compliant local model could route around, straight into route_decisions and outcome attribution on a category the classifier is supposed to have stopped emitting. The cost is stated where the next person will look for it. With the classifier unable to emit tool_use_agentic, no new POST /outcome attributes to the category, so its scores freeze at today's values (qwen3.6-35b 0.902 over 154 samples) and the filter reads frozen-but-real data. Considered and REJECTED as the fix: attributing outcomes to the category whenever the request carried a tools array. opencode sends tools on essentially every turn, so the score would converge on each model's overall pass rate and stop discriminating the one failure it exists to catch -- deepseek-v4-flash calling two tools to subtract 1:20pm from 3pm on a prompt that was not agentic at all. A frozen honest score beats a live meaningless one. The full reasoning is in config.yaml's comment block and docs/routing.md. The readout lives on the Classifier card: the candidate labels, and the excluded half with why it stopped accumulating. Read-only on purpose -- candidate_categories is a structural list in the same class as proficiency.categories and models.eligible_categories, neither of which has an editing surface, and the knob-coverage tripwire scopes to scalar leaves. What an operator needs from the portal is the answer to "why does nothing ever classify as X", which was invisible before this row existed. CLAUDE.md's min_tool_proficiency experiment note is corrected rather than left to send someone down a path that can no longer produce data. Verified: 2100 passed offline. New test_classifier_candidate_categories.py pins the shipped-config exclusion as ONLY tool_use_agentic, the subset and empty-set load refusals, the omit-means-all default, both backends reading the narrowed list, and the binding correction; the frontend readout is pinned in test_admin_frontend.py. Live on a throwaway 8081: the readout renders the ten labels and the excluded note, wraps cleanly at 1400 and 800, console clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
37 KiB
Deep dive into routing internals. Back: README.
Quality-First Selection
Seven hard filters are applied before scoring (not weighted — outright disqualification):
effective_context_window ≥ required_context_tokenstier ≥ required_tier(from classifier)- Serving class compatible with request's
latency_tolerance;access_levelreachable tool_use_agenticproficiency ≥routing.min_tool_proficiency, only when the request carries atoolsarray (ships disabled —null)supports_vision = 1when the request carriesimage_urlparts;NULLfails closed (an unknown flag means the capability cannot be confirmed)supports_json_mode = 1whenresponse_format.typeisjson_objectorjson_schema;NULLalso fails closedeligible_categoriesrestricts only when set: a row with a non-empty category list is admitted only when the task's category is in that list;NULL/absent means unrestricted (all cloud rows today)
Filters 4–6 are read from the request body rather than inferred. A tools
array states whether tool definitions are on the table, image_url parts state
whether vision is needed, and response_format states whether JSON mode is
needed. Local classifiers identified unambiguous tool-use prompts only 1-2
times in 6, so inferring capability requirements does not work; the request
states them exactly and for free.
The tool gate is a quality measurement filter: a model with no measured
tool_use_agentic score is admitted, because "unproven" is not "proven bad".
Only a measured score below routing.min_tool_proficiency is dropped. The
vision and JSON-mode gates are capability flag filters. There every
routable catalog row has supports_tools = 1, so a flag gate would be inert;
vision and JSON mode are not universal, and an absent or NULL flag means
"cannot confirm the capability", so the model is dropped. A wrong guess on a
capability flag is a guaranteed provider 400.
The tool filter ships off (null), because agent clients send tools on
nearly every request, so enabling it excludes the cheapest model from ordinary
agent traffic. Whether that trade is worth it is an empirical question best
settled with POST /outcome data rather than a 3-task benchmark score. Set it
to 0.5 to turn it on. The vision and JSON-mode gates ship on, because a
wrong guess is a guaranteed 400.
Local Vision Fallback
Cloud vision is not universal in the catalog, and the most economical coding
rows do not declare supports_vision. Routing an image request through them
would earn a provider-side 400, so when routing.require_vision is true only
rows with supports_vision = 1 survive the hard filters. If no cloud
candidate survives, the router can fall back to a local vision model instead
of returning 422.
local_vision: in config/config.yaml controls this path:
| Key | Default | Purpose |
|---|---|---|
enabled |
true |
Whether the fallback runs at all |
base_url |
http://localhost:11434/v1 |
OpenAI-compatible Ollama endpoint |
api_key_env |
null |
Env var holding an API key, if the endpoint needs one |
model |
qwen3-vl:4b |
Vision model on that Ollama |
timeout_seconds |
60 |
Request timeout |
max_images |
4 |
Refuse requests with more image parts |
max_image_bytes |
9437184 (9 MiB) |
Refuse requests whose image payload exceeds this |
The fallback is enabled by default both in config/config.yaml and in
LocalVisionConfig, so omitting the section still turns it on. Disable it
explicitly (enabled: false) on a host with no local Ollama or one that has
not pulled the vision model.
When enabled and no cloud candidate is selected, _run_local_vision sends the
original message list — with image_url parts intact — to the configured
Ollama model. The local answer then replaces the cloud completion:
_local_vision_response returns a normal OpenAI-shaped response, including a
stream-wrapped version when the client asked for stream: true. It does not
inject a caption into a cloud call, because the streaming proxy cannot rewrite
bytes mid-stream.
Security and budget guards:
- Only inline
data:URIs are accepted. A remotehttp(s)image URL is declined, because pointing a local model at an arbitrary URL would let an unauthenticated caller make the router fetch internal resources (SSRF). - Image count and total payload size are bounded by
max_imagesandmax_image_bytesbefore the local call is made. - If the local call fails for any reason, the request falls through to the
ordinary
422 No model satisfies the hard filtersrather than returning a silent empty response.
Pull the model on whichever Ollama the fallback points at:
ollama pull qwen3-vl:4b
Config is strict (extra="forbid"): a misspelled or misplaced key fails at
load instead of being silently ignored.
Local dispatch branch
Tasks in a local model's eligible_categories (for example file_summarization
and diff_checking for the shipped qwen2.5-coder-router:14b) survive the hard
filters the same way cloud rows do. Once selected, the request is sent to the
configured Ollama endpoint (provider='ollama-local') instead of to
dispatch_providers. A local 502 trips the circuit breaker, so the next request
sees the local row as excluded and reroutes to a cloud candidate.
Local dispatch is quality-gated and dormant under the default profile.
The local row competes in the same ranking as every cloud row and must score
within objective.quality_tolerance of the category's cloud leader before its
price advantage is even consulted. Measured on 2026-09-02/03,
qwen2.5-coder-router:14b scores 0.767 on file_summarization against the
cloud leader's 0.95 — a 0.183 gap against a 0.10 tolerance — so under the
default profile the local row is dormant by design, today, until a local
model scores within tolerance. This is not a bug: quality is the objective,
and the local model has not yet earned the tiebreak (its 9.1x cost advantage
applies mainly to prompt tokens, so it wins most where large-context
summarization dominates — and it is assigned to summarization precisely
because it is not good enough at code).
Do NOT widen objective.quality_tolerance to fix this: the tolerance
describes measurement noise (bench sample sizes are small enough that a
smaller gap is genuinely indistinguishable), not preference. The project's
direction is named profiles (auto:batch is already one — a named widening
of the candidate set), and a locality profile is the right future home for
a per-category local preference, not a global tolerance change.
When every cloud candidate is exhausted, or the upstream returns an
account-level 4xx (any 4xx except 400/404/422 — those mean the request
itself is malformed and would fail locally too), an eligible routed
request degrades to the local dispatch model instead of surfacing the
cloud error. The trigger is deliberately conservative because the
upstream's exact credit-exhaustion signature is not yet observed
(full-body upstream-failure logging is in place to capture it); the
remaining cloud rows share one account, so no time is wasted walking
them after an account-level refusal. The fallback respects
eligible_categories even here: for everything else, the cloud error
fails loudly rather than handing a code task to a summarization-grade
model. The degraded answer is recorded as kind='local_dispatch_fallback'
in route_decisions so the proficiency loop never mistakes it for local
winning on merit, and if the local call itself fails the original cloud
error surfaces (the same fail-through contract as the local vision
fallback). There is no separate config knob — the fallback is armed
exactly by a local_dispatch_models entry listing the category in its
eligible_categories.
1. drop candidates whose measured energy exceeds objective.max_energy_per_request
2. rank by expected pass rate for the task's category (proficiency.blended_score)
3. treat differences smaller than objective.quality_tolerance as equal
4. among equals, pick the cheapest
5. on a small share of eligible requests, explore the least-evidenced candidate
| Setting | Value | Notes |
|---|---|---|
quality_tolerance |
0.10 | % of real success rate, not an abstract quality unit: a candidate must beat the best by more than this to outrank a cheaper alternative |
assumed_cache_rate |
0.917 | Share of prompt tokens served from the provider's prefix cache. Agent clients resend the conversation each turn, so most of it hits. Measured token-weighted over 40.7M tokens; measure your own against the provider's session view |
assumed_completion_tokens |
500 | Completion length assumed when pricing a candidate |
max_energy_per_request |
null | Per-request kWh ceiling — a wall, not a bill. Null disables it |
plan_kwh_per_period |
6.25 | Set this to your own plan's quota. Reported in /health as burn against the allowance; it does not gate anything |
This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4). Measurement retired it: turning the cost weight from 0.4 to zero changed the winner in only 2 of 6 categories, so the blend was never steering on quality — while 60% of every decision adjudicated fractions of a cent.
What the proficiency score means now
proficiency.blended_score is an expected pass rate on your traffic, not a
raw benchmark quality level. The benchmark (leaderboard + self-eval) provides a
prior. Client-reported outcomes from POST /outcome provide the per-model
signal. The two are combined through empirical-Bayes shrinkage:
- A category with no outcome traffic keeps the benchmark score verbatim.
- A trafficked category row with no own outcomes inherits a peer prior.
- A row with its own outcomes blends the observed rate with that prior;
outcome_prior_strength(default 20 pseudo-observations) pulls thin data toward the category mean so one lucky sample cannot dominate.
So quality_tolerance asks "is this candidate's expected real success rate at
least 10% better than the cheaper one?" If not, cost breaks the tie. The score
is calibrated against the same client reports that feedback.py folds in, so
routing learns from traffic rather than from the benchmark alone.
Cost is priced per request, from catalog token prices scaled to the
request's shape — prompt size, assumed_completion_tokens,
assumed_cache_rate. It is deliberately not a benchmark average.
Neuralwatt bills flat per-kWh rather than per-token, so list price is not what gets charged. Scoring used measured billed cost from a fixed 400-token reference sweep for exactly that reason — and that was wrong for real traffic, because the ranking depends on the workload's shape, not just the model. On a 400-token prompt one model looked 3.2× cheaper than another; on a realistic 70k-token prompt the same pair inverted and the second was 5.0× cheaper. The provider's attribution ratio moves with prompt size, so a fixed-shape benchmark cannot rank models for a workload of a different shape.
List price is still not what is billed, but billing is capped at a multiple of it, so it tracks the real ordering and bounds it — and it is free, needs no sweep, and refreshes whenever the poller runs.
Why cost ≠ eco: Cost tracks energy (kWh), but carbon is energy × grid
intensity. Grid intensity spanned ~49 gCO2/kWh (FI) to ~442
(US-MIDA-PJM) when measured — and it moves with time of day. The models
disagree: glm-5.2-fast is
2nd cheapest but 6th cleanest; kimi-k3-flex draws 3.7× less energy than
kimi-k2.7-code while emitting 3.6× more carbon. Collapsing them picks a
side.
Exploration
Routing runs an epsilon-greedy exploration pass after ranking. If enabled
(exploration.enabled: true, default), a random share ε = 0.03 of eligible
requests replaces the winner with the hard-filter-eligible candidate that has
the fewest outcome_samples, tie-broken by lowest cost. The alternative is
skipped if its cost exceeds max_cost_ratio × the winner's cost (default 4×).
Only tier 1 and tier 2 requests explore (max_tier: 2); tier 3 stays
exploitative so frontier work gets the best expected rate.
was_exploration is persisted to route_decisions.exploration so metrics can
tell real preference from forced discovery. This is the mechanism that breaks
the exposure-bias loop: without it, the router would send most requests to the
models already richest in outcome samples and the least-sampled rows would
never catch up.
Incumbency and cache pricing
The ranking above treats every candidate as if its prompt were fresh. On a real
session it is not: the provider caches the conversation prefix, so the model
that served the last turn re-serves it with most of its prompt tokens billed at
the cached rate, while a different model pays full price for the entire prefix
on its first turn. assumed_cache_rate already prices that discount uniformly
for every row. Incumbency pricing makes it per-row: the incumbent keeps its
measured cache rate, challengers are priced toward cold, and a challenger now
has to beat the incumbent by more than the cache it is about to discard, which
on a long prompt is most of the prompt. This is the first place the
router's own past decision feeds back into its cost model, and it is
deliberately narrow: it changes only the cost key, never the sort.
Four knobs under objective:, all shipping off
(incumbent_cache_pricing: false, dial null; enable only after the Wave 1
post-restart baseline day, the Wave 2 gate in plans/token-waste-waves.md):
| Knob | Default | Purpose |
|---|---|---|
incumbent_cache_pricing |
false |
The gate; off means byte-identical ranking (pinned by test) |
incumbent_challenger_cache_rate |
null |
The challenger dial: null follows assumed_cache_rate; 0.0 is fully cold; between is a partial penalty |
incumbent_rate_refresh_seconds |
300 |
TTL on the measured per-(provider, model) rate table |
incumbent_rate_min_observations |
25 |
Observations before a measured rate is trusted for pricing; independent of the warning floor |
What the incumbent is. The model that served the session's last chat turn,
read back from route_decisions by _session_incumbent_lookup in
dispatcher.py: the most recent row for the session key with a selected model.
The query is an allowlist (kind IN ('chat'), never a != denylist), so a
one-request model pin (passthrough), a local_dispatch_fallback row, or a
route probe can never set the incumbent, and neither can any future kind
until it argues its way in. The reason is that incumbency is a statement about
preference: this session, with its own history, chose that model
repeatedly on merit. A single pinned request says nothing about the session
and a degraded fallback is the opposite of a preference; letting either become
sticky would convert an escape hatch into a rut.
The rate ladder. When the feature is on and an incumbent is present, the
incumbent is priced at its measured per-(provider, model) cache rate whenever
that series has at least incumbent_rate_min_observations observations
(default 25); below that trust threshold it falls back to assumed_cache_rate.
The rates come from metrics.cache_rate_series through
_measured_cache_rates() in dispatcher.py, cached for
incumbent_rate_refresh_seconds (the TTL runs on time.monotonic() from the
first call after process start, so a brief post-restart cold period is
expected). The pricing floor is deliberately a different knob from the
alerting floor cache_rate_warn_min_observations: pricing would rather fall
back to the assumed rate on thin data than steer money on a thin measurement,
alerting wants the opposite trade, and coupling the two would couple operator
intents that should stay separate. Any failure in the measured-rate path fails
open to an empty table and every row prices at assumed_cache_rate;
incumbency pricing must never block routing.
The challenger dial. objective.incumbent_challenger_cache_rate is the
rate every non-incumbent row is priced at. Neutral is decided by semantic
equality, not identity: null follows assumed_cache_rate, and any dial
equal to assumed_cache_rate is the same neutral, so at a neutral setting
every row including the incumbent prices at assumed_cache_rate and the
ranking is byte-identical to today's, even with the flag on (config resolves
null to a concrete float at load; downstream code never sees both
representations). 0.0 prices challengers as fully cold prompts, the maximum
incumbent advantage; any value between is a partial cache penalty. The whole
feature tunes from off to full by moving this one config value, never by a
revert: if the switch gap does not narrow, the move is toward cold, then
re-measure.
The clamp, and why it is load-bearing. The challenger rate is
min(dial, incumbent_rate), and the invariant it buys is: the incumbent's
cache rate is always >= every challenger's cache rate. Without it, the dial
is denominated in an absolute cache rate, so any dial above the incumbent's
measured rate prices the challenger as having a better cache than the
incumbent: the feature inverts and penalises the incumbent for being the
incumbent. Because several routed models measure below the assumed rate (a
model measured at ~0.73 against an assumed 0.917, for instance), that inverted
zone is not a corner case: for such a model it spans essentially the whole walk
from neutral down toward full penalty, which is exactly the range an operator
is told to tune through, and from inside it the counterfactual report looks
like "the penalty does not pay here" when in fact the penalty is running
backwards. Do not remove the clamp as redundant. An alternative dial
denominated as a relative penalty (incumbent_rate * (1 - penalty)) cannot
invert by construction and was deliberately deferred rather than adopted (see
the wave's second review, wave2-review-2.md); until that redesign exists,
the clamp is what makes the absolute form safe, and a property test pins it:
for every dial and every measured rate, the incumbent's priced rate is >=
every challenger's.
How it composes. Cache pricing happens inside estimated_cost: the
ranking loop simply passes each row a different cache_rate argument, so
every consumer of cost, from the band tiebreak to cost_score to the
flex-twin swap (a distinct serving endpoint, so it prices as a challenger at
the same dial), sees one coherent number per row. provider_cost_multipliers
composes multiplicatively on top of the cache-adjusted estimate, exactly as
before: it inflates the comparison cost inside the quality-band tiebreak and
can never override a genuine quality gap, because the quality band is
computed before the cost key and the incumbent gets no quality advantage. A
model that is genuinely better still wins outright.
Eviction. The incumbent is a pricing fact, not a reservation.
_resolve_incumbent re-runs the incumbent's catalog row through
rejection_reason with this request's actual filters and checks the profile's
restrict_to: if a hard filter would drop it (the session outgrew its
context window, a profile switch restricts the candidate set, a circuit is
open), it simply loses incumbency for that turn, no error, and the turn prices
exactly as it did before the feature: every row at assumed_cache_rate.
Exploration becomes session-scoped. The coin described above now flips
only when no incumbent is present: route() adds incumbent is None to the
exploration condition, and exploration.py itself is unchanged (injected RNG,
no mutable state, per its module contract). The coin flips at session start
only, so one exploratory session costs one cold prompt instead of one per
turn; incumbency pricing then pins the explored model for the session's
remaining turns, which is tiebreak protection without new state. Re-exploring
per turn with an incumbent present would dump the cache every few turns, the
exact leak the feature exists to stop (plans/token-waste-waves.md item 2.3).
Expiry checks. Both premises carry their own. The assumed_cache_rate
fallback is checked by cache_rate_warnings in /metrics, which compares the
measured aggregate and per-(provider, model) rates against the assumed
constant and warns beyond cache_rate_warn_margin. A deployment whose real
rate has diverged is mispriced, not broken, and the fix is to re-measure. The
dial's own premise, "the penalty is paying," is evaluated offline rather than
assumed: baseline_report.py --incumbent-challenger <rate> replays recent
route_decisions against a cold-challenger ranking and reports what the
penalty changed and what other dial settings would have changed, and the
operator re-measures the same-model/switched cache-rate gap and billed µ$ per
prompt token each re-measurement window after enabling. When the penalty
stops paying, the named move is toward neutral on the dial. No hardcoded
expiry duration: the check is the measurement.
Shipping state. Off, as shipped: behavior is byte-identical to
pre-feature traffic (pinned by test). When on, every switch decision is
explainable post-hoc: the rank debug log carries the incumbent identity, its
rate and source (measured or assumed), and the dial; the incumbent row itself
carries an incumbent_pricing stamp in the ranked output.
What routing actually returns, and why it moves
Sweeping /route across every category and tier is the fastest way to see
whether your data is doing anything. On the deployment this was written
against, 9 categories × 3 tiers currently yields 5 distinct winners
(qwen3.6-35b, gemma-4-31b, deepseek-v4-flash, kimi-k3,
kimi-k3-fast).
That number is a diagnostic, not a target, and it is worth knowing what each outcome means:
- One winner everywhere is a legitimate answer, not a misconfiguration.
It happened here: with cost, eco and proficiency all populated, one model
was Pareto-dominant — cheapest and cleanest in the routable set while
scoring within
quality_toleranceof the best. No defensible weighting picks anything else. If you see this, check whether the leader really is dominant before reaching for the config. - Winners that change with context size are the hard filters working.
A model is dropped once
required_context_tokensexceeds its window, so a long session can change model mid-conversation. Past the largest window,/v1/chat/completionsreturns 422 naming the constraint rather than silently truncating. - Winners that change by category mean proficiency is live. That is the
only category-dependent term, so until the
proficiencytable has data,task_categorycannot change a decision at all — the classifier computes it and the router pays for it for nothing. The score is now outcome-calibrated, so category-level client reports are the fastest way to shift these choices.
The spread here widened for two reasons worth copying: cost became a
per-request estimate rather than a benchmark average, and tier stopped being
inferred from price alone. Both had been quietly excluding a cheap
large-context model from every request above tier 1. Outcome samples are now
arriving too; deepseek-v4-flash gained substantial samples in
coding_general, coding_refactor, and general_chat from post-stream
POST /outcome reports.
If you want a different balance, the levers are objective.quality_tolerance
(how big a quality gap must be before it outranks cost) and
objective.max_energy_per_request (a hard ceiling). There is no weight to
tune — quality is the objective and cost is the tiebreak, which replaced an
earlier weighted blend.
Circuit Breaker — circuit_breaker.py
A seventh, dynamic filter sits alongside the six hard filters above: a model
that 5xx'd recently is passively excluded from the candidate set. On by
default (circuit_breaker.enabled: true).
It records nothing until a model actually fails — a 5xx from the upstream
call marks (model_id, provider) down for initial_cooldown_seconds (30s
default), doubling on each further failure (backoff_multiplier, capped at
max_cooldown_seconds, 600s) and clearing on the next success. Recovery is
passive by design — no background poller, no health-check loop. The next
real request that would have picked the down model becomes its own recovery
probe once the cooldown has passed.
Two call sites: candidate selection excludes down models outright, and the
dispatch retry loop (see the retry budget in verification.md)
records the failure/success on every upstream call and fails over to the
next-ranked candidate on a 5xx — but only for auto-routed requests, since a
pinned request has no alternative to fail over to.
eval_proficiency.py deliberately calls providers directly, bypassing both
the circuit breaker and the retry loop — a transient eval-harness failure
against one model's edges shouldn't be able to trip the breaker against real
production traffic.
The classifier's own circuit breaker
Separate machinery, same idea, different table: circuit_breaker.py guards
upstream models, while _last_classifier_failure in dispatcher.py guards
the local classifier — the one blocking LLM call on the request path.
For its first release that timestamp was written by _record_failure() and
read by nothing: _classify_cascade gated only its cloud step, and on a
different timestamp (_last_account_refusal). So the router re-dialled a
known-dead local classifier on every request. A stopped Ollama refuses the
connection immediately and costs little; a hung one — or a VPN-bound one
that black-holes — costs the full classifier.timeout_seconds, 120s on this
deployment, per request for as long as the outage lasts.
_classifier_backoff_active() is the read, consulted in classify()
before _classifier_client() is constructed, because constructing it and
waiting out the timeout is the expensive part.
One detail is load-bearing and easy to undo by accident: recording the failure
lives in _classify_via_local_llm's own exception handlers, not in
_classify_cascade. The cascade is walked for reasons other than a fresh
failure — a skipped attempt inside an already-open window, or gaming mode —
and if those re-stamped the clock then every request during an outage would
push the deadline forward and the local classifier would never be re-probed
while traffic flowed. That is a permanent outage wearing a circuit breaker's
clothes. tests/test_classifier_backoff.py pins it directly, and asserts on
whether the client was constructed rather than on the value returned — a
test that only checked the result would pass even if the router had waited
out the timeout first.
A skip and a real failure both reach the classifier's fallback cascade, but
they are logically distinct and dispatcher.py keeps them that way: a skip
raises _ClassifierSkipped (logged once, specifically, as
classify_local_skipped) while a real failure raises the underlying
exception (logged as fallback). Conflating the two would tell an operator
reading raw logs that the local classifier is failing when gaming mode simply
turned it off.
Which classifier is PRIMARY — classifier.mode
Separate machinery from the circuit breaker above: that guards upstream dispatch models once a category has already been classified. This guards which implementation does the classifying in the first place.
classifier.mode (local_llm default, cloud_llm, local_encoder) picks
the primary attempt. It is a peer to the classifier's own fallback cascade
(local → stale session cache → session history → optional cloud_fallback
→ static guess), not a replacement for it — whichever mode is primary, a
failure in _classify_via_configured_mode still funnels into the exact
same, unmodified cascade in dispatcher.py.
Two things are load-bearing here and easy to get backwards in a future edit:
cloud_llmsuccess recordssource="classifier", not"classifier_cloud". The latter string means specifically "the cascade's post-failure backup step fired," and both the/metricsdegradation-share warning and the outcome-attribution set read it as a degraded signal. An intentionally configured cloud primary succeeding is the opposite of degraded, so reusing"classifier_cloud"for it would make healthy, on-purpose traffic look like an ongoing local outage.cloud_primary_autoresolves live, cached, not per-request.routing.cheapest_classifier_candidatereusesselect_candidates+estimated_cost— the same functions real dispatch ranking uses — priced for the classifier's own short-prompt/short-completion call shape rather than the task's.dispatcher._resolve_auto_classifiercaches the result forclassifier.cooldown_secondsso this does not add a DB scan to every request's latency floor; the admin portal's GET endpoint calls the resolver fresh on every load instead, since a human loading a settings page is not on that latency floor and a stale "live" reading would be exactly the kind of bug the profiles page's zero-admit badge was.
local_encoder mode only ever produces task_category — task_tier falls
back to classifier.fallback_tier, a documented limitation rather than a
second heuristic. See local-models.md for why zero-shot
rather than fine-tuned (this router never stores raw task text).
Checking whether scoring earns its complexity
baseline_report.py automates the "check whether the leader really is
dominant" step above: it replays recent route_decisions against two trivial
counterfactuals (always cheapest, always highest proficiency) using the
current catalog and proficiency table, and reports how often real scoring
picked something a trivial baseline wouldn't have.
PYTHONPATH=src python -m baseline_report --since 2026-08-01
PYTHONPATH=src python -m baseline_report --since 2026-08-01 --category coding_refactor --csv
A high dominance share paired with a near-zero proficiency delta against
always_cheapest means the quality-first ranking isn't earning its complexity
for that slice of traffic. Read-only — adds no schema, spends no quota.
Tiering
Tier on reasoning_default_enabled (from metadata.reasoning.default_enabled,
falling back to capabilities.reasoning), not supports_reasoning.
capabilities.reasoning only means "the endpoint accepts a reasoning
param" — it is true for 17 of 19 rows, and tiering on it put 17 models in
tier 3 and left tier 1 empty. A -fast row does not inherit its sibling's
tier 3. Cost is checked before the reasoning rule so $0.28/1M models can
reach tier 1.
Cheapness is not a capability ceiling (tier1_context_max, default
512000). Tier is a floor — routing.py drops any row with
tier < required_tier — so tier 1 means "simple work only", not "cheap".
Deciding that on price alone put deepseek-v4-flash in tier 1 for no reason
but its $0.28/1M completion price, which excluded it outright from every
tier-2 request. It has a 1M advertised window and scores 1.00 on all three
coding categories. That was the same substitution the cost axis already had
to unlearn: price is a market signal, not a capability measurement.
Tier 1 now requires the model to be small and cheap. The gate reads the
advertised context_window, whose catalog values are the clean market
classes — 131056 / 199984 / 262128 / 1048560 — rather than
effective_context_window, which varies within a class. 512000 sits in the
empty band between the 256K and 1M classes with a 2x margin either side, so it
is not fitted to any one model. The gate only ever demotes; a huge window
never promotes an expensive model into tier 1, and a missing window does not
block it, since absent evidence should not decide anything.
Distribution moved 4 / 6 / 9 -> 1 / 9 / 9 — only the three deepseek rows
changed. With cost priced per-request, deepseek-v4-flash now wins
coding_general at every context size (16.5x cheaper than kimi-k3 at 200k)
and is still correctly absent from tool_use_agentic, where its measured 0.33
drops it out of the quality band. That is the eval data earning it the slot
rather than a thumb on the scale — no model_tiers override was needed.
Why tools and reasoning stay on their existing signals
Tools stay on the measured tool_use_agentic proficiency gate, not a
supports_tools flag gate. Every routable catalog row already has
supports_tools = 1, so a flag gate would be inert. The real signal is the
measured proficiency, because the observed failure is a model over-reaching for
tools on a non-agentic prompt. routing.min_tool_proficiency captures that
measurement and only applies when the request carries a tools array.
Reasoning stays on tiering (reasoning_default_enabled), not on a new flag gate.
supports_reasoning only means the endpoint accepts a reasoning parameter, and
that is true for 17 of 19 rows — nearly the whole catalog. has_reasoning_request
is detected purely for observation. Making it a gate would add no useful
filtering, because the decision of whether a request needs reasoning is already
encoded in the requested tier.
The fail-closed asymmetry, stated plainly: capability flags fail closed on
unknown; quality measurements admit on absent evidence. A missing
supports_vision or supports_json_mode flag means "cannot confirm", so the
model is dropped. A missing tool_use_agentic score or energy measurement means
"unproven, not bad", so the model is admitted. The first wrong guess is a
guaranteed 400; the second is just an empty data point that the neutral default
handles.
The classifier's labels are not the proficiency scoring axis
classifier.candidate_categories is the set of labels a classifier may
return. proficiency.categories is the axis models are scored on. They were
one list, and the two jobs are not the same job: the axis answers "what is
this model good at", the candidate set answers "what should a classifier be
asked to distinguish". Config load validates the candidate set as a subset
of the axis — a label outside it joins against nothing in the proficiency
table, so routing would rank on NULLs, which is the same phantom-join failure
routing.tool_use_category is validated against, arriving from the other
side. Omitting the key means "every category", i.e. the pre-split behaviour.
Both classifier backends read the candidate set: dispatcher.classify builds
its allowed-values prompt from it, and local_encoder.classify_zero_shot
receives it as its candidate labels. _classify_once also validates the
returned label against it, so the exclusion is binding rather than advisory —
a local model that ignores the prompt's list is corrected, exactly as an
invented label always was. The scoring-axis consumers were deliberately left
on the full list: leaderboard.py (priors are per scoring category),
eval_proficiency.py (the benchmark still evaluates every category),
admin.py's profile-coverage count and /metrics' category list (both
describe the scoring axis), and config.py's tool_use_category /
eligible_categories validators.
tool_use_agentic is the excluded category, and why is the point. It
describes what a turn mechanically does rather than what it is for, and
every agent turn does it. Measured 2026-09-14 on live traffic with the session
cache off so every turn classified for real: 31 of 31 consecutive turns
classified tool_use_agentic, all carrying tools, all routed to
z-ai/glm-5.3-flash — 14.8s time-to-first-token, p95 40s, the slowest model in
the catalog. Per-turn classification became accurate and routing got worse.
Classifying for tool use also duplicates a signal the router already has
exactly: the request carries a tools array, reading it is free, and
routing.min_tool_proficiency is the mechanism that acts on it.
What the exclusion costs, written down so the next person does not have to
rediscover it: no new POST /outcome report can attribute to
tool_use_agentic, because nothing classifies as it any more. Those scores
freeze at their current values (qwen3.6-35b at 0.902 over 154 samples).
The tool filter keeps reading them and the category keeps ranking; it simply
stops accumulating. That is accepted for now, not fixed.
Considered and rejected: attributing outcomes to tool_use_agentic
whenever the request carried a tools array. opencode sends tools on
essentially every request, so the score would converge on each model's overall
pass rate and stop discriminating the one thing the category exists to
detect — a model that over-reaches for tools on work that did not need them
(deepseek-v4-flash calling two tools to subtract 1:20pm from 3pm). A frozen
honest score beats a live meaningless one.
The readout is on the admin portal's Classifier card: the candidate list, and
the excluded categories with a note that they stop accumulating outcomes.
Read-only on purpose — it is a structural list in the same class as
proficiency.categories and models.eligible_categories, neither of which
has an editing surface, and what an operator needs from the portal is the
answer to "why does nothing ever classify as X".