Files
6krrt/docs/routing.md
2026-09-20 20:28:22 -04:00

37 KiB
Raw Permalink Blame History

Deep dive into routing internals. Back: README.

Quality-First Selection

Seven hard filters are applied before scoring (not weighted — outright disqualification):

  1. effective_context_window ≥ required_context_tokens
  2. tier ≥ required_tier (from classifier)
  3. Serving class compatible with request's latency_tolerance; access_level reachable
  4. tool_use_agentic proficiency ≥ routing.min_tool_proficiency, only when the request carries a tools array (ships disabled — null)
  5. supports_vision = 1 when the request carries image_url parts; NULL fails closed (an unknown flag means the capability cannot be confirmed)
  6. supports_json_mode = 1 when response_format.type is json_object or json_schema; NULL also fails closed
  7. eligible_categories restricts only when set: a row with a non-empty category list is admitted only when the task's category is in that list; NULL/absent means unrestricted (all cloud rows today)

Filters 4–6 are read from the request body rather than inferred. A tools array states whether tool definitions are on the table, image_url parts state whether vision is needed, and response_format states whether JSON mode is needed. Local classifiers identified unambiguous tool-use prompts only 1-2 times in 6, so inferring capability requirements does not work; the request states them exactly and for free.

The tool gate is a quality measurement filter: a model with no measured tool_use_agentic score is admitted, because "unproven" is not "proven bad". Only a measured score below routing.min_tool_proficiency is dropped. The vision and JSON-mode gates are capability flag filters. There every routable catalog row has supports_tools = 1, so a flag gate would be inert; vision and JSON mode are not universal, and an absent or NULL flag means "cannot confirm the capability", so the model is dropped. A wrong guess on a capability flag is a guaranteed provider 400.

The tool filter ships off (null), because agent clients send tools on nearly every request, so enabling it excludes the cheapest model from ordinary agent traffic. Whether that trade is worth it is an empirical question best settled with POST /outcome data rather than a 3-task benchmark score. Set it to 0.5 to turn it on. The vision and JSON-mode gates ship on, because a wrong guess is a guaranteed 400.

Local Vision Fallback

Cloud vision is not universal in the catalog, and the most economical coding rows do not declare supports_vision. Routing an image request through them would earn a provider-side 400, so when routing.require_vision is true only rows with supports_vision = 1 survive the hard filters. If no cloud candidate survives, the router can fall back to a local vision model instead of returning 422.

local_vision: in config/config.yaml controls this path:

Key Default Purpose
enabled true Whether the fallback runs at all
base_url http://localhost:11434/v1 OpenAI-compatible Ollama endpoint
api_key_env null Env var holding an API key, if the endpoint needs one
model qwen3-vl:4b Vision model on that Ollama
timeout_seconds 60 Request timeout
max_images 4 Refuse requests with more image parts
max_image_bytes 9437184 (9 MiB) Refuse requests whose image payload exceeds this

The fallback is enabled by default both in config/config.yaml and in LocalVisionConfig, so omitting the section still turns it on. Disable it explicitly (enabled: false) on a host with no local Ollama or one that has not pulled the vision model.

When enabled and no cloud candidate is selected, _run_local_vision sends the original message list — with image_url parts intact — to the configured Ollama model. The local answer then replaces the cloud completion: _local_vision_response returns a normal OpenAI-shaped response, including a stream-wrapped version when the client asked for stream: true. It does not inject a caption into a cloud call, because the streaming proxy cannot rewrite bytes mid-stream.

Security and budget guards:

  • Only inline data: URIs are accepted. A remote http(s) image URL is declined, because pointing a local model at an arbitrary URL would let an unauthenticated caller make the router fetch internal resources (SSRF).
  • Image count and total payload size are bounded by max_images and max_image_bytes before the local call is made.
  • If the local call fails for any reason, the request falls through to the ordinary 422 No model satisfies the hard filters rather than returning a silent empty response.

Pull the model on whichever Ollama the fallback points at:

ollama pull qwen3-vl:4b

Config is strict (extra="forbid"): a misspelled or misplaced key fails at load instead of being silently ignored.

Local dispatch branch

Tasks in a local model's eligible_categories (for example file_summarization and diff_checking for the shipped qwen2.5-coder-router:14b) survive the hard filters the same way cloud rows do. Once selected, the request is sent to the configured Ollama endpoint (provider='ollama-local') instead of to dispatch_providers. A local 502 trips the circuit breaker, so the next request sees the local row as excluded and reroutes to a cloud candidate.

Local dispatch is quality-gated and dormant under the default profile. The local row competes in the same ranking as every cloud row and must score within objective.quality_tolerance of the category's cloud leader before its price advantage is even consulted. Measured on 2026-09-02/03, qwen2.5-coder-router:14b scores 0.767 on file_summarization against the cloud leader's 0.95 — a 0.183 gap against a 0.10 tolerance — so under the default profile the local row is dormant by design, today, until a local model scores within tolerance. This is not a bug: quality is the objective, and the local model has not yet earned the tiebreak (its 9.1x cost advantage applies mainly to prompt tokens, so it wins most where large-context summarization dominates — and it is assigned to summarization precisely because it is not good enough at code).

Do NOT widen objective.quality_tolerance to fix this: the tolerance describes measurement noise (bench sample sizes are small enough that a smaller gap is genuinely indistinguishable), not preference. The project's direction is named profiles (auto:batch is already one — a named widening of the candidate set), and a locality profile is the right future home for a per-category local preference, not a global tolerance change.

When every cloud candidate is exhausted, or the upstream returns an account-level 4xx (any 4xx except 400/404/422 — those mean the request itself is malformed and would fail locally too), an eligible routed request degrades to the local dispatch model instead of surfacing the cloud error. The trigger is deliberately conservative because the upstream's exact credit-exhaustion signature is not yet observed (full-body upstream-failure logging is in place to capture it); the remaining cloud rows share one account, so no time is wasted walking them after an account-level refusal. The fallback respects eligible_categories even here: for everything else, the cloud error fails loudly rather than handing a code task to a summarization-grade model. The degraded answer is recorded as kind='local_dispatch_fallback' in route_decisions so the proficiency loop never mistakes it for local winning on merit, and if the local call itself fails the original cloud error surfaces (the same fail-through contract as the local vision fallback). There is no separate config knob — the fallback is armed exactly by a local_dispatch_models entry listing the category in its eligible_categories.

1. drop candidates whose measured energy exceeds objective.max_energy_per_request
2. rank by expected pass rate for the task's category (proficiency.blended_score)
3. treat differences smaller than objective.quality_tolerance as equal
4. among equals, pick the cheapest
5. on a small share of eligible requests, explore the least-evidenced candidate
Setting Value Notes
quality_tolerance 0.10 % of real success rate, not an abstract quality unit: a candidate must beat the best by more than this to outrank a cheaper alternative
assumed_cache_rate 0.917 Share of prompt tokens served from the provider's prefix cache. Agent clients resend the conversation each turn, so most of it hits. Measured token-weighted over 40.7M tokens; measure your own against the provider's session view
assumed_completion_tokens 500 Completion length assumed when pricing a candidate
max_energy_per_request null Per-request kWh ceiling — a wall, not a bill. Null disables it
plan_kwh_per_period 6.25 Set this to your own plan's quota. Reported in /health as burn against the allowance; it does not gate anything

This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4). Measurement retired it: turning the cost weight from 0.4 to zero changed the winner in only 2 of 6 categories, so the blend was never steering on quality — while 60% of every decision adjudicated fractions of a cent.

What the proficiency score means now

proficiency.blended_score is an expected pass rate on your traffic, not a raw benchmark quality level. The benchmark (leaderboard + self-eval) provides a prior. Client-reported outcomes from POST /outcome provide the per-model signal. The two are combined through empirical-Bayes shrinkage:

  • A category with no outcome traffic keeps the benchmark score verbatim.
  • A trafficked category row with no own outcomes inherits a peer prior.
  • A row with its own outcomes blends the observed rate with that prior; outcome_prior_strength (default 20 pseudo-observations) pulls thin data toward the category mean so one lucky sample cannot dominate.

So quality_tolerance asks "is this candidate's expected real success rate at least 10% better than the cheaper one?" If not, cost breaks the tie. The score is calibrated against the same client reports that feedback.py folds in, so routing learns from traffic rather than from the benchmark alone.

Cost is priced per request, from catalog token prices scaled to the request's shape — prompt size, assumed_completion_tokens, assumed_cache_rate. It is deliberately not a benchmark average.

Neuralwatt bills flat per-kWh rather than per-token, so list price is not what gets charged. Scoring used measured billed cost from a fixed 400-token reference sweep for exactly that reason — and that was wrong for real traffic, because the ranking depends on the workload's shape, not just the model. On a 400-token prompt one model looked 3.2× cheaper than another; on a realistic 70k-token prompt the same pair inverted and the second was 5.0× cheaper. The provider's attribution ratio moves with prompt size, so a fixed-shape benchmark cannot rank models for a workload of a different shape.

List price is still not what is billed, but billing is capped at a multiple of it, so it tracks the real ordering and bounds it — and it is free, needs no sweep, and refreshes whenever the poller runs.

Why cost ≠ eco: Cost tracks energy (kWh), but carbon is energy × grid intensity. Grid intensity spanned ~49 gCO2/kWh (FI) to ~442 (US-MIDA-PJM) when measured — and it moves with time of day. The models disagree: glm-5.2-fast is 2nd cheapest but 6th cleanest; kimi-k3-flex draws 3.7× less energy than kimi-k2.7-code while emitting 3.6× more carbon. Collapsing them picks a side.

Exploration

Routing runs an epsilon-greedy exploration pass after ranking. If enabled (exploration.enabled: true, default), a random share ε = 0.03 of eligible requests replaces the winner with the hard-filter-eligible candidate that has the fewest outcome_samples, tie-broken by lowest cost. The alternative is skipped if its cost exceeds max_cost_ratio × the winner's cost (default 4×). Only tier 1 and tier 2 requests explore (max_tier: 2); tier 3 stays exploitative so frontier work gets the best expected rate.

was_exploration is persisted to route_decisions.exploration so metrics can tell real preference from forced discovery. This is the mechanism that breaks the exposure-bias loop: without it, the router would send most requests to the models already richest in outcome samples and the least-sampled rows would never catch up.

Incumbency and cache pricing

The ranking above treats every candidate as if its prompt were fresh. On a real session it is not: the provider caches the conversation prefix, so the model that served the last turn re-serves it with most of its prompt tokens billed at the cached rate, while a different model pays full price for the entire prefix on its first turn. assumed_cache_rate already prices that discount uniformly for every row. Incumbency pricing makes it per-row: the incumbent keeps its measured cache rate, challengers are priced toward cold, and a challenger now has to beat the incumbent by more than the cache it is about to discard, which on a long prompt is most of the prompt. This is the first place the router's own past decision feeds back into its cost model, and it is deliberately narrow: it changes only the cost key, never the sort.

The incumbent is per conversation when the client sends X-Router-Conversation: session_key is then the namespaced conversation id, so each conversation tracks its own incumbent. Without that header the incumbent is per system-prompt fingerprint, and the fingerprint's hashing merges concurrent same-prompt agents into one incumbent so they share the cache discount.

Four knobs under objective:, all shipping off (incumbent_cache_pricing: false, dial null; enable only after the Wave 1 post-restart baseline day, the Wave 2 gate in plans/token-waste-waves.md):

Knob Default Purpose
incumbent_cache_pricing false The gate; off means byte-identical ranking (pinned by test)
incumbent_challenger_cache_rate null The challenger dial: null follows assumed_cache_rate; 0.0 is fully cold; between is a partial penalty
incumbent_rate_refresh_seconds 300 TTL on the measured per-(provider, model) rate table
incumbent_rate_min_observations 25 Observations before a measured rate is trusted for pricing; independent of the warning floor

What the incumbent is. The model that served the session's last chat turn, read back from route_decisions by _session_incumbent_lookup in dispatcher.py: the most recent row for the session key with a selected model. The query is an allowlist (kind IN ('chat'), never a != denylist), so a one-request model pin (passthrough), a local_dispatch_fallback row, or a route probe can never set the incumbent, and neither can any future kind until it argues its way in. The reason is that incumbency is a statement about preference: this session, with its own history, chose that model repeatedly on merit. A single pinned request says nothing about the session and a degraded fallback is the opposite of a preference; letting either become sticky would convert an escape hatch into a rut.

The rate ladder. When the feature is on and an incumbent is present, the incumbent is priced at its measured per-(provider, model) cache rate whenever that series has at least incumbent_rate_min_observations observations (default 25); below that trust threshold it falls back to assumed_cache_rate. The rates come from metrics.cache_rate_series through _measured_cache_rates() in dispatcher.py, cached for incumbent_rate_refresh_seconds (the TTL runs on time.monotonic() from the first call after process start, so a brief post-restart cold period is expected). The pricing floor is deliberately a different knob from the alerting floor cache_rate_warn_min_observations: pricing would rather fall back to the assumed rate on thin data than steer money on a thin measurement, alerting wants the opposite trade, and coupling the two would couple operator intents that should stay separate. Any failure in the measured-rate path fails open to an empty table and every row prices at assumed_cache_rate; incumbency pricing must never block routing.

The challenger dial. objective.incumbent_challenger_cache_rate is the rate every non-incumbent row is priced at. Neutral is decided by semantic equality, not identity: null follows assumed_cache_rate, and any dial equal to assumed_cache_rate is the same neutral, so at a neutral setting every row including the incumbent prices at assumed_cache_rate and the ranking is byte-identical to today's, even with the flag on (config resolves null to a concrete float at load; downstream code never sees both representations). 0.0 prices challengers as fully cold prompts, the maximum incumbent advantage; any value between is a partial cache penalty. The whole feature tunes from off to full by moving this one config value, never by a revert: if the switch gap does not narrow, the move is toward cold, then re-measure.

The clamp, and why it is load-bearing. The challenger rate is min(dial, incumbent_rate), and the invariant it buys is: the incumbent's cache rate is always >= every challenger's cache rate. Without it, the dial is denominated in an absolute cache rate, so any dial above the incumbent's measured rate prices the challenger as having a better cache than the incumbent: the feature inverts and penalises the incumbent for being the incumbent. Because several routed models measure below the assumed rate (a model measured at ~0.73 against an assumed 0.917, for instance), that inverted zone is not a corner case: for such a model it spans essentially the whole walk from neutral down toward full penalty, which is exactly the range an operator is told to tune through, and from inside it the counterfactual report looks like "the penalty does not pay here" when in fact the penalty is running backwards. Do not remove the clamp as redundant. An alternative dial denominated as a relative penalty (incumbent_rate * (1 - penalty)) cannot invert by construction and was deliberately deferred rather than adopted (see the wave's second review, wave2-review-2.md); until that redesign exists, the clamp is what makes the absolute form safe, and a property test pins it: for every dial and every measured rate, the incumbent's priced rate is >= every challenger's.

How it composes. Cache pricing happens inside estimated_cost: the ranking loop simply passes each row a different cache_rate argument, so every consumer of cost, from the band tiebreak to cost_score to the flex-twin swap (a distinct serving endpoint, so it prices as a challenger at the same dial), sees one coherent number per row. provider_cost_multipliers composes multiplicatively on top of the cache-adjusted estimate, exactly as before: it inflates the comparison cost inside the quality-band tiebreak and can never override a genuine quality gap, because the quality band is computed before the cost key and the incumbent gets no quality advantage. A model that is genuinely better still wins outright.

Eviction. The incumbent is a pricing fact, not a reservation. _resolve_incumbent re-runs the incumbent's catalog row through rejection_reason with this request's actual filters and checks the profile's restrict_to: if a hard filter would drop it (the session outgrew its context window, a profile switch restricts the candidate set, a circuit is open), it simply loses incumbency for that turn, no error, and the turn prices exactly as it did before the feature: every row at assumed_cache_rate.

Exploration becomes session-scoped. The coin described above now flips only when no incumbent is present: route() adds incumbent is None to the exploration condition, and exploration.py itself is unchanged (injected RNG, no mutable state, per its module contract). The coin flips at session start only, so one exploratory session costs one cold prompt instead of one per turn; incumbency pricing then pins the explored model for the session's remaining turns, which is tiebreak protection without new state. Re-exploring per turn with an incumbent present would dump the cache every few turns, the exact leak the feature exists to stop (plans/token-waste-waves.md item 2.3).

Expiry checks. Both premises carry their own. The assumed_cache_rate fallback is checked by cache_rate_warnings in /metrics, which compares the measured aggregate and per-(provider, model) rates against the assumed constant and warns beyond cache_rate_warn_margin. A deployment whose real rate has diverged is mispriced, not broken, and the fix is to re-measure. The dial's own premise, "the penalty is paying," is evaluated offline rather than assumed: baseline_report.py --incumbent-challenger <rate> replays recent route_decisions against a cold-challenger ranking and reports what the penalty changed and what other dial settings would have changed, and the operator re-measures the same-model/switched cache-rate gap and billed µ$ per prompt token each re-measurement window after enabling. When the penalty stops paying, the named move is toward neutral on the dial. No hardcoded expiry duration: the check is the measurement.

Shipping state. Off, as shipped: behavior is byte-identical to pre-feature traffic (pinned by test). When on, every switch decision is explainable post-hoc: the rank debug log carries the incumbent identity, its rate and source (measured or assumed), and the dial; the incumbent row itself carries an incumbent_pricing stamp in the ranked output.

What routing actually returns, and why it moves

Sweeping /route across every category and tier is the fastest way to see whether your data is doing anything. On the deployment this was written against, 9 categories × 3 tiers currently yields 5 distinct winners (qwen3.6-35b, gemma-4-31b, deepseek-v4-flash, kimi-k3, kimi-k3-fast).

That number is a diagnostic, not a target, and it is worth knowing what each outcome means:

  • One winner everywhere is a legitimate answer, not a misconfiguration. It happened here: with cost, eco and proficiency all populated, one model was Pareto-dominant — cheapest and cleanest in the routable set while scoring within quality_tolerance of the best. No defensible weighting picks anything else. If you see this, check whether the leader really is dominant before reaching for the config.
  • Winners that change with context size are the hard filters working. A model is dropped once required_context_tokens exceeds its window, so a long session can change model mid-conversation. Past the largest window, /v1/chat/completions returns 422 naming the constraint rather than silently truncating.
  • Winners that change by category mean proficiency is live. That is the only category-dependent term, so until the proficiency table has data, task_category cannot change a decision at all — the classifier computes it and the router pays for it for nothing. The score is now outcome-calibrated, so category-level client reports are the fastest way to shift these choices.

The spread here widened for two reasons worth copying: cost became a per-request estimate rather than a benchmark average, and tier stopped being inferred from price alone. Both had been quietly excluding a cheap large-context model from every request above tier 1. Outcome samples are now arriving too; deepseek-v4-flash gained substantial samples in coding_general, coding_refactor, and general_chat from post-stream POST /outcome reports.

If you want a different balance, the levers are objective.quality_tolerance (how big a quality gap must be before it outranks cost) and objective.max_energy_per_request (a hard ceiling). There is no weight to tune — quality is the objective and cost is the tiebreak, which replaced an earlier weighted blend.

Circuit Breaker — circuit_breaker.py

A seventh, dynamic filter sits alongside the six hard filters above: a model that 5xx'd recently is passively excluded from the candidate set. On by default (circuit_breaker.enabled: true).

It records nothing until a model actually fails — a 5xx from the upstream call marks (model_id, provider) down for initial_cooldown_seconds (30s default), doubling on each further failure (backoff_multiplier, capped at max_cooldown_seconds, 600s) and clearing on the next success. Recovery is passive by design — no background poller, no health-check loop. The next real request that would have picked the down model becomes its own recovery probe once the cooldown has passed.

Two call sites: candidate selection excludes down models outright, and the dispatch retry loop (see the retry budget in verification.md) records the failure/success on every upstream call and fails over to the next-ranked candidate on a 5xx — but only for auto-routed requests, since a pinned request has no alternative to fail over to.

eval_proficiency.py deliberately calls providers directly, bypassing both the circuit breaker and the retry loop — a transient eval-harness failure against one model's edges shouldn't be able to trip the breaker against real production traffic.

The classifier's own circuit breaker

Separate machinery, same idea, different table: circuit_breaker.py guards upstream models, while _last_classifier_failure in dispatcher.py guards the local classifier — the one blocking LLM call on the request path.

For its first release that timestamp was written by _record_failure() and read by nothing: _classify_cascade gated only its cloud step, and on a different timestamp (_last_account_refusal). So the router re-dialled a known-dead local classifier on every request. A stopped Ollama refuses the connection immediately and costs little; a hung one — or a VPN-bound one that black-holes — costs the full classifier.timeout_seconds, 120s on this deployment, per request for as long as the outage lasts.

_classifier_backoff_active() is the read, consulted in classify() before _classifier_client() is constructed, because constructing it and waiting out the timeout is the expensive part.

One detail is load-bearing and easy to undo by accident: recording the failure lives in _classify_via_local_llm's own exception handlers, not in _classify_cascade. The cascade is walked for reasons other than a fresh failure — a skipped attempt inside an already-open window, or gaming mode — and if those re-stamped the clock then every request during an outage would push the deadline forward and the local classifier would never be re-probed while traffic flowed. That is a permanent outage wearing a circuit breaker's clothes. tests/test_classifier_backoff.py pins it directly, and asserts on whether the client was constructed rather than on the value returned — a test that only checked the result would pass even if the router had waited out the timeout first.

A skip and a real failure both reach the classifier's fallback cascade, but they are logically distinct and dispatcher.py keeps them that way: a skip raises _ClassifierSkipped (logged once, specifically, as classify_local_skipped) while a real failure raises the underlying exception (logged as fallback). Conflating the two would tell an operator reading raw logs that the local classifier is failing when gaming mode simply turned it off.

Which classifier is PRIMARY — classifier.mode

Separate machinery from the circuit breaker above: that guards upstream dispatch models once a category has already been classified. This guards which implementation does the classifying in the first place.

classifier.mode (local_llm default, cloud_llm, local_encoder) picks the primary attempt. It is a peer to the classifier's own fallback cascade (local → stale session cache → session history → optional cloud_fallback → static guess), not a replacement for it — whichever mode is primary, a failure in _classify_via_configured_mode still funnels into the exact same, unmodified cascade in dispatcher.py.

Two things are load-bearing here and easy to get backwards in a future edit:

  • cloud_llm success records source="classifier", not "classifier_cloud". The latter string means specifically "the cascade's post-failure backup step fired," and both the /metrics degradation-share warning and the outcome-attribution set read it as a degraded signal. An intentionally configured cloud primary succeeding is the opposite of degraded, so reusing "classifier_cloud" for it would make healthy, on-purpose traffic look like an ongoing local outage.
  • cloud_primary_auto resolves live, cached, not per-request. routing.cheapest_classifier_candidate reuses select_candidates + estimated_cost — the same functions real dispatch ranking uses — priced for the classifier's own short-prompt/short-completion call shape rather than the task's. dispatcher._resolve_auto_classifier caches the result for classifier.cooldown_seconds so this does not add a DB scan to every request's latency floor; the admin portal's GET endpoint calls the resolver fresh on every load instead, since a human loading a settings page is not on that latency floor and a stale "live" reading would be exactly the kind of bug the profiles page's zero-admit badge was.

local_encoder mode only ever produces task_category — task_tier falls back to classifier.fallback_tier, a documented limitation rather than a second heuristic. See local-models.md for why zero-shot rather than fine-tuned (this router never stores raw task text).

Checking whether scoring earns its complexity

baseline_report.py automates the "check whether the leader really is dominant" step above: it replays recent route_decisions against two trivial counterfactuals (always cheapest, always highest proficiency) using the current catalog and proficiency table, and reports how often real scoring picked something a trivial baseline wouldn't have.

PYTHONPATH=src python -m baseline_report --since 2026-08-01
PYTHONPATH=src python -m baseline_report --since 2026-08-01 --category coding_refactor --csv

A high dominance share paired with a near-zero proficiency delta against always_cheapest means the quality-first ranking isn't earning its complexity for that slice of traffic. Read-only — adds no schema, spends no quota.

Tiering

Tier on reasoning_default_enabled (from metadata.reasoning.default_enabled, falling back to capabilities.reasoning), not supports_reasoning. capabilities.reasoning only means "the endpoint accepts a reasoning param" — it is true for 17 of 19 rows, and tiering on it put 17 models in tier 3 and left tier 1 empty. A -fast row does not inherit its sibling's tier 3. Cost is checked before the reasoning rule so $0.28/1M models can reach tier 1.

Cheapness is not a capability ceiling (tier1_context_max, default 512000). Tier is a floor — routing.py drops any row with tier < required_tier — so tier 1 means "simple work only", not "cheap". Deciding that on price alone put deepseek-v4-flash in tier 1 for no reason but its $0.28/1M completion price, which excluded it outright from every tier-2 request. It has a 1M advertised window and scores 1.00 on all three coding categories. That was the same substitution the cost axis already had to unlearn: price is a market signal, not a capability measurement.

Tier 1 now requires the model to be small and cheap. The gate reads the advertised context_window, whose catalog values are the clean market classes — 131056 / 199984 / 262128 / 1048560 — rather than effective_context_window, which varies within a class. 512000 sits in the empty band between the 256K and 1M classes with a 2x margin either side, so it is not fitted to any one model. The gate only ever demotes; a huge window never promotes an expensive model into tier 1, and a missing window does not block it, since absent evidence should not decide anything.

Distribution moved 4 / 6 / 9 -> 1 / 9 / 9 — only the three deepseek rows changed. With cost priced per-request, deepseek-v4-flash now wins coding_general at every context size (16.5x cheaper than kimi-k3 at 200k) and is still correctly absent from tool_use_agentic, where its measured 0.33 drops it out of the quality band. That is the eval data earning it the slot rather than a thumb on the scale — no model_tiers override was needed.

Why tools and reasoning stay on their existing signals

Tools stay on the measured tool_use_agentic proficiency gate, not a supports_tools flag gate. Every routable catalog row already has supports_tools = 1, so a flag gate would be inert. The real signal is the measured proficiency, because the observed failure is a model over-reaching for tools on a non-agentic prompt. routing.min_tool_proficiency captures that measurement and only applies when the request carries a tools array.

Reasoning stays on tiering (reasoning_default_enabled), not on a new flag gate. supports_reasoning only means the endpoint accepts a reasoning parameter, and that is true for 17 of 19 rows — nearly the whole catalog. has_reasoning_request is detected purely for observation. Making it a gate would add no useful filtering, because the decision of whether a request needs reasoning is already encoded in the requested tier.

The fail-closed asymmetry, stated plainly: capability flags fail closed on unknown; quality measurements admit on absent evidence. A missing supports_vision or supports_json_mode flag means "cannot confirm", so the model is dropped. A missing tool_use_agentic score or energy measurement means "unproven, not bad", so the model is admitted. The first wrong guess is a guaranteed 400; the second is just an empty data point that the neutral default handles.

The classifier's labels are not the proficiency scoring axis

classifier.candidate_categories is the set of labels a classifier may return. proficiency.categories is the axis models are scored on. They were one list, and the two jobs are not the same job: the axis answers "what is this model good at", the candidate set answers "what should a classifier be asked to distinguish". Config load validates the candidate set as a subset of the axis — a label outside it joins against nothing in the proficiency table, so routing would rank on NULLs, which is the same phantom-join failure routing.tool_use_category is validated against, arriving from the other side. Omitting the key means "every category", i.e. the pre-split behaviour.

Both classifier backends read the candidate set: dispatcher.classify builds its allowed-values prompt from it, and local_encoder.classify_zero_shot receives it as its candidate labels. _classify_once also validates the returned label against it, so the exclusion is binding rather than advisory — a local model that ignores the prompt's list is corrected, exactly as an invented label always was. The scoring-axis consumers were deliberately left on the full list: leaderboard.py (priors are per scoring category), eval_proficiency.py (the benchmark still evaluates every category), admin.py's profile-coverage count and /metrics' category list (both describe the scoring axis), and config.py's tool_use_category / eligible_categories validators.

tool_use_agentic is the excluded category, and why is the point. It describes what a turn mechanically does rather than what it is for, and every agent turn does it. Measured 2026-09-14 on live traffic with the session cache off so every turn classified for real: 31 of 31 consecutive turns classified tool_use_agentic, all carrying tools, all routed to z-ai/glm-5.3-flash — 14.8s time-to-first-token, p95 40s, the slowest model in the catalog. Per-turn classification became accurate and routing got worse. Classifying for tool use also duplicates a signal the router already has exactly: the request carries a tools array, reading it is free, and routing.min_tool_proficiency is the mechanism that acts on it.

What the exclusion costs, written down so the next person does not have to rediscover it: no new POST /outcome report can attribute to tool_use_agentic, because nothing classifies as it any more. Those scores freeze at their current values (qwen3.6-35b at 0.902 over 154 samples). The tool filter keeps reading them and the category keeps ranking; it simply stops accumulating. That is accepted for now, not fixed.

Considered and rejected: attributing outcomes to tool_use_agentic whenever the request carried a tools array. opencode sends tools on essentially every request, so the score would converge on each model's overall pass rate and stop discriminating the one thing the category exists to detect — a model that over-reaches for tools on work that did not need them (deepseek-v4-flash calling two tools to subtract 1:20pm from 3pm). A frozen honest score beats a live meaningless one.

The readout is on the admin portal's Classifier card: the candidate list, and the excluded categories with a note that they stop accumulating outcomes. Read-only on purpose — it is a structural list in the same class as proficiency.categories and models.eligible_categories, neither of which has an editing surface, and what an operator needs from the portal is the answer to "why does nothing ever classify as X".