> Deep dive into routing internals. Back: [README](../README.md). ## Quality-First Selection Seven hard filters are applied **before** scoring (not weighted — outright disqualification): 1. `effective_context_window ≥ required_context_tokens` 2. `tier ≥ required_tier` (from classifier) 3. Serving class compatible with request's `latency_tolerance`; `access_level` reachable 4. `tool_use_agentic` proficiency ≥ `routing.min_tool_proficiency`, **only when the request carries a `tools` array** (ships disabled — `null`) 5. `supports_vision = 1` when the request carries `image_url` parts; `NULL` fails closed (an unknown flag means the capability cannot be confirmed) 6. `supports_json_mode = 1` when `response_format.type` is `json_object` or `json_schema`; `NULL` also fails closed 7. `eligible_categories` restricts only when set: a row with a non-empty category list is admitted only when the task's category is in that list; `NULL`/absent means unrestricted (all cloud rows today) Filters 4–6 are read from the request body rather than inferred. A `tools` array states whether tool definitions are on the table, `image_url` parts state whether vision is needed, and `response_format` states whether JSON mode is needed. Local classifiers identified unambiguous tool-use prompts only 1-2 times in 6, so inferring capability requirements does not work; the request states them exactly and for free. The tool gate is a **quality measurement** filter: a model with no measured `tool_use_agentic` score is admitted, because "unproven" is not "proven bad". Only a measured score below `routing.min_tool_proficiency` is dropped. The vision and JSON-mode gates are **capability flag** filters. There every routable catalog row has `supports_tools = 1`, so a flag gate would be inert; vision and JSON mode are not universal, and an absent or `NULL` flag means "cannot confirm the capability", so the model is dropped. A wrong guess on a capability flag is a guaranteed provider 400. The tool filter ships **off** (`null`), because agent clients send `tools` on nearly every request, so enabling it excludes the cheapest model from ordinary agent traffic. Whether that trade is worth it is an empirical question best settled with `POST /outcome` data rather than a 3-task benchmark score. Set it to `0.5` to turn it on. The vision and JSON-mode gates ship **on**, because a wrong guess is a guaranteed 400. ### Local Vision Fallback Cloud vision is not universal in the catalog, and the most economical coding rows do not declare `supports_vision`. Routing an image request through them would earn a provider-side 400, so when `routing.require_vision` is true only rows with `supports_vision = 1` survive the hard filters. If **no** cloud candidate survives, the router can fall back to a local vision model instead of returning 422. `local_vision:` in `config/config.yaml` controls this path: | Key | Default | Purpose | |---|---|---| | `enabled` | `true` | Whether the fallback runs at all | | `base_url` | `http://localhost:11434/v1` | OpenAI-compatible Ollama endpoint | | `api_key_env` | `null` | Env var holding an API key, if the endpoint needs one | | `model` | `qwen3-vl:4b` | Vision model on that Ollama | | `timeout_seconds` | `60` | Request timeout | | `max_images` | `4` | Refuse requests with more image parts | | `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this | The fallback is **enabled by default** both in `config/config.yaml` and in `LocalVisionConfig`, so omitting the section still turns it on. Disable it explicitly (`enabled: false`) on a host with no local Ollama or one that has not pulled the vision model. When enabled and no cloud candidate is selected, `_run_local_vision` sends the original message list — with `image_url` parts intact — to the configured Ollama model. The local answer then **replaces** the cloud completion: `_local_vision_response` returns a normal OpenAI-shaped response, including a stream-wrapped version when the client asked for `stream: true`. It does *not* inject a caption into a cloud call, because the streaming proxy cannot rewrite bytes mid-stream. Security and budget guards: - **Only inline `data:` URIs are accepted.** A remote `http(s)` image URL is declined, because pointing a local model at an arbitrary URL would let an unauthenticated caller make the router fetch internal resources (SSRF). - Image count and total payload size are bounded by `max_images` and `max_image_bytes` before the local call is made. - If the local call fails for any reason, the request falls through to the ordinary `422 No model satisfies the hard filters` rather than returning a silent empty response. Pull the model on whichever Ollama the fallback points at: ```bash ollama pull qwen3-vl:4b ``` Config is strict (`extra="forbid"`): a misspelled or misplaced key fails at load instead of being silently ignored. ### Local dispatch branch Tasks in a local model's `eligible_categories` (for example `file_summarization` and `diff_checking` for the shipped `qwen2.5-coder-router:14b`) survive the hard filters the same way cloud rows do. Once selected, the request is sent to the configured Ollama endpoint (`provider='ollama-local'`) instead of to `dispatch_providers`. A local 502 trips the circuit breaker, so the next request sees the local row as excluded and reroutes to a cloud candidate. Local dispatch is **quality-gated and dormant under the default profile**. The local row competes in the same ranking as every cloud row and must score within `objective.quality_tolerance` of the category's cloud leader before its price advantage is even consulted. Measured on 2026-09-02/03, `qwen2.5-coder-router:14b` scores **0.767** on `file_summarization` against the cloud leader's **0.95** — a 0.183 gap against a 0.10 tolerance — so under the default profile the local row is dormant **by design**, today, until a local model scores within tolerance. This is not a bug: quality is the objective, and the local model has not yet earned the tiebreak (its 9.1x cost advantage applies mainly to prompt tokens, so it wins most where large-context summarization dominates — and it is assigned to summarization precisely because it is not good enough at code). Do NOT widen `objective.quality_tolerance` to fix this: the tolerance describes measurement noise (bench sample sizes are small enough that a smaller gap is genuinely indistinguishable), not preference. The project's direction is named profiles (`auto:batch` is already one — a named widening of the candidate set), and a `locality` profile is the right future home for a per-category local preference, not a global tolerance change. When every cloud candidate is exhausted, or the upstream returns an account-level 4xx (any 4xx except 400/404/422 — those mean the request itself is malformed and would fail locally too), an eligible routed request degrades to the local dispatch model instead of surfacing the cloud error. The trigger is deliberately conservative because the upstream's exact credit-exhaustion signature is not yet observed (full-body upstream-failure logging is in place to capture it); the remaining cloud rows share one account, so no time is wasted walking them after an account-level refusal. The fallback respects `eligible_categories` even here: for everything else, the cloud error fails loudly rather than handing a code task to a summarization-grade model. The degraded answer is recorded as `kind='local_dispatch_fallback'` in `route_decisions` so the proficiency loop never mistakes it for local winning on merit, and if the local call itself fails the original cloud error surfaces (the same fail-through contract as the local vision fallback). There is no separate config knob — the fallback is armed exactly by a `local_dispatch_models` entry listing the category in its `eligible_categories`. ``` 1. drop candidates whose measured energy exceeds objective.max_energy_per_request 2. rank by expected pass rate for the task's category (proficiency.blended_score) 3. treat differences smaller than objective.quality_tolerance as equal 4. among equals, pick the cheapest 5. on a small share of eligible requests, explore the least-evidenced candidate ``` | Setting | Value | Notes | |---|---|---| | `quality_tolerance` | 0.10 | % of real success rate, not an abstract quality unit: a candidate must beat the best by more than this to outrank a cheaper alternative | | `assumed_cache_rate` | 0.917 | Share of prompt tokens served from the provider's prefix cache. Agent clients resend the conversation each turn, so most of it hits. Measured token-weighted over 40.7M tokens; measure your own against the provider's session view | | `assumed_completion_tokens` | 500 | Completion length assumed when pricing a candidate | | `max_energy_per_request` | null | Per-request kWh ceiling — a wall, not a bill. Null disables it | | `plan_kwh_per_period` | 6.25 | **Set this to your own plan's quota.** Reported in `/health` as burn against the allowance; it does not gate anything | This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4). Measurement retired it: turning the cost weight from 0.4 to **zero** changed the winner in only 2 of 6 categories, so the blend was never steering on quality — while 60% of every decision adjudicated fractions of a cent. ### What the proficiency score means now `proficiency.blended_score` is an **expected pass rate on your traffic**, not a raw benchmark quality level. The benchmark (leaderboard + self-eval) provides a prior. Client-reported outcomes from `POST /outcome` provide the per-model signal. The two are combined through empirical-Bayes shrinkage: - A category with no outcome traffic keeps the benchmark score verbatim. - A trafficked category row with no own outcomes inherits a peer prior. - A row with its own outcomes blends the observed rate with that prior; `outcome_prior_strength` (default 20 pseudo-observations) pulls thin data toward the category mean so one lucky sample cannot dominate. So `quality_tolerance` asks "is this candidate's expected real success rate at least 10% better than the cheaper one?" If not, cost breaks the tie. The score is calibrated against the same client reports that `feedback.py` folds in, so routing learns from traffic rather than from the benchmark alone. **Cost is priced per request, from catalog token prices scaled to the request's shape** — prompt size, `assumed_completion_tokens`, `assumed_cache_rate`. It is deliberately *not* a benchmark average. Neuralwatt bills flat per-kWh rather than per-token, so list price is not what gets charged. Scoring used measured billed cost from a fixed 400-token reference sweep for exactly that reason — and that was wrong for real traffic, because the ranking depends on the workload's *shape*, not just the model. On a 400-token prompt one model looked 3.2× cheaper than another; on a realistic 70k-token prompt the same pair inverted and the second was 5.0× cheaper. The provider's attribution ratio moves with prompt size, so a fixed-shape benchmark cannot rank models for a workload of a different shape. List price is still not what is billed, but billing is capped at a multiple of it, so it tracks the real ordering and bounds it — and it is free, needs no sweep, and refreshes whenever the poller runs. **Why cost ≠ eco:** Cost tracks energy (kWh), but carbon is energy × grid intensity. Grid intensity spanned ~49 gCO2/kWh (`FI`) to ~442 (`US-MIDA-PJM`) when measured — and it moves with time of day. The models disagree: `glm-5.2-fast` is 2nd cheapest but 6th cleanest; `kimi-k3-flex` draws 3.7× less energy than `kimi-k2.7-code` while emitting 3.6× more carbon. Collapsing them picks a side. ### Exploration Routing runs an epsilon-greedy exploration pass after ranking. If enabled (`exploration.enabled: true`, default), a random share `ε = 0.03` of eligible requests replaces the winner with the hard-filter-eligible candidate that has the fewest `outcome_samples`, tie-broken by lowest cost. The alternative is skipped if its cost exceeds `max_cost_ratio ×` the winner's cost (default 4×). Only tier 1 and tier 2 requests explore (`max_tier: 2`); tier 3 stays exploitative so frontier work gets the best expected rate. `was_exploration` is persisted to `route_decisions.exploration` so metrics can tell real preference from forced discovery. This is the mechanism that breaks the exposure-bias loop: without it, the router would send most requests to the models already richest in outcome samples and the least-sampled rows would never catch up. ### Incumbency and cache pricing The ranking above treats every candidate as if its prompt were fresh. On a real session it is not: the provider caches the conversation prefix, so the model that served the last turn re-serves it with most of its prompt tokens billed at the cached rate, while a different model pays full price for the entire prefix on its first turn. `assumed_cache_rate` already prices that discount uniformly for every row. Incumbency pricing makes it per-row: the incumbent keeps its measured cache rate, challengers are priced toward cold, and a challenger now has to beat the incumbent by more than the cache it is about to discard, which on a long prompt is most of the prompt. This is the first place the router's own past decision feeds back into its cost model, and it is deliberately narrow: it changes only the cost key, never the sort. Four knobs under `objective:`, all shipping **off** (`incumbent_cache_pricing: false`, dial `null`; enable only after the Wave 1 post-restart baseline day, the Wave 2 gate in `plans/token-waste-waves.md`): | Knob | Default | Purpose | |---|---|---| | `incumbent_cache_pricing` | `false` | The gate; off means byte-identical ranking (pinned by test) | | `incumbent_challenger_cache_rate` | `null` | The challenger dial: `null` follows `assumed_cache_rate`; `0.0` is fully cold; between is a partial penalty | | `incumbent_rate_refresh_seconds` | `300` | TTL on the measured per-(provider, model) rate table | | `incumbent_rate_min_observations` | `25` | Observations before a measured rate is trusted **for pricing**; independent of the warning floor | **What the incumbent is.** The model that served the session's last chat turn, read back from `route_decisions` by `_session_incumbent_lookup` in `dispatcher.py`: the most recent row for the session key with a selected model. The query is an allowlist (`kind IN ('chat')`, never a `!=` denylist), so a one-request `model` pin (passthrough), a `local_dispatch_fallback` row, or a `route` probe can never set the incumbent, and neither can any future `kind` until it argues its way in. The reason is that incumbency is a statement about *preference*: this session, with its own history, chose that model repeatedly on merit. A single pinned request says nothing about the session and a degraded fallback is the opposite of a preference; letting either become sticky would convert an escape hatch into a rut. **The rate ladder.** When the feature is on and an incumbent is present, the incumbent is priced at its measured per-(provider, model) cache rate whenever that series has at least `incumbent_rate_min_observations` observations (default 25); below that trust threshold it falls back to `assumed_cache_rate`. The rates come from `metrics.cache_rate_series` through `_measured_cache_rates()` in `dispatcher.py`, cached for `incumbent_rate_refresh_seconds` (the TTL runs on `time.monotonic()` from the first call after process start, so a brief post-restart cold period is expected). The pricing floor is deliberately a different knob from the alerting floor `cache_rate_warn_min_observations`: pricing would rather fall back to the assumed rate on thin data than steer money on a thin measurement, alerting wants the opposite trade, and coupling the two would couple operator intents that should stay separate. Any failure in the measured-rate path fails open to an empty table and every row prices at `assumed_cache_rate`; incumbency pricing must never block routing. **The challenger dial.** `objective.incumbent_challenger_cache_rate` is the rate every non-incumbent row is priced at. Neutral is decided by **semantic equality**, not identity: `null` follows `assumed_cache_rate`, and any dial equal to `assumed_cache_rate` is the same neutral, so at a neutral setting every row *including the incumbent* prices at `assumed_cache_rate` and the ranking is byte-identical to today's, even with the flag on (config resolves `null` to a concrete float at load; downstream code never sees both representations). `0.0` prices challengers as fully cold prompts, the maximum incumbent advantage; any value between is a partial cache penalty. The whole feature tunes from off to full by moving this one config value, never by a revert: if the switch gap does not narrow, the move is toward cold, then re-measure. **The clamp, and why it is load-bearing.** The challenger rate is `min(dial, incumbent_rate)`, and the invariant it buys is: **the incumbent's cache rate is always >= every challenger's cache rate.** Without it, the dial is denominated in an absolute cache rate, so any dial above the incumbent's measured rate prices the challenger as having a *better* cache than the incumbent: the feature inverts and penalises the incumbent for being the incumbent. Because several routed models measure below the assumed rate (a model measured at ~0.73 against an assumed 0.917, for instance), that inverted zone is not a corner case: for such a model it spans essentially the whole walk from neutral down toward full penalty, which is exactly the range an operator is told to tune through, and from inside it the counterfactual report looks like "the penalty does not pay here" when in fact the penalty is running backwards. Do not remove the clamp as redundant. An alternative dial denominated as a relative penalty (`incumbent_rate * (1 - penalty)`) cannot invert by construction and was deliberately deferred rather than adopted (see the wave's second review, `wave2-review-2.md`); until that redesign exists, the clamp is what makes the absolute form safe, and a property test pins it: for every dial and every measured rate, the incumbent's priced rate is >= every challenger's. **How it composes.** Cache pricing happens inside `estimated_cost`: the ranking loop simply passes each row a different `cache_rate` argument, so every consumer of cost, from the band tiebreak to `cost_score` to the flex-twin swap (a distinct serving endpoint, so it prices as a challenger at the same dial), sees one coherent number per row. `provider_cost_multipliers` composes multiplicatively on top of the cache-adjusted estimate, exactly as before: it inflates the comparison cost inside the quality-band tiebreak and can never override a genuine quality gap, because the quality band is computed before the cost key and the incumbent gets no quality advantage. A model that is genuinely better still wins outright. **Eviction.** The incumbent is a pricing fact, not a reservation. `_resolve_incumbent` re-runs the incumbent's catalog row through `rejection_reason` with this request's actual filters and checks the profile's `restrict_to`: if a hard filter would drop it (the session outgrew its context window, a profile switch restricts the candidate set, a circuit is open), it simply loses incumbency for that turn, no error, and the turn prices exactly as it did before the feature: every row at `assumed_cache_rate`. **Exploration becomes session-scoped.** The coin described above now flips only when no incumbent is present: `route()` adds `incumbent is None` to the exploration condition, and `exploration.py` itself is unchanged (injected RNG, no mutable state, per its module contract). The coin flips at session start only, so one exploratory session costs one cold prompt instead of one per turn; incumbency pricing then pins the explored model for the session's remaining turns, which is tiebreak protection without new state. Re-exploring per turn with an incumbent present would dump the cache every few turns, the exact leak the feature exists to stop (`plans/token-waste-waves.md` item 2.3). **Expiry checks.** Both premises carry their own. The `assumed_cache_rate` fallback is checked by `cache_rate_warnings` in `/metrics`, which compares the measured aggregate and per-(provider, model) rates against the assumed constant and warns beyond `cache_rate_warn_margin`. A deployment whose real rate has diverged is mispriced, not broken, and the fix is to re-measure. The dial's own premise, "the penalty is paying," is evaluated offline rather than assumed: `baseline_report.py --incumbent-challenger ` replays recent `route_decisions` against a cold-challenger ranking and reports what the penalty changed and what other dial settings would have changed, and the operator re-measures the same-model/switched cache-rate gap and billed µ$ per prompt token each re-measurement window after enabling. When the penalty stops paying, the named move is toward neutral on the dial. No hardcoded expiry duration: the check is the measurement. **Shipping state.** Off, as shipped: behavior is byte-identical to pre-feature traffic (pinned by test). When on, every switch decision is explainable post-hoc: the rank debug log carries the incumbent identity, its rate and source (measured or assumed), and the dial; the incumbent row itself carries an `incumbent_pricing` stamp in the ranked output. ### What routing actually returns, and why it moves Sweeping `/route` across every category and tier is the fastest way to see whether your data is doing anything. On the deployment this was written against, 9 categories × 3 tiers currently yields **5 distinct winners** (`qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `kimi-k3`, `kimi-k3-fast`). That number is a diagnostic, not a target, and it is worth knowing what each outcome means: - **One winner everywhere** is a legitimate answer, not a misconfiguration. It happened here: with cost, eco and proficiency all populated, one model was Pareto-dominant — cheapest *and* cleanest in the routable set while scoring within `quality_tolerance` of the best. No defensible weighting picks anything else. If you see this, check whether the leader really is dominant before reaching for the config. - **Winners that change with context size** are the hard filters working. A model is dropped once `required_context_tokens` exceeds its window, so a long session can change model mid-conversation. Past the largest window, `/v1/chat/completions` returns 422 naming the constraint rather than silently truncating. - **Winners that change by category** mean proficiency is live. That is the only category-dependent term, so until the `proficiency` table has data, `task_category` cannot change a decision at all — the classifier computes it and the router pays for it for nothing. The score is now outcome-calibrated, so category-level client reports are the fastest way to shift these choices. The spread here widened for two reasons worth copying: cost became a per-request estimate rather than a benchmark average, and tier stopped being inferred from price alone. Both had been quietly excluding a cheap large-context model from every request above tier 1. Outcome samples are now arriving too; `deepseek-v4-flash` gained substantial samples in `coding_general`, `coding_refactor`, and `general_chat` from post-stream `POST /outcome` reports. If you want a different balance, the levers are `objective.quality_tolerance` (how big a quality gap must be before it outranks cost) and `objective.max_energy_per_request` (a hard ceiling). There is no weight to tune — quality is the objective and cost is the tiebreak, which replaced an earlier weighted blend. ### Circuit Breaker — `circuit_breaker.py` A seventh, dynamic filter sits alongside the six hard filters above: a model that 5xx'd recently is passively excluded from the candidate set. On by default (`circuit_breaker.enabled: true`). It records nothing until a model actually fails — a 5xx from the upstream call marks `(model_id, provider)` down for `initial_cooldown_seconds` (30s default), doubling on each further failure (`backoff_multiplier`, capped at `max_cooldown_seconds`, 600s) and clearing on the next success. Recovery is **passive by design** — no background poller, no health-check loop. The next real request that would have picked the down model becomes its own recovery probe once the cooldown has passed. Two call sites: candidate selection excludes down models outright, and the dispatch retry loop (see the retry budget in [verification.md](verification.md)) records the failure/success on every upstream call and fails over to the next-ranked candidate on a 5xx — but only for `auto`-routed requests, since a pinned request has no alternative to fail over to. `eval_proficiency.py` deliberately calls providers directly, bypassing both the circuit breaker and the retry loop — a transient eval-harness failure against one model's edges shouldn't be able to trip the breaker against real production traffic. ### The classifier's own circuit breaker Separate machinery, same idea, different table: `circuit_breaker.py` guards *upstream models*, while `_last_classifier_failure` in `dispatcher.py` guards the **local classifier** — the one blocking LLM call on the request path. For its first release that timestamp was written by `_record_failure()` and read by nothing: `_classify_cascade` gated only its *cloud* step, and on a different timestamp (`_last_account_refusal`). So the router re-dialled a known-dead local classifier on every request. A *stopped* Ollama refuses the connection immediately and costs little; a **hung** one — or a VPN-bound one that black-holes — costs the full `classifier.timeout_seconds`, 120s on this deployment, per request for as long as the outage lasts. `_classifier_backoff_active()` is the read, consulted in `classify()` **before** `_classifier_client()` is constructed, because constructing it and waiting out the timeout is the expensive part. One detail is load-bearing and easy to undo by accident: recording the failure lives in `_classify_via_local_llm`'s own exception handlers, **not** in `_classify_cascade`. The cascade is walked for reasons other than a fresh failure — a skipped attempt inside an already-open window, or gaming mode — and if those re-stamped the clock then every request during an outage would push the deadline forward and the local classifier would never be re-probed while traffic flowed. That is a permanent outage wearing a circuit breaker's clothes. `tests/test_classifier_backoff.py` pins it directly, and asserts on whether the client was *constructed* rather than on the value returned — a test that only checked the result would pass even if the router had waited out the timeout first. A skip and a real failure both reach the classifier's fallback cascade, but they are logically distinct and `dispatcher.py` keeps them that way: a skip raises `_ClassifierSkipped` (logged once, specifically, as `classify_local_skipped`) while a real failure raises the underlying exception (logged as `fallback`). Conflating the two would tell an operator reading raw logs that the local classifier is failing when gaming mode simply turned it off. ### Which classifier is PRIMARY — `classifier.mode` Separate machinery from the circuit breaker above: that guards *upstream dispatch models* once a category has already been classified. This guards *which implementation does the classifying in the first place*. `classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks the primary attempt. It is a peer to the classifier's own fallback cascade (local → stale session cache → session history → optional `cloud_fallback` → static guess), not a replacement for it — whichever mode is primary, a failure in `_classify_via_configured_mode` still funnels into the exact same, unmodified cascade in `dispatcher.py`. Two things are load-bearing here and easy to get backwards in a future edit: - **`cloud_llm` success records `source="classifier"`, not `"classifier_cloud"`.** The latter string means specifically "the cascade's post-failure backup step fired," and both the `/metrics` degradation-share warning and the outcome-attribution set read it as a *degraded* signal. An intentionally configured cloud primary succeeding is the opposite of degraded, so reusing `"classifier_cloud"` for it would make healthy, on-purpose traffic look like an ongoing local outage. - **`cloud_primary_auto` resolves live, cached, not per-request.** `routing.cheapest_classifier_candidate` reuses `select_candidates` + `estimated_cost` — the same functions real dispatch ranking uses — priced for the classifier's own short-prompt/short-completion call shape rather than the task's. `dispatcher._resolve_auto_classifier` caches the result for `classifier.cooldown_seconds` so this does not add a DB scan to every request's latency floor; the admin portal's GET endpoint calls the resolver fresh on every load instead, since a human loading a settings page is not on that latency floor and a stale "live" reading would be exactly the kind of bug the profiles page's zero-admit badge was. `local_encoder` mode only ever produces `task_category` — `task_tier` falls back to `classifier.fallback_tier`, a documented limitation rather than a second heuristic. See [local-models.md](local-models.md) for why zero-shot rather than fine-tuned (this router never stores raw task text). ### Checking whether scoring earns its complexity `baseline_report.py` automates the "check whether the leader really is dominant" step above: it replays recent `route_decisions` against two trivial counterfactuals (always cheapest, always highest proficiency) using the *current* catalog and proficiency table, and reports how often real scoring picked something a trivial baseline wouldn't have. ```bash PYTHONPATH=src python -m baseline_report --since 2026-08-01 PYTHONPATH=src python -m baseline_report --since 2026-08-01 --category coding_refactor --csv ``` A high dominance share paired with a near-zero proficiency delta against `always_cheapest` means the quality-first ranking isn't earning its complexity for that slice of traffic. Read-only — adds no schema, spends no quota. ### Tiering Tier on `reasoning_default_enabled` (from `metadata.reasoning.default_enabled`, falling back to `capabilities.reasoning`), **not** `supports_reasoning`. `capabilities.reasoning` only means "the endpoint accepts a reasoning param" — it is true for 17 of 19 rows, and tiering on it put 17 models in tier 3 and left tier 1 empty. A `-fast` row does not inherit its sibling's tier 3. Cost is checked before the reasoning rule so $0.28/1M models can reach tier 1. **Cheapness is not a capability ceiling** (`tier1_context_max`, default 512000). Tier is a *floor* — `routing.py` drops any row with `tier < required_tier` — so tier 1 means "simple work only", not "cheap". Deciding that on price alone put `deepseek-v4-flash` in tier 1 for no reason but its $0.28/1M completion price, which excluded it **outright** from every tier-2 request. It has a 1M advertised window and scores 1.00 on all three coding categories. That was the same substitution the cost axis already had to unlearn: price is a market signal, not a capability measurement. Tier 1 now requires the model to be small **and** cheap. The gate reads the **advertised** `context_window`, whose catalog values are the clean market classes — 131056 / 199984 / 262128 / 1048560 — rather than `effective_context_window`, which varies within a class. 512000 sits in the empty band between the 256K and 1M classes with a 2x margin either side, so it is not fitted to any one model. The gate only ever demotes; a huge window never promotes an expensive model into tier 1, and a *missing* window does not block it, since absent evidence should not decide anything. Distribution moved **4 / 6 / 9 -> 1 / 9 / 9** — only the three deepseek rows changed. With cost priced per-request, `deepseek-v4-flash` now wins `coding_general` at every context size (16.5x cheaper than `kimi-k3` at 200k) and is still correctly absent from `tool_use_agentic`, where its measured 0.33 drops it out of the quality band. That is the eval data earning it the slot rather than a thumb on the scale — no `model_tiers` override was needed. ### Why tools and reasoning stay on their existing signals Tools stay on the measured `tool_use_agentic` proficiency gate, not a `supports_tools` flag gate. Every routable catalog row already has `supports_tools = 1`, so a flag gate would be inert. The real signal is the measured proficiency, because the observed failure is a model over-reaching for tools on a non-agentic prompt. `routing.min_tool_proficiency` captures that measurement and only applies when the request carries a `tools` array. Reasoning stays on tiering (`reasoning_default_enabled`), not on a new flag gate. `supports_reasoning` only means the endpoint accepts a reasoning parameter, and that is true for 17 of 19 rows — nearly the whole catalog. `has_reasoning_request` is detected purely for observation. Making it a gate would add no useful filtering, because the decision of whether a request needs reasoning is already encoded in the requested tier. The fail-closed asymmetry, stated plainly: **capability flags fail closed on unknown; quality measurements admit on absent evidence.** A missing `supports_vision` or `supports_json_mode` flag means "cannot confirm", so the model is dropped. A missing `tool_use_agentic` score or energy measurement means "unproven, not bad", so the model is admitted. The first wrong guess is a guaranteed 400; the second is just an empty data point that the neutral default handles. ### The classifier's labels are not the proficiency scoring axis `classifier.candidate_categories` is the set of labels a classifier may return. `proficiency.categories` is the axis models are scored on. They were one list, and the two jobs are not the same job: the axis answers "what is this model good at", the candidate set answers "what should a classifier be asked to distinguish". Config load validates the candidate set as a **subset** of the axis — a label outside it joins against nothing in the `proficiency` table, so routing would rank on NULLs, which is the same phantom-join failure `routing.tool_use_category` is validated against, arriving from the other side. Omitting the key means "every category", i.e. the pre-split behaviour. Both classifier backends read the candidate set: `dispatcher.classify` builds its allowed-values prompt from it, and `local_encoder.classify_zero_shot` receives it as its candidate labels. `_classify_once` also validates the returned label against it, so the exclusion is binding rather than advisory — a local model that ignores the prompt's list is corrected, exactly as an invented label always was. The scoring-axis consumers were deliberately left on the full list: `leaderboard.py` (priors are per scoring category), `eval_proficiency.py` (the benchmark still evaluates every category), `admin.py`'s profile-coverage count and `/metrics`' category list (both describe the scoring axis), and `config.py`'s `tool_use_category` / `eligible_categories` validators. **`tool_use_agentic` is the excluded category, and why is the point.** It describes what a turn mechanically *does* rather than what it is *for*, and every agent turn does it. Measured 2026-09-14 on live traffic with the session cache off so every turn classified for real: **31 of 31 consecutive turns** classified `tool_use_agentic`, all carrying `tools`, all routed to `z-ai/glm-5.3-flash` — 14.8s time-to-first-token, p95 40s, the slowest model in the catalog. Per-turn classification became accurate and routing got worse. Classifying *for* tool use also duplicates a signal the router already has exactly: the request carries a `tools` array, reading it is free, and `routing.min_tool_proficiency` is the mechanism that acts on it. **What the exclusion costs, written down so the next person does not have to rediscover it:** no new `POST /outcome` report can attribute to `tool_use_agentic`, because nothing classifies as it any more. Those scores **freeze** at their current values (`qwen3.6-35b` at 0.902 over 154 samples). The tool filter keeps reading them and the category keeps ranking; it simply stops accumulating. That is accepted for now, not fixed. **Considered and rejected:** attributing outcomes to `tool_use_agentic` whenever the request carried a `tools` array. opencode sends `tools` on essentially every request, so the score would converge on each model's overall pass rate and stop discriminating the one thing the category exists to detect — a model that *over-reaches* for tools on work that did not need them (`deepseek-v4-flash` calling two tools to subtract 1:20pm from 3pm). A frozen honest score beats a live meaningless one. The readout is on the admin portal's Classifier card: the candidate list, and the excluded categories with a note that they stop accumulating outcomes. Read-only on purpose — it is a structural list in the same class as `proficiency.categories` and `models.eligible_categories`, neither of which has an editing surface, and what an operator needs from the portal is the answer to "why does nothing ever classify as X".