643 lines
37 KiB
Markdown
643 lines
37 KiB
Markdown
> Deep dive into routing internals. Back: [README](../README.md).
|
||
|
||
## Quality-First Selection
|
||
|
||
Seven hard filters are applied **before** scoring (not weighted — outright disqualification):
|
||
|
||
1. `effective_context_window ≥ required_context_tokens`
|
||
2. `tier ≥ required_tier` (from classifier)
|
||
3. Serving class compatible with request's `latency_tolerance`; `access_level` reachable
|
||
4. `tool_use_agentic` proficiency ≥ `routing.min_tool_proficiency`, **only when
|
||
the request carries a `tools` array** (ships disabled — `null`)
|
||
5. `supports_vision = 1` when the request carries `image_url` parts; `NULL`
|
||
fails closed (an unknown flag means the capability cannot be confirmed)
|
||
6. `supports_json_mode = 1` when `response_format.type` is `json_object` or
|
||
`json_schema`; `NULL` also fails closed
|
||
7. `eligible_categories` restricts only when set: a row with a non-empty
|
||
category list is admitted only when the task's category is in that list;
|
||
`NULL`/absent means unrestricted (all cloud rows today)
|
||
|
||
Filters 4–6 are read from the request body rather than inferred. A `tools`
|
||
array states whether tool definitions are on the table, `image_url` parts state
|
||
whether vision is needed, and `response_format` states whether JSON mode is
|
||
needed. Local classifiers identified unambiguous tool-use prompts only 1-2
|
||
times in 6, so inferring capability requirements does not work; the request
|
||
states them exactly and for free.
|
||
|
||
The tool gate is a **quality measurement** filter: a model with no measured
|
||
`tool_use_agentic` score is admitted, because "unproven" is not "proven bad".
|
||
Only a measured score below `routing.min_tool_proficiency` is dropped. The
|
||
vision and JSON-mode gates are **capability flag** filters. There every
|
||
routable catalog row has `supports_tools = 1`, so a flag gate would be inert;
|
||
vision and JSON mode are not universal, and an absent or `NULL` flag means
|
||
"cannot confirm the capability", so the model is dropped. A wrong guess on a
|
||
capability flag is a guaranteed provider 400.
|
||
|
||
The tool filter ships **off** (`null`), because agent clients send `tools` on
|
||
nearly every request, so enabling it excludes the cheapest model from ordinary
|
||
agent traffic. Whether that trade is worth it is an empirical question best
|
||
settled with `POST /outcome` data rather than a 3-task benchmark score. Set it
|
||
to `0.5` to turn it on. The vision and JSON-mode gates ship **on**, because a
|
||
wrong guess is a guaranteed 400.
|
||
|
||
### Local Vision Fallback
|
||
|
||
Cloud vision is not universal in the catalog, and the most economical coding
|
||
rows do not declare `supports_vision`. Routing an image request through them
|
||
would earn a provider-side 400, so when `routing.require_vision` is true only
|
||
rows with `supports_vision = 1` survive the hard filters. If **no** cloud
|
||
candidate survives, the router can fall back to a local vision model instead
|
||
of returning 422.
|
||
|
||
`local_vision:` in `config/config.yaml` controls this path:
|
||
|
||
| Key | Default | Purpose |
|
||
|---|---|---|
|
||
| `enabled` | `true` | Whether the fallback runs at all |
|
||
| `base_url` | `http://localhost:11434/v1` | OpenAI-compatible Ollama endpoint |
|
||
| `api_key_env` | `null` | Env var holding an API key, if the endpoint needs one |
|
||
| `model` | `qwen3-vl:4b` | Vision model on that Ollama |
|
||
| `timeout_seconds` | `60` | Request timeout |
|
||
| `max_images` | `4` | Refuse requests with more image parts |
|
||
| `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this |
|
||
|
||
The fallback is **enabled by default** both in `config/config.yaml` and in
|
||
`LocalVisionConfig`, so omitting the section still turns it on. Disable it
|
||
explicitly (`enabled: false`) on a host with no local Ollama or one that has
|
||
not pulled the vision model.
|
||
|
||
When enabled and no cloud candidate is selected, `_run_local_vision` sends the
|
||
original message list — with `image_url` parts intact — to the configured
|
||
Ollama model. The local answer then **replaces** the cloud completion:
|
||
`_local_vision_response` returns a normal OpenAI-shaped response, including a
|
||
stream-wrapped version when the client asked for `stream: true`. It does *not*
|
||
inject a caption into a cloud call, because the streaming proxy cannot rewrite
|
||
bytes mid-stream.
|
||
|
||
Security and budget guards:
|
||
- **Only inline `data:` URIs are accepted.** A remote `http(s)` image URL is
|
||
declined, because pointing a local model at an arbitrary URL would let an
|
||
unauthenticated caller make the router fetch internal resources (SSRF).
|
||
- Image count and total payload size are bounded by `max_images` and
|
||
`max_image_bytes` before the local call is made.
|
||
- If the local call fails for any reason, the request falls through to the
|
||
ordinary `422 No model satisfies the hard filters` rather than returning a
|
||
silent empty response.
|
||
|
||
Pull the model on whichever Ollama the fallback points at:
|
||
|
||
```bash
|
||
ollama pull qwen3-vl:4b
|
||
```
|
||
|
||
Config is strict (`extra="forbid"`): a misspelled or misplaced key fails at
|
||
load instead of being silently ignored.
|
||
|
||
### Local dispatch branch
|
||
|
||
Tasks in a local model's `eligible_categories` (for example `file_summarization`
|
||
and `diff_checking` for the shipped `qwen2.5-coder-router:14b`) survive the hard
|
||
filters the same way cloud rows do. Once selected, the request is sent to the
|
||
configured Ollama endpoint (`provider='ollama-local'`) instead of to
|
||
`dispatch_providers`. A local 502 trips the circuit breaker, so the next request
|
||
sees the local row as excluded and reroutes to a cloud candidate.
|
||
|
||
Local dispatch is **quality-gated and dormant under the default profile**.
|
||
The local row competes in the same ranking as every cloud row and must score
|
||
within `objective.quality_tolerance` of the category's cloud leader before its
|
||
price advantage is even consulted. Measured on 2026-09-02/03,
|
||
`qwen2.5-coder-router:14b` scores **0.767** on `file_summarization` against the
|
||
cloud leader's **0.95** — a 0.183 gap against a 0.10 tolerance — so under the
|
||
default profile the local row is dormant **by design**, today, until a local
|
||
model scores within tolerance. This is not a bug: quality is the objective,
|
||
and the local model has not yet earned the tiebreak (its 9.1x cost advantage
|
||
applies mainly to prompt tokens, so it wins most where large-context
|
||
summarization dominates — and it is assigned to summarization precisely
|
||
because it is not good enough at code).
|
||
|
||
Do NOT widen `objective.quality_tolerance` to fix this: the tolerance
|
||
describes measurement noise (bench sample sizes are small enough that a
|
||
smaller gap is genuinely indistinguishable), not preference. The project's
|
||
direction is named profiles (`auto:batch` is already one — a named widening
|
||
of the candidate set), and a `locality` profile is the right future home for
|
||
a per-category local preference, not a global tolerance change.
|
||
|
||
When every cloud candidate is exhausted, or the upstream returns an
|
||
account-level 4xx (any 4xx except 400/404/422 — those mean the request
|
||
itself is malformed and would fail locally too), an eligible routed
|
||
request degrades to the local dispatch model instead of surfacing the
|
||
cloud error. The trigger is deliberately conservative because the
|
||
upstream's exact credit-exhaustion signature is not yet observed
|
||
(full-body upstream-failure logging is in place to capture it); the
|
||
remaining cloud rows share one account, so no time is wasted walking
|
||
them after an account-level refusal. The fallback respects
|
||
`eligible_categories` even here: for everything else, the cloud error
|
||
fails loudly rather than handing a code task to a summarization-grade
|
||
model. The degraded answer is recorded as `kind='local_dispatch_fallback'`
|
||
in `route_decisions` so the proficiency loop never mistakes it for local
|
||
winning on merit, and if the local call itself fails the original cloud
|
||
error surfaces (the same fail-through contract as the local vision
|
||
fallback). There is no separate config knob — the fallback is armed
|
||
exactly by a `local_dispatch_models` entry listing the category in its
|
||
`eligible_categories`.
|
||
|
||
```
|
||
1. drop candidates whose measured energy exceeds objective.max_energy_per_request
|
||
2. rank by expected pass rate for the task's category (proficiency.blended_score)
|
||
3. treat differences smaller than objective.quality_tolerance as equal
|
||
4. among equals, pick the cheapest
|
||
5. on a small share of eligible requests, explore the least-evidenced candidate
|
||
```
|
||
|
||
| Setting | Value | Notes |
|
||
|---|---|---|
|
||
| `quality_tolerance` | 0.10 | % of real success rate, not an abstract quality unit: a candidate must beat the best by more than this to outrank a cheaper alternative |
|
||
| `assumed_cache_rate` | 0.917 | Share of prompt tokens served from the provider's prefix cache. Agent clients resend the conversation each turn, so most of it hits. Measured token-weighted over 40.7M tokens; measure your own against the provider's session view |
|
||
| `assumed_completion_tokens` | 500 | Completion length assumed when pricing a candidate |
|
||
| `max_energy_per_request` | null | Per-request kWh ceiling — a wall, not a bill. Null disables it |
|
||
| `plan_kwh_per_period` | 6.25 | **Set this to your own plan's quota.** Reported in `/health` as burn against the allowance; it does not gate anything |
|
||
|
||
This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4).
|
||
Measurement retired it: turning the cost weight from 0.4 to **zero** changed
|
||
the winner in only 2 of 6 categories, so the blend was never steering on
|
||
quality — while 60% of every decision adjudicated fractions of a cent.
|
||
|
||
### What the proficiency score means now
|
||
|
||
`proficiency.blended_score` is an **expected pass rate on your traffic**, not a
|
||
raw benchmark quality level. The benchmark (leaderboard + self-eval) provides a
|
||
prior. Client-reported outcomes from `POST /outcome` provide the per-model
|
||
signal. The two are combined through empirical-Bayes shrinkage:
|
||
|
||
- A category with no outcome traffic keeps the benchmark score verbatim.
|
||
- A trafficked category row with no own outcomes inherits a peer prior.
|
||
- A row with its own outcomes blends the observed rate with that prior;
|
||
`outcome_prior_strength` (default 20 pseudo-observations) pulls thin data
|
||
toward the category mean so one lucky sample cannot dominate.
|
||
|
||
So `quality_tolerance` asks "is this candidate's expected real success rate at
|
||
least 10% better than the cheaper one?" If not, cost breaks the tie. The score
|
||
is calibrated against the same client reports that `feedback.py` folds in, so
|
||
routing learns from traffic rather than from the benchmark alone.
|
||
|
||
**Cost is priced per request, from catalog token prices scaled to the
|
||
request's shape** — prompt size, `assumed_completion_tokens`,
|
||
`assumed_cache_rate`. It is deliberately *not* a benchmark average.
|
||
|
||
Neuralwatt bills flat per-kWh rather than per-token, so list price is not what
|
||
gets charged. Scoring used measured billed cost from a fixed 400-token
|
||
reference sweep for exactly that reason — and that was wrong for real traffic,
|
||
because the ranking depends on the workload's *shape*, not just the model. On
|
||
a 400-token prompt one model looked 3.2× cheaper than another; on a realistic
|
||
70k-token prompt the same pair inverted and the second was 5.0× cheaper. The
|
||
provider's attribution ratio moves with prompt size, so a fixed-shape
|
||
benchmark cannot rank models for a workload of a different shape.
|
||
|
||
List price is still not what is billed, but billing is capped at a multiple of
|
||
it, so it tracks the real ordering and bounds it — and it is free, needs no
|
||
sweep, and refreshes whenever the poller runs.
|
||
|
||
**Why cost ≠ eco:** Cost tracks energy (kWh), but carbon is energy × grid
|
||
intensity. Grid intensity spanned ~49 gCO2/kWh (`FI`) to ~442
|
||
(`US-MIDA-PJM`) when measured — and it moves with time of day. The models
|
||
disagree: `glm-5.2-fast` is
|
||
2nd cheapest but 6th cleanest; `kimi-k3-flex` draws 3.7× less energy than
|
||
`kimi-k2.7-code` while emitting 3.6× more carbon. Collapsing them picks a
|
||
side.
|
||
|
||
### Exploration
|
||
|
||
Routing runs an epsilon-greedy exploration pass after ranking. If enabled
|
||
(`exploration.enabled: true`, default), a random share `ε = 0.03` of eligible
|
||
requests replaces the winner with the hard-filter-eligible candidate that has
|
||
the fewest `outcome_samples`, tie-broken by lowest cost. The alternative is
|
||
skipped if its cost exceeds `max_cost_ratio ×` the winner's cost (default 4×).
|
||
Only tier 1 and tier 2 requests explore (`max_tier: 2`); tier 3 stays
|
||
exploitative so frontier work gets the best expected rate.
|
||
|
||
`was_exploration` is persisted to `route_decisions.exploration` so metrics can
|
||
tell real preference from forced discovery. This is the mechanism that breaks
|
||
the exposure-bias loop: without it, the router would send most requests to the
|
||
models already richest in outcome samples and the least-sampled rows would
|
||
never catch up.
|
||
|
||
### Incumbency and cache pricing
|
||
|
||
The ranking above treats every candidate as if its prompt were fresh. On a real
|
||
session it is not: the provider caches the conversation prefix, so the model
|
||
that served the last turn re-serves it with most of its prompt tokens billed at
|
||
the cached rate, while a different model pays full price for the entire prefix
|
||
on its first turn. `assumed_cache_rate` already prices that discount uniformly
|
||
for every row. Incumbency pricing makes it per-row: the incumbent keeps its
|
||
measured cache rate, challengers are priced toward cold, and a challenger now
|
||
has to beat the incumbent by more than the cache it is about to discard, which
|
||
on a long prompt is most of the prompt. This is the first place the
|
||
router's own past decision feeds back into its cost model, and it is
|
||
deliberately narrow: it changes only the cost key, never the sort.
|
||
|
||
The incumbent is **per conversation** when the client sends
|
||
`X-Router-Conversation`: `session_key` is then the namespaced conversation id,
|
||
so each conversation tracks its own incumbent. Without that header the
|
||
incumbent is per system-prompt fingerprint, and the fingerprint's hashing
|
||
merges concurrent same-prompt agents into one incumbent so they share the
|
||
cache discount.
|
||
|
||
Four knobs under `objective:`, all shipping **off**
|
||
(`incumbent_cache_pricing: false`, dial `null`; enable only after the Wave 1
|
||
post-restart baseline day, the Wave 2 gate in `plans/token-waste-waves.md`):
|
||
|
||
| Knob | Default | Purpose |
|
||
|---|---|---|
|
||
| `incumbent_cache_pricing` | `false` | The gate; off means byte-identical ranking (pinned by test) |
|
||
| `incumbent_challenger_cache_rate` | `null` | The challenger dial: `null` follows `assumed_cache_rate`; `0.0` is fully cold; between is a partial penalty |
|
||
| `incumbent_rate_refresh_seconds` | `300` | TTL on the measured per-(provider, model) rate table |
|
||
| `incumbent_rate_min_observations` | `25` | Observations before a measured rate is trusted **for pricing**; independent of the warning floor |
|
||
|
||
**What the incumbent is.** The model that served the session's last chat turn,
|
||
read back from `route_decisions` by `_session_incumbent_lookup` in
|
||
`dispatcher.py`: the most recent row for the session key with a selected model.
|
||
The query is an allowlist (`kind IN ('chat')`, never a `!=` denylist), so a
|
||
one-request `model` pin (passthrough), a `local_dispatch_fallback` row, or a
|
||
`route` probe can never set the incumbent, and neither can any future `kind`
|
||
until it argues its way in. The reason is that incumbency is a statement about
|
||
*preference*: this session, with its own history, chose that model
|
||
repeatedly on merit. A single pinned request says nothing about the session
|
||
and a degraded fallback is the opposite of a preference; letting either become
|
||
sticky would convert an escape hatch into a rut.
|
||
|
||
**The rate ladder.** When the feature is on and an incumbent is present, the
|
||
incumbent is priced at its measured per-(provider, model) cache rate whenever
|
||
that series has at least `incumbent_rate_min_observations` observations
|
||
(default 25); below that trust threshold it falls back to `assumed_cache_rate`.
|
||
The rates come from `metrics.cache_rate_series` through
|
||
`_measured_cache_rates()` in `dispatcher.py`, cached for
|
||
`incumbent_rate_refresh_seconds` (the TTL runs on `time.monotonic()` from the
|
||
first call after process start, so a brief post-restart cold period is
|
||
expected). The pricing floor is deliberately a different knob from the
|
||
alerting floor `cache_rate_warn_min_observations`: pricing would rather fall
|
||
back to the assumed rate on thin data than steer money on a thin measurement,
|
||
alerting wants the opposite trade, and coupling the two would couple operator
|
||
intents that should stay separate. Any failure in the measured-rate path fails
|
||
open to an empty table and every row prices at `assumed_cache_rate`;
|
||
incumbency pricing must never block routing.
|
||
|
||
**The challenger dial.** `objective.incumbent_challenger_cache_rate` is the
|
||
rate every non-incumbent row is priced at. Neutral is decided by **semantic
|
||
equality**, not identity: `null` follows `assumed_cache_rate`, and any dial
|
||
equal to `assumed_cache_rate` is the same neutral, so at a neutral setting
|
||
every row *including the incumbent* prices at `assumed_cache_rate` and the
|
||
ranking is byte-identical to today's, even with the flag on (config resolves
|
||
`null` to a concrete float at load; downstream code never sees both
|
||
representations). `0.0` prices challengers as fully cold prompts, the maximum
|
||
incumbent advantage; any value between is a partial cache penalty. The whole
|
||
feature tunes from off to full by moving this one config value, never by a
|
||
revert: if the switch gap does not narrow, the move is toward cold, then
|
||
re-measure.
|
||
|
||
**The clamp, and why it is load-bearing.** The challenger rate is
|
||
`min(dial, incumbent_rate)`, and the invariant it buys is: **the incumbent's
|
||
cache rate is always >= every challenger's cache rate.** Without it, the dial
|
||
is denominated in an absolute cache rate, so any dial above the incumbent's
|
||
measured rate prices the challenger as having a *better* cache than the
|
||
incumbent: the feature inverts and penalises the incumbent for being the
|
||
incumbent. Because several routed models measure below the assumed rate (a
|
||
model measured at ~0.73 against an assumed 0.917, for instance), that inverted
|
||
zone is not a corner case: for such a model it spans essentially the whole walk
|
||
from neutral down toward full penalty, which is exactly the range an operator
|
||
is told to tune through, and from inside it the counterfactual report looks
|
||
like "the penalty does not pay here" when in fact the penalty is running
|
||
backwards. Do not remove the clamp as redundant. An alternative dial
|
||
denominated as a relative penalty (`incumbent_rate * (1 - penalty)`) cannot
|
||
invert by construction and was deliberately deferred rather than adopted (see
|
||
the wave's second review, `wave2-review-2.md`); until that redesign exists,
|
||
the clamp is what makes the absolute form safe, and a property test pins it:
|
||
for every dial and every measured rate, the incumbent's priced rate is >=
|
||
every challenger's.
|
||
|
||
**How it composes.** Cache pricing happens inside `estimated_cost`: the
|
||
ranking loop simply passes each row a different `cache_rate` argument, so
|
||
every consumer of cost, from the band tiebreak to `cost_score` to the
|
||
flex-twin swap (a distinct serving endpoint, so it prices as a challenger at
|
||
the same dial), sees one coherent number per row. `provider_cost_multipliers`
|
||
composes multiplicatively on top of the cache-adjusted estimate, exactly as
|
||
before: it inflates the comparison cost inside the quality-band tiebreak and
|
||
can never override a genuine quality gap, because the quality band is
|
||
computed before the cost key and the incumbent gets no quality advantage. A
|
||
model that is genuinely better still wins outright.
|
||
|
||
**Eviction.** The incumbent is a pricing fact, not a reservation.
|
||
`_resolve_incumbent` re-runs the incumbent's catalog row through
|
||
`rejection_reason` with this request's actual filters and checks the profile's
|
||
`restrict_to`: if a hard filter would drop it (the session outgrew its
|
||
context window, a profile switch restricts the candidate set, a circuit is
|
||
open), it simply loses incumbency for that turn, no error, and the turn prices
|
||
exactly as it did before the feature: every row at `assumed_cache_rate`.
|
||
|
||
**Exploration becomes session-scoped.** The coin described above now flips
|
||
only when no incumbent is present: `route()` adds `incumbent is None` to the
|
||
exploration condition, and `exploration.py` itself is unchanged (injected RNG,
|
||
no mutable state, per its module contract). The coin flips at session start
|
||
only, so one exploratory session costs one cold prompt instead of one per
|
||
turn; incumbency pricing then pins the explored model for the session's
|
||
remaining turns, which is tiebreak protection without new state. Re-exploring
|
||
per turn with an incumbent present would dump the cache every few turns, the
|
||
exact leak the feature exists to stop (`plans/token-waste-waves.md` item 2.3).
|
||
|
||
**Expiry checks.** Both premises carry their own. The `assumed_cache_rate`
|
||
fallback is checked by `cache_rate_warnings` in `/metrics`, which compares the
|
||
measured aggregate and per-(provider, model) rates against the assumed
|
||
constant and warns beyond `cache_rate_warn_margin`. A deployment whose real
|
||
rate has diverged is mispriced, not broken, and the fix is to re-measure. The
|
||
dial's own premise, "the penalty is paying," is evaluated offline rather than
|
||
assumed: `baseline_report.py --incumbent-challenger <rate>` replays recent
|
||
`route_decisions` against a cold-challenger ranking and reports what the
|
||
penalty changed and what other dial settings would have changed, and the
|
||
operator re-measures the same-model/switched cache-rate gap and billed µ$ per
|
||
prompt token each re-measurement window after enabling. When the penalty
|
||
stops paying, the named move is toward neutral on the dial. No hardcoded
|
||
expiry duration: the check is the measurement.
|
||
|
||
**Shipping state.** Off, as shipped: behavior is byte-identical to
|
||
pre-feature traffic (pinned by test). When on, every switch decision is
|
||
explainable post-hoc: the rank debug log carries the incumbent identity, its
|
||
rate and source (measured or assumed), and the dial; the incumbent row itself
|
||
carries an `incumbent_pricing` stamp in the ranked output.
|
||
|
||
### What routing actually returns, and why it moves
|
||
|
||
Sweeping `/route` across every category and tier is the fastest way to see
|
||
whether your data is doing anything. On the deployment this was written
|
||
against, 9 categories × 3 tiers currently yields **5 distinct winners**
|
||
(`qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `kimi-k3`,
|
||
`kimi-k3-fast`).
|
||
|
||
That number is a diagnostic, not a target, and it is worth knowing what each
|
||
outcome means:
|
||
|
||
- **One winner everywhere** is a legitimate answer, not a misconfiguration.
|
||
It happened here: with cost, eco and proficiency all populated, one model
|
||
was Pareto-dominant — cheapest *and* cleanest in the routable set while
|
||
scoring within `quality_tolerance` of the best. No defensible weighting
|
||
picks anything else. If you see this, check whether the leader really is
|
||
dominant before reaching for the config.
|
||
- **Winners that change with context size** are the hard filters working.
|
||
A model is dropped once `required_context_tokens` exceeds its window, so a
|
||
long session can change model mid-conversation. Past the largest window,
|
||
`/v1/chat/completions` returns 422 naming the constraint rather than
|
||
silently truncating.
|
||
- **Winners that change by category** mean proficiency is live. That is the
|
||
only category-dependent term, so until the `proficiency` table has data,
|
||
`task_category` cannot change a decision at all — the classifier computes
|
||
it and the router pays for it for nothing. The score is now
|
||
outcome-calibrated, so category-level client reports are the fastest way to
|
||
shift these choices.
|
||
|
||
The spread here widened for two reasons worth copying: cost became a
|
||
per-request estimate rather than a benchmark average, and tier stopped being
|
||
inferred from price alone. Both had been quietly excluding a cheap
|
||
large-context model from every request above tier 1. Outcome samples are now
|
||
arriving too; `deepseek-v4-flash` gained substantial samples in
|
||
`coding_general`, `coding_refactor`, and `general_chat` from post-stream
|
||
`POST /outcome` reports.
|
||
|
||
If you want a different balance, the levers are `objective.quality_tolerance`
|
||
(how big a quality gap must be before it outranks cost) and
|
||
`objective.max_energy_per_request` (a hard ceiling). There is no weight to
|
||
tune — quality is the objective and cost is the tiebreak, which replaced an
|
||
earlier weighted blend.
|
||
|
||
### Circuit Breaker — `circuit_breaker.py`
|
||
|
||
A seventh, dynamic filter sits alongside the six hard filters above: a model
|
||
that 5xx'd recently is passively excluded from the candidate set. On by
|
||
default (`circuit_breaker.enabled: true`).
|
||
|
||
It records nothing until a model actually fails — a 5xx from the upstream
|
||
call marks `(model_id, provider)` down for `initial_cooldown_seconds` (30s
|
||
default), doubling on each further failure (`backoff_multiplier`, capped at
|
||
`max_cooldown_seconds`, 600s) and clearing on the next success. Recovery is
|
||
**passive by design** — no background poller, no health-check loop. The next
|
||
real request that would have picked the down model becomes its own recovery
|
||
probe once the cooldown has passed.
|
||
|
||
Two call sites: candidate selection excludes down models outright, and the
|
||
dispatch retry loop (see the retry budget in [verification.md](verification.md))
|
||
records the failure/success on every upstream call and fails over to the
|
||
next-ranked candidate on a 5xx — but only for `auto`-routed requests, since a
|
||
pinned request has no alternative to fail over to.
|
||
|
||
`eval_proficiency.py` deliberately calls providers directly, bypassing both
|
||
the circuit breaker and the retry loop — a transient eval-harness failure
|
||
against one model's edges shouldn't be able to trip the breaker against real
|
||
production traffic.
|
||
|
||
### The classifier's own circuit breaker
|
||
|
||
Separate machinery, same idea, different table: `circuit_breaker.py` guards
|
||
*upstream models*, while `_last_classifier_failure` in `dispatcher.py` guards
|
||
the **local classifier** — the one blocking LLM call on the request path.
|
||
|
||
For its first release that timestamp was written by `_record_failure()` and
|
||
read by nothing: `_classify_cascade` gated only its *cloud* step, and on a
|
||
different timestamp (`_last_account_refusal`). So the router re-dialled a
|
||
known-dead local classifier on every request. A *stopped* Ollama refuses the
|
||
connection immediately and costs little; a **hung** one — or a VPN-bound one
|
||
that black-holes — costs the full `classifier.timeout_seconds`, 120s on this
|
||
deployment, per request for as long as the outage lasts.
|
||
|
||
`_classifier_backoff_active()` is the read, consulted in `classify()`
|
||
**before** `_classifier_client()` is constructed, because constructing it and
|
||
waiting out the timeout is the expensive part.
|
||
|
||
One detail is load-bearing and easy to undo by accident: recording the failure
|
||
lives in `_classify_via_local_llm`'s own exception handlers, **not** in
|
||
`_classify_cascade`. The cascade is walked for reasons other than a fresh
|
||
failure — a skipped attempt inside an already-open window, or gaming mode —
|
||
and if those re-stamped the clock then every request during an outage would
|
||
push the deadline forward and the local classifier would never be re-probed
|
||
while traffic flowed. That is a permanent outage wearing a circuit breaker's
|
||
clothes. `tests/test_classifier_backoff.py` pins it directly, and asserts on
|
||
whether the client was *constructed* rather than on the value returned — a
|
||
test that only checked the result would pass even if the router had waited
|
||
out the timeout first.
|
||
|
||
A skip and a real failure both reach the classifier's fallback cascade, but
|
||
they are logically distinct and `dispatcher.py` keeps them that way: a skip
|
||
raises `_ClassifierSkipped` (logged once, specifically, as
|
||
`classify_local_skipped`) while a real failure raises the underlying
|
||
exception (logged as `fallback`). Conflating the two would tell an operator
|
||
reading raw logs that the local classifier is failing when gaming mode simply
|
||
turned it off.
|
||
|
||
### Which classifier is PRIMARY — `classifier.mode`
|
||
|
||
Separate machinery from the circuit breaker above: that guards *upstream
|
||
dispatch models* once a category has already been classified. This guards
|
||
*which implementation does the classifying in the first place*.
|
||
|
||
`classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks
|
||
the primary attempt. It is a peer to the classifier's own fallback cascade
|
||
(local → stale session cache → session history → optional `cloud_fallback`
|
||
→ static guess), not a replacement for it — whichever mode is primary, a
|
||
failure in `_classify_via_configured_mode` still funnels into the exact
|
||
same, unmodified cascade in `dispatcher.py`.
|
||
|
||
Two things are load-bearing here and easy to get backwards in a future edit:
|
||
|
||
- **`cloud_llm` success records `source="classifier"`, not
|
||
`"classifier_cloud"`.** The latter string means specifically "the
|
||
cascade's post-failure backup step fired," and both the `/metrics`
|
||
degradation-share warning and the outcome-attribution set read it as a
|
||
*degraded* signal. An intentionally configured cloud primary succeeding
|
||
is the opposite of degraded, so reusing `"classifier_cloud"` for it would
|
||
make healthy, on-purpose traffic look like an ongoing local outage.
|
||
- **`cloud_primary_auto` resolves live, cached, not per-request.**
|
||
`routing.cheapest_classifier_candidate` reuses `select_candidates` +
|
||
`estimated_cost` — the same functions real dispatch ranking uses — priced
|
||
for the classifier's own short-prompt/short-completion call shape rather
|
||
than the task's. `dispatcher._resolve_auto_classifier` caches the result
|
||
for `classifier.cooldown_seconds` so this does not add a DB scan to every
|
||
request's latency floor; the admin portal's GET endpoint calls the
|
||
resolver fresh on every load instead, since a human loading a settings
|
||
page is not on that latency floor and a stale "live" reading would be
|
||
exactly the kind of bug the profiles page's zero-admit badge was.
|
||
|
||
`local_encoder` mode only ever produces `task_category` — `task_tier` falls
|
||
back to `classifier.fallback_tier`, a documented limitation rather than a
|
||
second heuristic. See [local-models.md](local-models.md) for why zero-shot
|
||
rather than fine-tuned (this router never stores raw task text).
|
||
|
||
### Checking whether scoring earns its complexity
|
||
|
||
`baseline_report.py` automates the "check whether the leader really is
|
||
dominant" step above: it replays recent `route_decisions` against two trivial
|
||
counterfactuals (always cheapest, always highest proficiency) using the
|
||
*current* catalog and proficiency table, and reports how often real scoring
|
||
picked something a trivial baseline wouldn't have.
|
||
|
||
```bash
|
||
PYTHONPATH=src python -m baseline_report --since 2026-08-01
|
||
PYTHONPATH=src python -m baseline_report --since 2026-08-01 --category coding_refactor --csv
|
||
```
|
||
|
||
A high dominance share paired with a near-zero proficiency delta against
|
||
`always_cheapest` means the quality-first ranking isn't earning its complexity
|
||
for that slice of traffic. Read-only — adds no schema, spends no quota.
|
||
|
||
### Tiering
|
||
|
||
Tier on `reasoning_default_enabled` (from `metadata.reasoning.default_enabled`,
|
||
falling back to `capabilities.reasoning`), **not** `supports_reasoning`.
|
||
`capabilities.reasoning` only means "the endpoint accepts a reasoning
|
||
param" — it is true for 17 of 19 rows, and tiering on it put 17 models in
|
||
tier 3 and left tier 1 empty. A `-fast` row does not inherit its sibling's
|
||
tier 3. Cost is checked before the reasoning rule so $0.28/1M models can
|
||
reach tier 1.
|
||
|
||
**Cheapness is not a capability ceiling** (`tier1_context_max`, default
|
||
512000). Tier is a *floor* — `routing.py` drops any row with
|
||
`tier < required_tier` — so tier 1 means "simple work only", not "cheap".
|
||
Deciding that on price alone put `deepseek-v4-flash` in tier 1 for no reason
|
||
but its $0.28/1M completion price, which excluded it **outright** from every
|
||
tier-2 request. It has a 1M advertised window and scores 1.00 on all three
|
||
coding categories. That was the same substitution the cost axis already had
|
||
to unlearn: price is a market signal, not a capability measurement.
|
||
|
||
Tier 1 now requires the model to be small **and** cheap. The gate reads the
|
||
**advertised** `context_window`, whose catalog values are the clean market
|
||
classes — 131056 / 199984 / 262128 / 1048560 — rather than
|
||
`effective_context_window`, which varies within a class. 512000 sits in the
|
||
empty band between the 256K and 1M classes with a 2x margin either side, so it
|
||
is not fitted to any one model. The gate only ever demotes; a huge window
|
||
never promotes an expensive model into tier 1, and a *missing* window does not
|
||
block it, since absent evidence should not decide anything.
|
||
|
||
Distribution moved **4 / 6 / 9 -> 1 / 9 / 9** — only the three deepseek rows
|
||
changed. With cost priced per-request, `deepseek-v4-flash` now wins
|
||
`coding_general` at every context size (16.5x cheaper than `kimi-k3` at 200k)
|
||
and is still correctly absent from `tool_use_agentic`, where its measured 0.33
|
||
drops it out of the quality band. That is the eval data earning it the slot
|
||
rather than a thumb on the scale — no `model_tiers` override was needed.
|
||
|
||
### Why tools and reasoning stay on their existing signals
|
||
|
||
Tools stay on the measured `tool_use_agentic` proficiency gate, not a
|
||
`supports_tools` flag gate. Every routable catalog row already has
|
||
`supports_tools = 1`, so a flag gate would be inert. The real signal is the
|
||
measured proficiency, because the observed failure is a model over-reaching for
|
||
tools on a non-agentic prompt. `routing.min_tool_proficiency` captures that
|
||
measurement and only applies when the request carries a `tools` array.
|
||
|
||
Reasoning stays on tiering (`reasoning_default_enabled`), not on a new flag gate.
|
||
`supports_reasoning` only means the endpoint accepts a reasoning parameter, and
|
||
that is true for 17 of 19 rows — nearly the whole catalog. `has_reasoning_request`
|
||
is detected purely for observation. Making it a gate would add no useful
|
||
filtering, because the decision of whether a request needs reasoning is already
|
||
encoded in the requested tier.
|
||
|
||
The fail-closed asymmetry, stated plainly: **capability flags fail closed on
|
||
unknown; quality measurements admit on absent evidence.** A missing
|
||
`supports_vision` or `supports_json_mode` flag means "cannot confirm", so the
|
||
model is dropped. A missing `tool_use_agentic` score or energy measurement means
|
||
"unproven, not bad", so the model is admitted. The first wrong guess is a
|
||
guaranteed 400; the second is just an empty data point that the neutral default
|
||
handles.
|
||
|
||
### The classifier's labels are not the proficiency scoring axis
|
||
|
||
`classifier.candidate_categories` is the set of labels a classifier may
|
||
return. `proficiency.categories` is the axis models are scored on. They were
|
||
one list, and the two jobs are not the same job: the axis answers "what is
|
||
this model good at", the candidate set answers "what should a classifier be
|
||
asked to distinguish". Config load validates the candidate set as a **subset**
|
||
of the axis — a label outside it joins against nothing in the `proficiency`
|
||
table, so routing would rank on NULLs, which is the same phantom-join failure
|
||
`routing.tool_use_category` is validated against, arriving from the other
|
||
side. Omitting the key means "every category", i.e. the pre-split behaviour.
|
||
|
||
Both classifier backends read the candidate set: `dispatcher.classify` builds
|
||
its allowed-values prompt from it, and `local_encoder.classify_zero_shot`
|
||
receives it as its candidate labels. `_classify_once` also validates the
|
||
returned label against it, so the exclusion is binding rather than advisory —
|
||
a local model that ignores the prompt's list is corrected, exactly as an
|
||
invented label always was. The scoring-axis consumers were deliberately left
|
||
on the full list: `leaderboard.py` (priors are per scoring category),
|
||
`eval_proficiency.py` (the benchmark still evaluates every category),
|
||
`admin.py`'s profile-coverage count and `/metrics`' category list (both
|
||
describe the scoring axis), and `config.py`'s `tool_use_category` /
|
||
`eligible_categories` validators.
|
||
|
||
**`tool_use_agentic` is the excluded category, and why is the point.** It
|
||
describes what a turn mechanically *does* rather than what it is *for*, and
|
||
every agent turn does it. Measured 2026-09-14 on live traffic with the session
|
||
cache off so every turn classified for real: **31 of 31 consecutive turns**
|
||
classified `tool_use_agentic`, all carrying `tools`, all routed to
|
||
`z-ai/glm-5.3-flash` — 14.8s time-to-first-token, p95 40s, the slowest model in
|
||
the catalog. Per-turn classification became accurate and routing got worse.
|
||
Classifying *for* tool use also duplicates a signal the router already has
|
||
exactly: the request carries a `tools` array, reading it is free, and
|
||
`routing.min_tool_proficiency` is the mechanism that acts on it.
|
||
|
||
**What the exclusion costs, written down so the next person does not have to
|
||
rediscover it:** no new `POST /outcome` report can attribute to
|
||
`tool_use_agentic`, because nothing classifies as it any more. Those scores
|
||
**freeze** at their current values (`qwen3.6-35b` at 0.902 over 154 samples).
|
||
The tool filter keeps reading them and the category keeps ranking; it simply
|
||
stops accumulating. That is accepted for now, not fixed.
|
||
|
||
**Considered and rejected:** attributing outcomes to `tool_use_agentic`
|
||
whenever the request carried a `tools` array. opencode sends `tools` on
|
||
essentially every request, so the score would converge on each model's overall
|
||
pass rate and stop discriminating the one thing the category exists to
|
||
detect — a model that *over-reaches* for tools on work that did not need them
|
||
(`deepseek-v4-flash` calling two tools to subtract 1:20pm from 3pm). A frozen
|
||
honest score beats a live meaningless one.
|
||
|
||
The readout is on the admin portal's Classifier card: the candidate list, and
|
||
the excluded categories with a note that they stop accumulating outcomes.
|
||
Read-only on purpose — it is a structural list in the same class as
|
||
`proficiency.categories` and `models.eligible_categories`, neither of which
|
||
has an editing surface, and what an operator needs from the portal is the
|
||
answer to "why does nothing ever classify as X".
|
||
|