Reverts2d145d6, which appended "CLIREF.md updated to CLAUDE.md for current project state mapping" to the end of CLAUDE.md. It is a stray sentence describing a rename that never happened. Correction forb774d55: its subject and body say "CLIREF.md". No such file exists in this repository; the 17 lines that commit describes (the North Star 1 classifier-coverage paragraphs) are in CLAUDE.md. History is not rewritten. Also rewords the first of those paragraphs: "partially covered" understated it. Every classifier scalar is now a control, card-backed, or a recorded excuse, and fallback_category / max_input_chars are persisted-only. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
1413 lines
84 KiB
Markdown
1413 lines
84 KiB
Markdown
# Local LLM Model Router — project brief
|
||
|
||
`README.md` is the concise front door. Deep-dive module-by-module reference
|
||
docs live in **`docs/`**.
|
||
`design/local-llm-model-router.md` holds the architecture and rationale,
|
||
including parts still unbuilt. This file is the working state + immediate
|
||
next steps, and is the one to trust on what is currently true.
|
||
|
||
## NORTH STAR GUIDELINES
|
||
|
||
Four rules that outrank local cleverness. Each exists because it was broken
|
||
first and the breakage was expensive to find. When a change conflicts with one
|
||
of these, the change is wrong — not the rule.
|
||
|
||
### 1. Every config knob is reachable from the admin page
|
||
|
||
Every scalar knob in `config.yaml` gets an admin control — runtime, persisted,
|
||
or both as its mechanism warrants — **or** a recorded decision saying why it
|
||
must not. The absence of a control has to be a decision someone made, never an
|
||
oversight nobody noticed.
|
||
|
||
`objective.credit_attenuation.enabled` is the model exception: deliberately
|
||
absent, because enabling it must be a config edit plus a restart. That is a
|
||
recorded choice, not a gap.
|
||
|
||
The classifier section is now fully accounted for under the gate. Seven
|
||
classifier scalars (`context_framing`, `cooldown_seconds`, `fallback_tier`,
|
||
`fallback_category`, `max_input_chars`, `degraded_warn_min`,
|
||
`degraded_warn_threshold`) have a runtime control, a persisted control, or both
|
||
(`fallback_category` and `max_input_chars` are persisted-only). Every other
|
||
classifier scalar is either card-backed (`_CARD_BACKED_PATHS` in
|
||
`src/admin.py`: mode, cloud primary/fallback, encoder, and local-decision
|
||
fields) or excused in `DELIBERATELY_NOT_IN_ADMIN` (timeout, temperature,
|
||
`max_output_tokens`, `encoder.tier_from_features`).
|
||
|
||
Outside-gate knobs are still pending. Sections not yet reached by the portal
|
||
(database, dispatch providers/settings, local dispatch models, profiles,
|
||
local energy, tiers/tiering/proficiency/context, and deployment wiring inside
|
||
classifier/verification/local vision) are tracked in
|
||
`plans/deferred-knobs.md`. The first control added under any of those sections
|
||
will drag every scalar under it into scope at once, per the coverage test's
|
||
clause 1.
|
||
|
||
**Why:** Wave 2 shipped `incumbent_cache_pricing` and
|
||
`incumbent_challenger_cache_rate` with no control at all. Nobody decided that;
|
||
it just never came up. The dial's entire purpose is tuning from neutral to full
|
||
*without reverting code* — and hand-editing a tracked file and restarting is a
|
||
loop nobody walks, so the knob's rationale evaporated on contact with reality.
|
||
|
||
Enforced by `tests/test_admin_knob_coverage.py`, which fails naming the knob.
|
||
Honour it; do not add a `DELIBERATELY_NOT_IN_ADMIN` entry to silence it unless
|
||
the reason is true.
|
||
|
||
### 2. Fix classifier latency by improving the classifier, never by reusing a
|
||
stale decision
|
||
|
||
Classification latency is a real problem and the answer is a faster or better
|
||
classifier — a smaller model, a warmer process, a cheaper backend. The answer
|
||
is **never** to let one classification stand in for later, different work.
|
||
|
||
A stale label does not merely add noise. It silently redirects money: the whole
|
||
point of this router is sending each task to the model that fits *that task*,
|
||
and a replayed label routes the next task to whatever fitted the last one.
|
||
|
||
**Why:** the session cache reached **96.6% of all classifications**. One real
|
||
classification drove **107 consecutive turns** across 14 minutes on 2026-09-16.
|
||
It was introduced to avoid paying a 2-5s round trip per turn, which was a
|
||
reasonable trade in isolation and became the dominant path without anyone
|
||
choosing that.
|
||
|
||
Two costs, and the second is worse than the first:
|
||
|
||
- **Routing** decides on what the session was doing when it started, not what
|
||
this turn is. A mid-session pivot — docs, then bug fixes — routes the bug
|
||
fixes on the docs label.
|
||
- **Proficiency** is trained on those labels, because `POST /outcome`
|
||
attributes to `(model, task_category)`. That is part of how
|
||
`deepseek-v4-flash` came to hold a saturated 1.000 on `file_summarization`
|
||
while failing 50% of them on real traffic.
|
||
|
||
It also made a working classifier look broken: replaying a handful of session
|
||
labels across hundreds of turns produced "pages and pages of `diff_checking`",
|
||
which read as a classifier stuck on one label when the underlying
|
||
classifications were a reasonable mix.
|
||
|
||
If latency forces a cache, that is a **measured, time-boxed concession with an
|
||
expiry condition written next to it** — not a default.
|
||
|
||
### 3. Check worktrees and other agents' work before touching any file
|
||
|
||
Before editing, run `git worktree list` and check for other agents or sessions
|
||
working in the same tree. Do not assume a worktree is yours. When parallel work
|
||
is unavoidable, isolate it — a separate worktree per agent — and commit by
|
||
explicit path, never `git add -A`.
|
||
|
||
**Why:** two agents were once launched into the *same* worktree while each was
|
||
told the other was elsewhere. It survived only because the files were disjoint
|
||
and one of them committed by explicit path rather than sweeping the tree. The
|
||
second agent's test count was also measured against a tree containing the
|
||
first's uncommitted work, so its "verified green" was not a clean signal.
|
||
|
||
Related: a `router.db` inside a worktree is a **stale copy**, not the live
|
||
database. The live one is `/home/alee/Sources/6krrt/router.db`. Check mtimes
|
||
before measuring anything, and open it read-only:
|
||
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
|
||
not accept `-uri`.
|
||
|
||
### 4. Waste is surfaced in the admin portal, and stopping it is one click
|
||
|
||
When the router can see money being wasted, it shows the operator where they
|
||
already look, with the evidence and the lever to stop it side by side. A
|
||
detector that only writes a log line, or a fix that needs a config edit, a
|
||
restart or a long table scan, has not met this rule.
|
||
|
||
"Waste" here means **spend with no concrete change landing**: an agent session
|
||
looping, re-reading, or retrying the same failure, and a model that keeps
|
||
producing such sessions. It does **not** mean steady spend. A healthy agent run
|
||
can burn for hours, and a spend-rate alarm cannot tell the two apart.
|
||
|
||
In practice:
|
||
- **Stalled sessions are visible** in the portal with their evidence: turns and
|
||
$ since the last landed change, and the top repeated target. They also reach
|
||
the operator when nobody is watching (desktop alert).
|
||
- **A model that keeps producing them is visible** as a per-model rollup, with
|
||
**Block model** beside the evidence. The block records its reason, shows in a
|
||
Blocked list, and is one click to undo.
|
||
- **Automatic responses come after the visible one,** never instead of it, and
|
||
every automatic action shows up in the same place.
|
||
|
||
**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on
|
||
2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec
|
||
68 times its length, and a model confabulated truncation that was not there.
|
||
Every existing check stayed green. It was caught only by a human, or a Claude
|
||
session, reading opencode's session store by hand. Pulling the model took a trip
|
||
through a dropdown whose vocabulary is catalog `deprecated`. The detection
|
||
existed nowhere, and the lever existed only for someone who already knew.
|
||
|
||
## What this is
|
||
|
||
A router that uses a local model (served via Ollama) to classify incoming
|
||
coding/documentation tasks — category, tier, required context size — and
|
||
dispatch each task to the best-fit open-weight model on **Neuralwatt Cloud**,
|
||
under a per-request cost ceiling, tiebroken by price, and ranked by
|
||
category-level expected pass rate.
|
||
|
||
Every measurement in this file was taken on one deployment against one
|
||
provider account. They are recorded because the reasoning is worth more than
|
||
the conclusion, but treat them as observations with a date on them, not as
|
||
constants — the catalog, prices, grid intensity and pool load all move. When
|
||
a number here decides something, re-run the measurement before trusting it.
|
||
|
||
Neuralwatt remains the primary provider. OpenRouter is back as an opt-in
|
||
provider gated by the `provider_model_allowlist` table — a provider with
|
||
`require_allowlist: true` is ignored unless its model_id is explicitly
|
||
allowlisted, and any previously-upserted row that drops off the list is
|
||
marked `deprecated` on the next poll. The `provider` column and the
|
||
`(model_id, provider)` primary key are unchanged, so this adds no migration.
|
||
|
||
## Stack
|
||
|
||
- Python (chosen over Rust — this is I/O-bound against provider APIs, not
|
||
CPU-bound; iteration speed on the scoring/weighting logic matters more than
|
||
raw execution speed at this scale)
|
||
- SQLite for the decision table
|
||
- Ollama for local classification, via any OpenAI-compatible endpoint —
|
||
`localhost:11434/v1`, or an Ollama on another machine across a VPN
|
||
(`classifier.base_url`)
|
||
- FastAPI for the dispatcher service
|
||
|
||
## Billing is per-kWh, not per-token — and neither is what scoring uses
|
||
|
||
**Measured against the live API, 2026-08-11.** Neuralwatt bills a flat
|
||
**$8.00 per kWh** and the catalog's `input_per_million` /
|
||
`output_per_million` prices are not what this account is charged. Confirmed
|
||
across five models; `cost_usd / energy_kwh` came back 8.00 every time:
|
||
|
||
| model | list $/1M out | completion tokens | billed USD | kWh | $/kWh |
|
||
|---|---|---|---|---|---|
|
||
| deepseek-v4-flash | 0.28 | 600 | 5.00e-06 | 5.88e-07 | 8.50\* |
|
||
| gemma-4-31b | 0.42 | 540 | 3.80e-04 | 4.75e-05 | 8.00 |
|
||
| qwen3.6-35b-fast | 1.15 | 600 | 2.53e-04 | 3.16e-05 | 8.01 |
|
||
| kimi-k2.7-code-fast | 4.00 | 600 | 2.25e-03 | 2.81e-04 | 8.00 |
|
||
| kimi-k3-fast | 15.00 | 539 | 2.17e-04 | 2.71e-05 | 8.00 |
|
||
|
||
\* rounding — billed cost is quantized to ~1e-06.
|
||
|
||
The precise rule, validated against all 65 samples of the reference sweep
|
||
(61/65 within 2%; the 4 outliers are microdollar rounding, not misses):
|
||
|
||
```
|
||
cost_usd = min( $8.00/kWh x energy_kwh , 3 x list token price )
|
||
```
|
||
|
||
The ceiling bound in only 2 of 65 samples, both `deepseek-v4-flash` energy
|
||
spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x
|
||
$0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349.
|
||
|
||
**List price ranks models backwards.** Not approximately — invertedly:
|
||
|
||
| | list $/1M | actually billed | gCO2eq |
|
||
|---|---|---|---|
|
||
| `deepseek-v4-flash` | **0.28** | 9.80e-05 | 8.90e-04 |
|
||
| `gemma-4-31b` | 0.42 | **5.02e-05** | **2.32e-04** |
|
||
|
||
deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon.
|
||
|
||
**But cost and eco are NOT the same axis** — the tempting simplification, and
|
||
it is wrong. Cost tracks energy, but carbon is energy x the serving region's
|
||
grid intensity, and models run in different regions:
|
||
|
||
| grid | gCO2/kWh | models |
|
||
|---|---|---|
|
||
| `FI` | ~49-50 | most of the catalog; varies by time |
|
||
| `FI` (reported) | 475 | `glm-5.2-fast`, `glm-5.2-flex` |
|
||
| `US-MIDA-PJM` | ~442 | the `kimi-k3` family |
|
||
|
||
A 13.6x spread, so the two axes disagree: `glm-5.2-fast` is the 2nd cheapest
|
||
model and only the 6th cleanest; `kimi-k3-flex` draws 3.7x *less* energy than
|
||
`kimi-k2.7-code` while emitting 3.6x *more* carbon. Weighting them separately
|
||
is load-bearing, and `tests/test_routing.py` pins it.
|
||
|
||
**Superseded — cost no longer comes from the sweep at all.** `cost` was the
|
||
*median* measured USD over the reference sweep. That was measured to be WRONG
|
||
for real traffic, because the reference workload is the wrong shape.
|
||
|
||
The sweep sends a 400-token prompt with a 400-token completion. Real agent
|
||
traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache
|
||
hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the
|
||
figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts:
|
||
|
||
| workload | winner |
|
||
|---|---|
|
||
| reference sweep (400/400) | `glm-5.2-fast`, 3.2x cheaper |
|
||
| realistic (70k prompt, short answer) | **`deepseek-v4-flash`, 5.0x cheaper** |
|
||
|
||
Same two models, opposite answer. `glm-5.2-fast` sits at attribution 0.006 on
|
||
a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts
|
||
and badly on real ones. `deepseek-v4-flash` barely moves (0.21 -> 0.25).
|
||
|
||
So `routing.estimated_cost` prices each request from **catalog token prices,
|
||
scaled to that request's actual shape** (prompt size, assumed completion
|
||
length, `assumed_cache_rate`). List price is not what gets billed, but billing
|
||
is capped at 3x list, so it tracks the real ordering and bounds it — and on
|
||
the one case that was checked live it agrees with the measurement in direction
|
||
and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no
|
||
sweep, and refreshes whenever the poller runs.
|
||
|
||
`objective.plan_kwh_per_period` is a planning figure only: per-request traffic
|
||
is **never** refused for exceeding it — it **gates nothing**. Overage is billed
|
||
against the account's credit balance (`allowance_remaining_usd` from the
|
||
provider). The `/admin/api/snapshot` endpoint reports balance, estimated burn
|
||
rate, and projected runway **per provider** inside `quota.accounts[]` — a
|
||
list of per-provider billing shapes (`metered_plan`, `prepaid_credit`,
|
||
`self_hosted`, or `unmetered`) with `plan`, `pool`, `burn`, and `credit`
|
||
blocks as appropriate. `quota.spend` aggregates provider spend and a
|
||
list-price estimate. The old flat keys and `by_provider`/`total_balance_usd`
|
||
shape were removed; every consumer was updated in the same change, so there
|
||
are no deprecated aliases.
|
||
|
||
Three signals said `deepseek-v4-flash` — catalog token price (7.8x cheaper),
|
||
NeuralWatt's own published per-request energy (~10x lower), and a live 70k
|
||
measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the
|
||
workload you actually run.
|
||
|
||
`eco` still comes from the sweep's median gCO2eq, and is still not an
|
||
objective. `flex_cost_multiplier` is gone: a flex row's measured cost already
|
||
is its flex cost.
|
||
|
||
**Open, and worth knowing:** NeuralWatt's model cards publish *gross* energy
|
||
(~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is
|
||
gross x attribution. GLM burns roughly 7x more actual electricity per request
|
||
and charges ~5x less, because far more tenants share its GPUs. Anything built
|
||
on `eco` inherits that inversion — the attributed carbon figure answers "what
|
||
is my share", not "what was burned".
|
||
|
||
## Energy attribution: signal that looks like noise
|
||
|
||
Billed energy decomposes exactly:
|
||
|
||
```
|
||
energy_kwh = avg_power_watts x duration_seconds x attribution_ratio
|
||
```
|
||
|
||
`attribution_ratio` is the request's share of a shared multi-tenant GPU pool.
|
||
Up close it looks like pure noise — eight rapid identical calls to one model
|
||
spanned 20x in billed energy, correlating **+0.997** with the ratio while
|
||
power and duration held steady. Two sweeps of the same 13 models with the
|
||
same prompt disagreed by up to 36x.
|
||
|
||
Scoring on the pre-attribution product (`power x duration`) was tried, and it
|
||
is **wrong**. Across the sweep:
|
||
|
||
| | spread |
|
||
|---|---|
|
||
| median attribution, **between** models | **750x** |
|
||
| typical spread **within** one model | **1.8x** |
|
||
|
||
The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving
|
||
concurrency, a stable per-model property, not weather. A model whose GPUs
|
||
carry far more concurrent requests genuinely costs less per request, and
|
||
that is most of the real cost difference in the catalog: `deepseek-v4-flash`
|
||
bills ~1000x under its share of pool gross. Stripping attribution discards a
|
||
750x real signal to suppress a 1.8x one.
|
||
|
||
So scoring reads the attributed figures, and the **median** absorbs what
|
||
noise remains. A split-half check on the 7-sample sweep (median of first
|
||
three vs last four) shows that working:
|
||
|
||
- **10 of 13 models agree within 1.4x** — stable enough to route on
|
||
- **3 do not**: `kimi-k2.7-code-fast` (29x), `kimi-k3` (14x),
|
||
`glm-5.2-flex` (2.2x). Those need more samples before their position is
|
||
trustworthy.
|
||
|
||
`dispatcher.gross_energy_kwh` remains as a diagnostic on the identity, not a
|
||
scoring input.
|
||
|
||
### Attribution drifts across hours, so sampling must too
|
||
|
||
Within about 30 minutes the billed figures reproduce (0.3-1.1x on a
|
||
spot-check). Across hours they do not: between two sweeps,
|
||
`deepseek-v4-flash` moved roughly 50x and `qwen3.6-35b` about 7x the other
|
||
way — enough to **invert their cost ranking**. Attribution tracks pool load,
|
||
and pool load tracks time of day.
|
||
|
||
More samples inside one sweep does not fix this; it measures one moment more
|
||
precisely. Coverage across time does. `load_candidates` already takes the
|
||
median over ALL `seed_reference` rows, so repeated sweeps accumulate into a
|
||
median-across-time for free — hence `llm-router-seed.timer`, which runs a
|
||
small sweep every 6 hours.
|
||
|
||
Until several sweeps have accumulated, treat the eco ordering as provisional.
|
||
A single sweep's ranking is one sample of a moving quantity.
|
||
|
||
**And none have accumulated since 6e729ad.** That commit moved
|
||
`log_observation`'s trailing arguments to keyword-only without updating
|
||
`seed_energy.py`, so every timer run since spent one billed completion and then
|
||
died on `TypeError` — which is not a `RequestException`, so the per-sample
|
||
`except` did not catch it. Fixed, and the sweep now has an offline end-to-end
|
||
test, but the accumulation this section describes starts from the next run
|
||
rather than from months of history.
|
||
|
||
## What's built and working
|
||
|
||
- `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md).
|
||
- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below.
|
||
- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md).
|
||
- `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md).
|
||
- `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md).
|
||
- `tiering.py` / `tier.py` — pure tier resolver + DB pass. Why tier on `reasoning_default_enabled`, cheapness-not-ceiling, `tier1_context_max`: [routing#tiering](docs/routing.md#tiering).
|
||
- `routing.py` — pure hard filters + ranking, plus the request-side capability gates (fail-closed asymmetry). [routing](docs/routing.md).
|
||
- `routing.py` `rank_candidates` incumbency: prices the session's last chat model at its measured cache rate and every challenger at the `objective.incumbent_challenger_cache_rate` dial, behind a load-bearing `min(dial, incumbent_rate)` clamp; gated off by default (`incumbent_cache_pricing: false`), tunable from off to full in config. [routing#incumbency-and-cache-pricing](docs/routing.md#incumbency-and-cache-pricing).
|
||
- `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates.
|
||
- `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below).
|
||
- `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md).
|
||
- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded.
|
||
- `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md). `--dry-run` no longer just describes the fold, it **projects** it: `feedback_preview.py` copies the DB into memory, runs the real `add_outcome` against the copy, and reports the per-row before/after, opening the source `mode=ro` so a preview cannot write to what it is previewing. `deploy/llm-router-feedback.{service,timer}` gives this loop the timer it never had — and ships **not enabled**, because the fold is irreversible and the first one against an accumulated backlog is an operator decision. See `deploy/README.md`.
|
||
- `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md).
|
||
- `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md).
|
||
- `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable.
|
||
- `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md).
|
||
- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. [api](docs/api.md).
|
||
- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. The decision table leads with a `time` column and carries `profile` plus an `E` flag for exploratory picks; the quota panel now shows per-account billing shapes (`metered_plan`, `prepaid_credit`, etc.) with plan/pool/burn/credit blocks, spend aggregates, and an alarm line. [architecture](docs/architecture.md).
|
||
- `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted.
|
||
- `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.
|
||
- `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md).
|
||
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor.
|
||
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
|
||
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
|
||
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
|
||
- `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md).
|
||
- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md).
|
||
- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md).
|
||
- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md).
|
||
- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md).
|
||
- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md).
|
||
- Agent guardrails — `scripts/verify_commit.py` (is this commit good?),
|
||
`scripts/oc_dispatch_audit.py` (audit an orchestrator session's dispatches),
|
||
and the opencode plugin `deploy/opencode-plugin/guardrails.js` whose
|
||
`tool.execute.before` hook blocks rule-breaking tool calls.
|
||
[agent-guardrails](docs/agent-guardrails.md).
|
||
|
||
## #45 — OpenRouter is an opt-in allowlist provider
|
||
|
||
OpenRouter used to be ingested whole, then removed, and is now back — but only
|
||
as an opt-in provider. The base config sets `openrouter.require_allowlist: true`
|
||
and ships a short seed allowlist. Models on that seed list are upserted and
|
||
kept active; anything else in the OpenRouter catalog is filtered out before
|
||
upsert and any previously-active OpenRouter row that is not on the list is
|
||
marked `deprecated` on the next poll.
|
||
|
||
This is deliberately different from the old ingest-everything behavior. The
|
||
previous approach once pulled in a non-chat model (`lyria/...`) that returned
|
||
HTTP 404 on dispatch because the endpoint expected chat completions. The router
|
||
had paid for the classification, selected the model, and then failed on the
|
||
provider call. Allowlist-gating prevents that class of failure by default: if a
|
||
model id has not been reviewed and explicitly added, the router acts as if it
|
||
does not exist.
|
||
|
||
The seed allowlist is short and has firm exclusions. It does NOT include:
|
||
|
||
- `x-ai/*` (Grok)
|
||
- `openai/*`
|
||
- `anthropic/*`
|
||
|
||
Those exclusions are non-negotiable. They are not "currently excluded" or
|
||
planned for future inclusion; they are deliberately absent from the seed list.
|
||
Adding one requires editing both the seed allowlist and this file.
|
||
|
||
The admin portal exposes the allowlist under `/admin`: the providers page shows
|
||
which providers require one, and each provider row links to an allowlist editor
|
||
where entries can be added or removed. The underlying four API endpoints are
|
||
`GET /admin/api/providers/{name}/allowlist`,
|
||
`POST /admin/api/providers/{name}/allowlist`,
|
||
`DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers
|
||
page itself surfaces `require_allowlist` with a link to the allowlist editor.
|
||
|
||
Neuralwatt remains the primary, ungated provider. The allowlist behavior only
|
||
fires for providers with `require_allowlist: true`.
|
||
|
||
## Proficiency: category now changes routing
|
||
|
||
`proficiency_score` is the ONLY category-dependent term in the ranking, so
|
||
until this table had data, `task_category` could not change a decision at
|
||
all — the classifier computed it, the router paid ~10s for it, and then it
|
||
made no difference. It does now. Two categories were added for local dispatch:
|
||
`file_summarization` and `diff_checking`; see [evaluation](docs/evaluation.md). The score is also no longer a raw benchmark
|
||
level: it has been converted into an **expected pass rate on real traffic**,
|
||
calibrated against 1,059 client-reported outcomes and shrunk with a
|
||
20-pseudo-observation prior so thin data does not dominate.
|
||
|
||
`proficiency` now holds 141 rows:
|
||
|
||
| source | count | meaning |
|
||
|---|---|---|
|
||
| `outcome_blended` | 33 | Fresh per-model outcome evidence |
|
||
| `outcome_prior` | 76 | Trafficked-sibling rows inheriting the peer-rate prior |
|
||
| `self_eval_thin` | 32 | Cold categories (summarization, translation); benchmark preserved verbatim |
|
||
|
||
A further 33 (model, category) pairs have direct outcome samples. The biggest
|
||
evidence gains went to `deepseek-v4-flash` (notably `coding_general`,
|
||
`coding_refactor`, and `general_chat`), `kimi-k2.7-code` (across 7
|
||
categories), and `qwen3.6-35b`.
|
||
|
||
Sweeping 9 categories x 3 tiers currently returns **5 distinct winners** at
|
||
both 50k and 120k of context.
|
||
|
||
| context | winners over 27 decisions |
|
||
|---|---|
|
||
| 50k | `qwen3.6-35b` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) |
|
||
| 120k | `kimi-k2.7-code` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) |
|
||
|
||
**This spread is recent, and how it got here is the useful part.** For a long
|
||
time all 27 decisions returned ONE model, and that was the correct answer at
|
||
the time rather than a bug: with cost and eco both populated, `qwen3.6-35b`
|
||
was Pareto-dominant — cheapest AND cleanest in the routable set, while
|
||
scoring within `quality_tolerance` of the best. No defensible weighting picks
|
||
anything else out of that.
|
||
|
||
Two corrections widened it, and neither was a tuning change:
|
||
|
||
- **Cost stopped being a benchmark average.** It is now priced per request
|
||
from catalog prices scaled to the request's shape, so the ranking depends
|
||
on the workload instead of on a 400-token reference sweep that no real
|
||
traffic resembles.
|
||
- **Tier stopped being inferred from price.** `deepseek-v4-flash` was pinned
|
||
to tier 1 for being cheap, which excluded it from every tier-2 request
|
||
regardless of what any score said.
|
||
|
||
A third shift is under way: the outcome backlog has been spent, so the score
|
||
now reflects real pass/fail reports rather than the benchmark alone. That
|
||
changes the numbers; it does not change the rule. Quality is still the
|
||
objective and cost is still the tiebreak within `quality_tolerance`.
|
||
|
||
Note what changes between the two rows above: only the leader, and only
|
||
because of the hard context filter. That is the filter working, not the
|
||
scoring disagreeing with itself.
|
||
|
||
**If you see one model win everything again, check for dominance before
|
||
reaching for config.** One winner is a legitimate outcome. The levers, if a
|
||
genuinely different balance is wanted, are `objective.quality_tolerance`
|
||
(how large an expected-success-rate gap must be before it outranks a cost
|
||
saving) or `objective.max_energy_per_request` (a hard ceiling). There is no
|
||
weight to tune.
|
||
|
||
### What the task set actually found
|
||
|
||
**The benchmark could not discriminate these models on coding.** Every row
|
||
scored exactly 1.00 on `coding_general`, `coding_refactor` and `debugging` —
|
||
and that is after the tasks were deliberately hardened with touching
|
||
intervals, full semver, present-but-falsy defaults, late-binding closures and
|
||
a binary search that infinite-loops. Every model in this catalog is simply
|
||
good at that class of problem, so cost decides coding routes, which is the
|
||
right outcome.
|
||
|
||
**Real traffic broke one of those ties, which the benchmark never could.**
|
||
`coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29
|
||
samples folded in by `feedback.py` from an actual agent session, and crossed
|
||
`self_eval_min_samples` on the way, so it reads `self_eval` rather than
|
||
`self_eval_thin`. That is the intended shape of this system — the 43-task
|
||
benchmark establishes a floor, and your own traffic is what refines it.
|
||
`coding_refactor` and `debugging` are still flat at 1.00, awaiting the same
|
||
treatment.
|
||
|
||
**Benchmark-sourced hardening landed, and it broke both remaining ties.** The
|
||
task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after
|
||
the `score_exact` literal-eval fix), and three Exercism refactor/debug pairs.
|
||
The CRUXEval-O rows split `coding_general` into a 0-1 mix across models — five
|
||
of six now fail at least one model — and the Exercism refactor rows moved
|
||
`coding_refactor` off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00,
|
||
affine 0.56-1.00. The debug pairs split `debugging` the same way (bowling
|
||
0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted:
|
||
`debug_affine_coprime` and most of the BFCL tasks sat flat at 1.00, so they do
|
||
not discriminate.
|
||
|
||
**A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2
|
||
samples per model the category read 0.70-1.00 with a model at the ceiling, and
|
||
the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more
|
||
benchmark passes moved every score and left NOTHING at 1.00:
|
||
|
||
| model | n=2 | n=11-14 |
|
||
|---|---|---|
|
||
| `kimi-k3` | 0.85 | **0.973** |
|
||
| `kimi-k2.7-code` | 0.85 | 0.886 |
|
||
| `deepseek-v4-flash` | 0.80 | 0.864 |
|
||
| `kimi-k3-fast` | **1.00** | 0.864 |
|
||
| `qwen3.6-35b` | 0.85 | 0.800 |
|
||
| `gemma-4-31b` | 0.85 | 0.786 |
|
||
|
||
The winner moved to `kimi-k2.7-code`, **3.2x cheaper** at 50k of context
|
||
($0.0441 -> $0.0136), with no config change — `kimi-k3` scores higher but sits
|
||
inside `quality_tolerance`, so cost breaks the tie. `deepseek-v4-flash`
|
||
($0.0024) misses the band by 0.009, which is the kind of margin the tolerance
|
||
exists to describe rather than a verdict.
|
||
|
||
**The whole spread rests on one rubric line, though.** `docs_function` is
|
||
effectively saturated — 1.00 on nine of every ten samples — and nearly every
|
||
`docs_gotcha` deduction is the same omission: the model documents that order is
|
||
preserved, that the first occurrence is kept, and what `key` does, then never
|
||
says elements must be hashable. That is real discrimination, since it is a real
|
||
property of the function, but one sentence is deciding a category. Treat this
|
||
ordering as thinner than n=14 makes it look.
|
||
|
||
**The self-judging guard costs sample density, and it shows up here.** Most
|
||
models reached n=14; `kimi-k3` and `kimi-k3-fast` reached only 11, because
|
||
those two are the ones diverted to the alternate judge `qwen3.6-35b`, which
|
||
returns unparseable JSON more often than `kimi-k3` does. The guard is still
|
||
right — a model grading its own family is worse than a thinner sample — but
|
||
the alternate judges should be picked for parseability, not just for being
|
||
someone else.
|
||
|
||
The reflex when a category looks flat is to reach for `quality_tolerance`.
|
||
Neither tie broken so far was broken that way: `coding_general` opened up when
|
||
`feedback.py` folded in real traffic, and `docs_writing` opened up on six more
|
||
benchmark passes. Both were samples, not settings. `coding_refactor` and
|
||
`debugging` are still flat at 1.00 on 2-3 samples each — which is now a state
|
||
this project has mistaken for a measurement once.
|
||
|
||
Current spread by category, widest first:
|
||
|
||
| category | spread |
|
||
|---|---|
|
||
| `tool_use_agentic` | 0.33 - 1.00 |
|
||
| `summarization` | 0.60 - 1.00 |
|
||
| `reasoning_math` | 0.67 - 1.00 |
|
||
| `docs_writing` | 0.66 - 0.97 |
|
||
| `general_chat` | 0.80 - 1.00 |
|
||
| `translation` | 0.85 - 1.00 |
|
||
| `coding_general` | 0.86 - 1.00 |
|
||
| `coding_refactor`, `debugging` | flat at 1.00 |
|
||
|
||
**What does discriminate is tool use, arithmetic traps, and prose.**
|
||
`deepseek-v4-flash` scores 1.00 on all three coding categories yet **0.33 on
|
||
`tool_use_agentic`** and 0.67 on `reasoning_math`. Verified live, not an
|
||
artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes
|
||
away?" — both times supplied — it calls *two* tools rather than subtracting.
|
||
It over-reaches for tools, which is exactly the failure mode that matters in
|
||
an agent loop. The router now avoids it for those categories while still
|
||
picking it for coding.
|
||
|
||
Most rows still read `source='self_eval_thin'` (118 of 132): real
|
||
measurement, but below `self_eval_min_samples` at 2-3 tasks per category per
|
||
run. The 14 that have crossed it are all `docs_writing`, from the six extra
|
||
passes above. Two paths thicken it, and they are complementary — re-run
|
||
`eval_proficiency.py` to accumulate benchmark samples, or just use the router
|
||
and let `feedback.py` fold in real outcomes. Both fold into a running mean
|
||
rather than replacing, so samples add up across runs.
|
||
|
||
### A score is only as fresh as the row it was copied to
|
||
|
||
Proficiency is a property of the weights, not the queue, so the eval harness
|
||
scores one row per family and `propagate_to_variants` copies the result onto
|
||
the serving variants — `kimi-k3-flex` gets `kimi-k3`'s number, because no
|
||
benchmark rates a `-flex` row separately.
|
||
|
||
That copy used to happen **exactly once per variant, ever.** The guard skipped
|
||
any row with `self_eval_samples > 0`, meaning "measured directly, do not
|
||
overwrite" — but inheritance copies the sample count too, so after the first
|
||
propagation an inherited row was indistinguishable from a measured one and was
|
||
never refreshed again. `kimi-k3-flex` sat at 0.85/n=2 while `kimi-k3` moved to
|
||
0.973/n=11.
|
||
|
||
`proficiency.inherited_from` records the provenance that was missing, and the
|
||
**migration** was the delicate half, not the fix: `ADD COLUMN` gives every
|
||
existing row NULL, which reads as "measured here", so shipping the guard alone
|
||
would have permanently frozen the exact rows it exists to unfreeze. The
|
||
backfill infers provenance from the harness's own selection rule rather than
|
||
guessing — `eval_identities` only ever evaluates standard rows plus flex rows
|
||
with **no** standard equivalent, so a flex row that has one was never a
|
||
candidate for direct evaluation, whatever its sample count claims. Everything
|
||
else keeps NULL, which fails safe: NULL means "do not overwrite", so no real
|
||
measurement can be lost to a wrong guess.
|
||
|
||
Confirmed on the live database, and on the catalog's one genuine exception —
|
||
`glm-5.2` is canary, so `glm-5.2-flex` is the routable row the harness scores
|
||
directly, and its NULL is correct.
|
||
|
||
### Harness bugs this shook out
|
||
|
||
Three separate defects, each of which scored the rig rather than the model,
|
||
and each caught by reading per-task detail rather than the summary:
|
||
|
||
- **Token budget.** `max_tokens` was shared between a reasoning model's trace
|
||
and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and
|
||
returned an EMPTY content field, scoring 0.00 on tasks it can plainly do.
|
||
Now 24000, clamped per model (gemma-4-31b caps at 16384), and
|
||
`finish_reason: length` skips the sample instead of scoring it.
|
||
- **One leading space.** kimi-k2.7-code returns `" def f(...)"`, which becomes
|
||
IndentationError once the harness prepends its imports — 0.00 across all
|
||
nine coding tasks for a model with "code" in its name.
|
||
- **Judge failures scored as model failures.** 44% of judge calls returned
|
||
unparseable output (the judge is itself a reasoning model and leaks its
|
||
thinking despite `response_format`). Each was recorded as 0.0. Now the JSON
|
||
is extracted from surrounding prose and an unusable reply yields no sample.
|
||
|
||
`tests/test_task_set.py` exists so this stops happening: it implements a
|
||
reference solution for every `code` task and asserts it passes every check,
|
||
recomputes every `exact` answer (one by brute force), and confirms each
|
||
refactor target already passes its own checks while each debugging target
|
||
fails. It immediately caught a check where the expected value was simply
|
||
wrong — which would have docked every model on a task and been
|
||
indistinguishable from genuine difficulty.
|
||
|
||
## Tool competence is read from the request, not guessed at
|
||
|
||
Neither local classifier can identify agentic work. Asked to label six
|
||
unambiguous tool-use prompts ("read the config then update the manifest",
|
||
"run the tests and fix what fails"), `qwen3.5` got 2/6 and `mistral-nemo`
|
||
1/6 — and `mistral-nemo`'s misses collapse to `general_chat`, which is also
|
||
the configured `fallback_category`, so qwen3.5's crashes land in the same
|
||
place.
|
||
|
||
That mattered because `tool_use_agentic` has the widest proficiency spread in
|
||
the table (0.33-1.00) and `deepseek-v4-flash` — the current winner on coding —
|
||
sits at the bottom of it.
|
||
|
||
**The fix was not a better classifier.** Whether tools are on the table is
|
||
stated in the request: every agent client sends a `tools` array, and
|
||
`chat_completions` never looked at it. Reading it is exact and free.
|
||
|
||
It is applied as a **hard filter**, not a category override, and the
|
||
distinction is load-bearing. The question is not "is this task agentic" but
|
||
"can this model be trusted with tools that exist". The recorded failure is
|
||
precisely the second one: `deepseek-v4-flash` was given a *non*-agentic prompt
|
||
("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times
|
||
supplied) and called two tools rather than subtracting. A model that
|
||
over-reaches is a hazard on every request where tools are available, whatever
|
||
a classifier would have labelled the task.
|
||
|
||
So `routing.min_tool_proficiency` drops any candidate whose measured
|
||
`tool_use_agentic` score is below it, but only when the request carries tools:
|
||
|
||
| request | winner on `coding_general` @ 50k |
|
||
|---|---|
|
||
| no tools | `deepseek-v4-flash` ($0.0024) |
|
||
| tools present | `qwen3.6-35b` ($0.0041) |
|
||
|
||
Verified live through `/v1/chat/completions` with identical bodies differing
|
||
only by the `tools` array. The cost of safety here is 1.7x on that route,
|
||
paid only where tools exist.
|
||
|
||
**It is currently set to `null`, i.e. OFF**, deliberately and pending
|
||
experiment. opencode sends `tools` on essentially every request, so with the
|
||
filter on, `deepseek-v4-flash` is excluded from ordinary agent traffic and its
|
||
~7x cost advantage goes unused; with it off, that advantage applies and a
|
||
model measured at 0.33 on tool use handles requests where tools are on the
|
||
table. Which is right is an empirical question and the benchmark cannot
|
||
answer it — the 0.33 comes from 3 tasks.
|
||
|
||
What settles it is `POST /outcome`: run with the filter off, let real pass/fail
|
||
reports accumulate, and compare `deepseek-v4-flash`'s `tool_use_agentic`
|
||
proficiency before and after. That is the one signal here that knows whether
|
||
the work actually worked, and `feedback.py` folds client outcomes in both
|
||
directions, so success counts too.
|
||
|
||
**Update 2026-09-15: that accumulation path is closed.** The classifier no
|
||
longer emits `tool_use_agentic` — per-turn classification collapsed onto it
|
||
(31 of 31 consecutive live turns, all routed to the slowest model in the
|
||
catalog) — so no new outcome attributes to the category and its scores are
|
||
frozen at today's values. Accepted, not fixed; the reasoning and the rejected
|
||
alternatives are in [docs/routing.md](docs/routing.md), "The classifier's
|
||
labels are not the proficiency scoring axis". The filter experiment now either
|
||
finds a different signal or reads frozen data.
|
||
|
||
0.5 sits in the empty band between the only two values the catalog holds
|
||
(0.33 and 1.00), so it is not fitted to either. A model with **no** measured
|
||
tool score is unproven rather than proven bad and is not dropped — the same
|
||
rule as the tier-1 context gate. Config load refuses a
|
||
`routing.tool_use_category` that is not a real category, because a name
|
||
matching nothing yields NULL for every row and NULL means "do not
|
||
disqualify": the filter would silently stop filtering.
|
||
|
||
## Tier is an iteration budget, not just a floor
|
||
|
||
A tier used to mean only "do not route below this". It now also buys
|
||
corrective attempts after a verification failure:
|
||
|
||
| tier | batch | interactive |
|
||
|---|---|---|
|
||
| 1 | 0 retries | 0 |
|
||
| 2 | 1 | 1 |
|
||
| 3 | 2 | 1 |
|
||
|
||
Interactive is capped below its tier because every retry doubles
|
||
time-to-answer, and in interactive use latency **is** a quality loss.
|
||
|
||
Retries are matched to the failure, since the causes differ:
|
||
|
||
- **truncated** — raise the token budget on the same model; a different one
|
||
would run out too. If there is no cap to raise, the model's own output
|
||
ceiling is the wall, so escalate to a candidate that can emit more.
|
||
- **malformed** — more tokens will not make unparseable output parse, so
|
||
escalate to the next-ranked candidate.
|
||
- **ok / unverifiable** — buy nothing. Retrying `unverifiable` would burn
|
||
quota across the majority of prose traffic for no signal.
|
||
|
||
`escalation.preemptive_on_low_confidence` is now **off by default**. Bumping
|
||
the tier because the classifier was unsure pays frontier prices before
|
||
anything has gone wrong; spending after a check has actually failed is better
|
||
on both mandates — the cheap attempt usually succeeds, and when it fails you
|
||
have evidence rather than a hunch.
|
||
|
||
## The only ground truth: `POST /outcome`
|
||
|
||
Everything else the router records is a proxy. Structural checks know whether
|
||
code *parses*. The local checker guesses whether prose *looks* right. Neither
|
||
knows whether the answer did the job — the client does, because it ran the
|
||
tests.
|
||
|
||
```bash
|
||
# id comes from the completion body, or any stream chunk
|
||
curl -s localhost:8080/outcome -H 'content-type: application/json' \
|
||
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
|
||
```
|
||
|
||
Two things make this the highest-value signal available:
|
||
|
||
- **It is the only quality signal that survives streaming.** A retry cannot
|
||
reach a streamed response because the bytes are already gone; a report
|
||
arrives afterwards and works either way. Every agent client streams.
|
||
- **Its successes count.** `feedback.py` folds `ok: true` and `ok: false`
|
||
into `proficiency.outcome_score` via `add_outcome()`. Structural and
|
||
`local_llm` verdicts are now diagnostics only; they no longer move the
|
||
score, because a structural failure is not a verified client outcome.
|
||
|
||
Client outcomes are what calibrate the routing score. The benchmark is a
|
||
prior; `POST /outcome` is the posterior.
|
||
|
||
`request_id` is written back to `route_decisions.request_id` on both the
|
||
streamed and buffered paths, so the report joins cleanly to the decision that
|
||
produced it. An unknown `request_id` returns 404 rather than being quietly
|
||
accepted — a client whose reports go nowhere should find out.
|
||
|
||
## Verification: what local compute is actually good for
|
||
|
||
Local inference is a poor substitute for cloud completions here — 7-82x the
|
||
energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s).
|
||
But it is very good at stopping a cloud completion from being wasted, and
|
||
completions are where all the money is: fitted on real traffic, a completion
|
||
token costs **201x** a prompt token.
|
||
|
||
| check | cost | what it catches |
|
||
|---|---|---|
|
||
| structural (`verification.py`) | **free** | truncation, malformed code/JSON/YAML, empty answers |
|
||
| local LLM (Ollama) | ~$6.9e-05, ~6s | refusals, wrong-question answers, incoherence |
|
||
|
||
Structural checks run on every response and never execute the code — they
|
||
parse it. The local LLM check runs only on answers above
|
||
`verification.min_completion_tokens`, in the background after the client has
|
||
its response, because it costs ~15% of a median 193-token answer and only
|
||
pays above ~600 tokens.
|
||
|
||
### An agent turn is not a prose answer
|
||
|
||
Both checkers had mirror halves of the same blind spot, and real traffic is
|
||
what found it. On the first genuine agent session (63 completions, a shipped
|
||
feature, 349 passing tests, clean mypy and ruff):
|
||
|
||
| checker | called it | actually |
|
||
|---|---|---|
|
||
| structural | `malformed: empty response` x29 | turns that ended in a tool call |
|
||
| local LLM | `cuts off mid-sentence` x8 of 9 | same turns, judged from the other side |
|
||
|
||
A turn that calls a tool has empty or half-finished text **by design**. Both
|
||
paths now take `has_tool_calls` and return `unverifiable`; in `verify_response`
|
||
that check outranks even `finish_reason == 'length'`, because stopping
|
||
mid-sentence at a call boundary is not a budget overrun. `worth_local_check`
|
||
declines outright, which also stops paying ~6s of local inference to
|
||
mis-grade a tool call.
|
||
|
||
Had `feedback.py` run before this, it would have applied ~12 false failures to
|
||
the models that had just shipped the feature. That is the **fourth** harness
|
||
bug in this project that would have scored the rig rather than the model, and
|
||
the first caught by real traffic instead of a synthetic test. The pre-fix rows
|
||
are kept but set `model_attributable = 0`, so the record survives without
|
||
steering routing.
|
||
|
||
**`client_capped` was also over-applied.** It marked *every* verdict
|
||
non-attributable whenever the client set `max_tokens` — and opencode always
|
||
does — so real failures were invisible to feedback. A client's token cap
|
||
explains a `truncated` verdict and nothing else; it is now scoped to exactly
|
||
that.
|
||
|
||
`feedback.py` folds observed failures into `proficiency`, so routing learns
|
||
from your traffic rather than only the 43-task benchmark. Only **failures**
|
||
are folded in: a structural 'ok' means the code parsed, not that it was
|
||
correct, and recording those as 1.0 would flatten every score toward the
|
||
ceiling. Failures the model did not cause — a client's own tight `max_tokens`
|
||
truncating the answer — are recorded but excluded.
|
||
|
||
## The classifier is the latency floor
|
||
|
||
Every routed request pays a full local classification round-trip before an
|
||
upstream token is requested — the classifier is the latency floor. Four settings keep it usable:
|
||
`max_retries=0` (SDK retry → 3x wall-clock), `max_output_tokens: 1024` (bounds reasoning), `max_input_chars: 8000` (doesn't need the document; head+tail clamp), `fallback_tier`/`fallback_category` (mid tier, not 502/503), plus `temperature: 0` (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not `0.0.0.0` (Ollama has no auth).
|
||
Full measurements: [docs/local-models.md](docs/local-models.md#the-classifier-is-the-latency-floor).
|
||
|
||
### A classifier failure no longer collapses to one fixed guess
|
||
|
||
`fallback_tier`/`fallback_category` above is now the LAST step, not the only
|
||
one. When the local classifier fails, `classify()` walks a cascade:
|
||
|
||
| # | step | cost |
|
||
|---|---|---|
|
||
| 1 | this session's cached classification, **staleness ignored** | free |
|
||
| 2 | this session's history in `route_decisions` | free |
|
||
| 3 | `classifier.cloud_fallback`, if configured | money |
|
||
| 4 | `fallback_tier` / `fallback_category` | free |
|
||
|
||
Steps 1 and 2 reuse a *real* classification rather than guessing. Step 2 is
|
||
what survives a restart, when the in-memory cache is gone but the session's
|
||
decisions are still on disk. Both filter degraded sources, so one outage's
|
||
guess cannot propagate through every later turn of a session and end up
|
||
looking like a measurement.
|
||
|
||
**Step 3 is absent by default and the whole feature costs nothing until it is
|
||
configured.** Three things bound it once it is:
|
||
|
||
- `classifier.cooldown_seconds` (30) is a **global** backoff, so a sustained
|
||
outage buys one cloud attempt per window across ALL requests, not one per
|
||
request;
|
||
- an account-level provider refusal suppresses step 3 for the same window —
|
||
an out-of-credit account makes that call a guaranteed wasted request;
|
||
- the `session_cache.put()` gate admits `classifier_cloud`, bounding a session
|
||
to one cloud call per staleness window instead of one per turn.
|
||
`session_stale` and `session_history` are deliberately NOT cached: borrowed
|
||
answers must not renew a staleness clock they never earned.
|
||
|
||
**The trap this had to avoid, and it is the reason to read this section.** A
|
||
degraded classification is recorded with `task_category = general_chat` —
|
||
which is a *fully scored* category (13 proficiency rows) as well as the
|
||
configured `fallback_category`. Without a guard, `POST /outcome` reports on
|
||
mislabelled outage traffic fold into `proficiency(model, general_chat)` and
|
||
drag real scores toward whatever happened to be flowing while the classifier
|
||
was down. So `report_outcome` marks outcomes of degraded-source decisions
|
||
`model_attributable = 0` — kept as a record, excluded from folding, the same
|
||
treatment the tool-call false-failures got. `feedback.py` is untouched; its
|
||
existing `AND model_attributable = 1` already does the work.
|
||
|
||
Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not:
|
||
`fallback`, `session_stale`, `session_history`. The asymmetry is deliberate —
|
||
a degraded source is good enough to route one visibly-flagged request, but a
|
||
proficiency score is consulted by every future request, so attribution must
|
||
not trust more than routing does. Unknown decision rows **fail open**; an
|
||
over-applied exclusion already starved feedback once (`client_capped`).
|
||
|
||
`/metrics` warns when the degraded share of the last 24h crosses
|
||
`classifier.degraded_warn_threshold` over at least `degraded_warn_min`
|
||
decisions. A survivable failure is exactly the kind that goes unnoticed for
|
||
weeks.
|
||
|
||
**Why a decline is now recorded, and what the numbers said (2026-10-04).**
|
||
Under `local_decision` on the 4b the classifier declines about a third of agent
|
||
turns on purpose: 769 of 2,557 (30%) that day, 8% the day before, against 0% for
|
||
`local_llm` and `local_encoder`. 22.5% were `session_history` replays and 7.6%
|
||
the static `general_chat` guess, and `session_history` has no age bound (p99 67
|
||
minutes, max 100). None of that could be tuned from the database, because
|
||
`route_decisions.confidence` is the override's hard-coded 1.0 on 99.9% of chat
|
||
rows. `classifier_confidence`, `classifier_coverage` and `classifier_reject` now
|
||
hold the classifier's own numbers and the reason (codes in
|
||
[data-model](docs/data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)),
|
||
and the degraded-share warning lists the reasons it saw. The warning's default
|
||
threshold (0.5) never fires at that rate, so the live overlay sets 0.2.
|
||
`degraded_warn_min` and `degraded_warn_threshold` have no admin control, and that
|
||
is a deferred gap, not a decision: the `classifier` section is outside the
|
||
knob-coverage gate because it has its own card, and adding any `classifier.*`
|
||
knob to the generic registries would pull every other classifier scalar into the
|
||
gate at once. Do it in one pass when classification is reworked.
|
||
|
||
### Which implementation is PRIMARY is now a config choice
|
||
|
||
`classifier.mode` in the live deployment is currently `local_encoder`, set in
|
||
`config.local.yaml` with `device: cpu` on the classifier host. That is the
|
||
mode answering real traffic right now. The explicit caveat is that a
|
||
CPU-vs-CUDA latency and confidence comparison on the live classifier host is
|
||
still pending; until that measurement exists, flipping the project default to
|
||
`local_encoder` is not decided.
|
||
|
||
**Turning that mode on for real found three real bugs in one afternoon
|
||
(2026-09-06), each a fresh instance of this project's own recurring lesson —
|
||
verify against the live system, not the plan.** First, the config-load
|
||
validator for `confidence_threshold` didn't exist yet: the admin UI saved a
|
||
raw `80` (meant as 80%) straight into the overlay with no conversion, which
|
||
would have made every real confidence score read as below-threshold on the
|
||
next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability
|
||
exceeds 1.0). Caught before the restart, not after — see `confidence_threshold`
|
||
in [local-models](docs/local-models.md) for the full incident and the fix.
|
||
Second, once that was corrected and the service actually restarted, it
|
||
crash-looped twice more before coming up clean: the shipped default model
|
||
(`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace
|
||
sometime after this project picked it (401 on an unauthenticated GET of its
|
||
own model page), and separately `HF_HOME`'s default cache path falls outside
|
||
this service's `ProtectHome=read-only` sandbox exception — both are now fixed
|
||
(switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo).
|
||
Third, and most consequential: the very first real classifications measured
|
||
only 5 of 9 test categories correct, because `classify_zero_shot` was passing
|
||
raw config identifiers like `tool_use_agentic` and `diff_checking` directly as
|
||
zero-shot candidate labels — HF's pipeline scores a label against a hypothesis
|
||
template ("This example is {}."), and an underscored code token is not a
|
||
sentence the model's NLI training ever saw. Mapping each category to a natural-
|
||
language description before scoring, plus `multi_label=True` (the pipeline's
|
||
single-label default forces every candidate to compete for the same
|
||
probability mass), brought that to 8 of 9 correct with confidence scores
|
||
0.77-0.999 on the hits — the one remaining miss scored below the configured
|
||
threshold and correctly fell through to the safe fallback rather than
|
||
mis-routing.
|
||
|
||
`classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`,
|
||
`local_decision`) picks what answers a classification request — a peer concept to the cascade above,
|
||
**not a replacement for it**. Whichever mode is primary, a failure still
|
||
walks the exact same cascade (stale session → session history →
|
||
`cloud_fallback` → the static guess), unmodified.
|
||
|
||
- **`cloud_llm`** makes a cloud model the primary attempt, not just the
|
||
cascade's post-failure backup. Either pin one (`classifier.cloud_primary`,
|
||
same shape as `cloud_fallback`) or set `cloud_primary_auto: true` to
|
||
resolve the cheapest currently-routable model **live** against the catalog
|
||
(`routing.cheapest_classifier_candidate`, priced for the classifier's own
|
||
short-prompt/short-completion shape — not the task's). A success here
|
||
records `source="classifier"`, deliberately the *same* string
|
||
`local_llm`'s success uses, not `"classifier_cloud"` — that string means
|
||
specifically "the cascade's backup step fired" and feeds the `/metrics`
|
||
degradation-share warning above as a degraded signal. An intentionally
|
||
configured primary succeeding is not degraded.
|
||
- **`local_encoder`** classifies with a small, non-generative zero-shot
|
||
model instead of an LLM — structurally immune to the runaway-reasoning
|
||
failure mode documented above, since there is no generation to run away.
|
||
Zero-shot rather than fine-tuned: this router never stores raw task text
|
||
anywhere, so there is no training corpus without a new, separate opt-in
|
||
capture feature (not built). Only produces `task_category`; `task_tier`
|
||
falls back to `fallback_tier` — a real limitation, not a bug. A
|
||
below-threshold confidence is treated as a failure and cascades exactly
|
||
like a local-LLM parse failure would.
|
||
- **`local_decision`** asks a small generative model (configured in
|
||
`classifier.decision`) to pick a category from a set of natural-language
|
||
descriptions, then optionally classifies tier and — when enabled in its
|
||
decision config — runs the same A/B confidence check on the chosen label.
|
||
Requires a `classifier.decision:` block in config (confidence threshold,
|
||
base URL, model name); it is an opt-in overlay mode, not the default, and
|
||
the config load will refuse to start without it when the mode is selected.
|
||
Only produces `task_category` and an optional `task_tier` when the
|
||
decision block's `tier_classification.enabled` is true. Below-threshold
|
||
confidence cascades like a parse failure would.
|
||
- **Neither is gated by `local_compute.enabled`** (gaming mode, below) the
|
||
way `local_llm` is: `cloud_llm` never touches local hardware, and
|
||
`local_encoder` is small enough to run on CPU, so neither competes for the
|
||
GPU gaming mode exists to free up.
|
||
|
||
Configured via `config.yaml` (global default) and overridable per-machine
|
||
in `config.local.yaml` — the existing overlay, not a new mechanism — or
|
||
through the admin portal's Classifier card, which reports the **live**
|
||
resolved primary for `cloud_primary_auto` rather than echoing the config
|
||
value (see [admin-portal](docs/admin-portal.md)).
|
||
|
||
The default in `config.yaml` is still `local_llm`. `local_encoder` is
|
||
intentionally an opt-in per-deployment choice rather than the repository
|
||
default until the pending CPU-vs-CUDA comparison on the live classifier host is
|
||
available. `local_decision` is also opt-in: it requires a `classifier.decision:`
|
||
block with its own base URL, model, and confidence thresholds — the config load
|
||
refuses to start without it when the mode is selected.
|
||
|
||
## Local dispatch model
|
||
|
||
A second local model can now be dispatched directly for specific categories.
|
||
`qwen2.5-coder-router:14b` is configured as a tier-1 local row with
|
||
`provider='ollama-local'`, gated by `models.eligible_categories`
|
||
(`file_summarization` and `diff_checking`). The poller refreshes the row each
|
||
run; `seed_local_dispatch_energy.py` derives its price from measured GPU draw
|
||
and the user's tariff. Routing treats a local row like any other candidate
|
||
once the category filter admits it, and the circuit breaker excludes it on a
|
||
local failure so the next request reroutes to cloud candidates.
|
||
|
||
Known limitations of the local dispatch branch right now:
|
||
|
||
- **No true streaming.** The response is shaped into an SSE stream, but the
|
||
local answer is generated before any bytes leave the router.
|
||
- **No verification rows.** Structural and local-LLM checks run but are not
|
||
written to `verifications` for local answers.
|
||
- **No within-request cloud failover into the upload path is gone.** An eligible
|
||
routed request degrades to the local dispatch model when the cloud account
|
||
refuses or is exhausted (a fallback, not a preference), so the local row is
|
||
no longer a dead-end before a client retry. The degraded answer is recorded
|
||
as `kind='local_dispatch_fallback'`.
|
||
- **Follow-ups are not special-cased.** A pinned or auto-routed follow-up to the
|
||
same local model works, but nothing caches the loaded model between turns.
|
||
|
||
Dormant under the default profile by design — see docs/routing.md § Local dispatch branch.
|
||
|
||
`POST /outcome` now attributes through the local energy ledger too. Local rows
|
||
include `request_id` and `session_dir` in `local_energy_observations`, so a
|
||
client report on a local answer resolves to the same `(model_id, provider,
|
||
task_category)` provider-agnostic record as a cloud one.
|
||
|
||
## Routing notes
|
||
|
||
Ranking is quality-first, cost as a tiebreak; cost is never allowed to override
|
||
a real quality gap. The optional `objective.credit_attenuation` block extends
|
||
that tiebreak without changing it: when the block is enabled, a per-provider
|
||
multiplier is applied to a candidate's comparison cost only, producing an
|
||
`effective_cost` that breaks ties. The multiplier is derived from the provider's
|
||
polled account balance (the `balance_url` path, such as OpenRouter), so a low
|
||
prepaid balance can nudge a near-tie toward a healthier provider. The logged
|
||
`est_cost_usd` and the decision history stay as raw catalog estimates. The
|
||
multiplier is 1.0 for providers whose balance comes from per-completion
|
||
`allowance_remaining_usd` telemetry (NeuralWatt), so normal overage readings do
|
||
not bias routing.
|
||
|
||
Two semantics matter when reading the numbers. `total_balance_usd` is a sum of
|
||
heterogeneous provider-reported readings: OpenRouter's prepaid credits plus
|
||
NeuralWatt's overage allowance, which normally reads near -$0.004. It can be
|
||
negative and it is not a single spendable figure. `credit_attenuation.enabled`
|
||
deliberately lives only in the config file; it is absent from the admin
|
||
persisted-config allowlist and from provider edits. Turning it on or off
|
||
requires editing `config/config.yaml` and `systemctl --user restart
|
||
llm-router.service`, because the dispatcher's `cfg` binds at import time.
|
||
|
||
## What's NOT built yet — pick up here
|
||
|
||
Built: session-directory attribution, the local energy ledger, local model
|
||
dispatch, admin profiles/proficiency/gaming-mode, and the configurable
|
||
classifier backend (`classifier.mode`, including its admin card — all listed
|
||
under "What's built and working" above).
|
||
|
||
**As of 2026-09-05, PRs #26-#36 landed** the gitignored config overlay, admin
|
||
profile CRUD writing to it, the provider literal cleanup, quota
|
||
balance/burn/runway, capability-aware ceiling and rejection warnings, the TUI
|
||
schema catch-up, the classifier fallback cascade, and the admin portal uplift
|
||
(proficiency page, profiles duplicate/coverage fix, gaming mode, the
|
||
classifier-backoff bug fix). `classifier.mode` (this document's own section
|
||
above) is a further, independent addition on top of that.
|
||
`plans/multi-provider-support.md` is PARKED on provider selection — Z.ai was
|
||
the recommendation and is no longer settled; the coupling surface in it is
|
||
measured and still valid.
|
||
|
||
Known follow-ups recorded but not specced, both small:
|
||
|
||
- The TUI decision table renders the literal `"None"` in the `ctx` cell when
|
||
`required_context_tokens` is absent — the same defect the `profile` cell was
|
||
written to avoid. See `.omo/notepads/tui-overhaul/issues.md`.
|
||
- A NULL `required_context_tokens` raises `TypeError` inside the
|
||
demand-ceiling comparison, and the column is nullable. Latent only: zero
|
||
such rows exist today, checked on the live DB.
|
||
|
||
The items below remain open.
|
||
|
||
1. **Fine-tuning `local_encoder` on real traffic.** Scoped, not built:
|
||
`classifier.training_capture.enabled` (opt-in, off by default — a
|
||
deliberate reversal of "never store task text", so it must be impossible
|
||
to enable by accident), a `classifier_training_samples` table gated the
|
||
same way `report_outcome` already filters proficiency (only rows whose
|
||
`classification_source` is in the attributable set), and a
|
||
`train_local_encoder.py` script matching `eval_proficiency.py`'s
|
||
conventions. `local_encoder.py` currently ships zero-shot only.
|
||
|
||
2. **Leaderboard priors are unfilled.** `leaderboards.yaml` ships empty on
|
||
purpose — inventing plausible-looking benchmark numbers would put
|
||
fabricated data straight into routing, the same failure as the provider's
|
||
`static_fallback` carbon constant this project already excludes. Until real
|
||
sourced figures go in, a newly listed NeuralWatt family has no prior and
|
||
relies entirely on self-eval accumulating. `python leaderboard.py --check`
|
||
lists what is missing.
|
||
|
||
3. **Sampling depth for three models — now `eco`-only.** 7 samples/model gives
|
||
split-half agreement within 1.4x for 10 of 13, but `kimi-k2.7-code-fast`
|
||
(29x), `kimi-k3` (14x) and `glm-5.2-flex` (2.2x) are still unsettled. This
|
||
no longer touches cost, which is priced per-request from the catalog, so it
|
||
only affects `eco` — which is not an objective. Low priority unless eco
|
||
comes back.
|
||
|
||
4. **Retry does not reach streaming.** The iteration budget (`iteration.py`)
|
||
retries after a failed check, but only on the non-streaming path — once
|
||
bytes have gone to the client there is nothing to take back. Buffering to
|
||
fix that would cost streaming itself, a worse trade for interactive work.
|
||
`POST /outcome` is the answer for streamed traffic: it arrives afterwards,
|
||
so it works identically either way.
|
||
|
||
5. **`local_encoder` noise isolation — built for the confirmed shapes; residuals below.**
|
||
The raw task's fenced code blocks, `Tool result:`-shaped lines, and closed
|
||
`<system-reminder>` spans are now stripped by `_isolate_task_text` (pure,
|
||
stdlib-only, deterministic) before the fit + embed pass — `classify_zero_shot`
|
||
runs isolate → fit → prefix, so cleaning happens first and a noisy task often
|
||
fits the token window outright. Measured on the real `BAAI/bge-large-en-v1.5`
|
||
(2026-09-19): 8 clean one-sentence tasks scored 8/8, but the same instructions
|
||
wrapped in that noise scored 2/8 with the truncation fix already in place —
|
||
and tail-biased windowing ALONE also scored 2/8 on long noisy pairs, so the
|
||
fit does not subsume isolation. Post-isolation: 8/8 on short and long noisy
|
||
pairs, clean-vs-noisy pair consistency 8/8 + 8/8; `'[code]'`/`'[elided]'`
|
||
placeholder tokens measured worse than pure removal (7/8, 6/8 on long pairs)
|
||
and were rejected; a 20%-ratio floor guard measured harmful (3/8 — it reverts
|
||
exactly the short noisy inputs isolation exists to fix) in favor of an
|
||
absolute 24-char floor that only catches near-all-code inputs. Offline
|
||
regression: a noise tripwire (same instruction bare vs wrapped must classify
|
||
identically) fails against pre-isolation code and passes after. Still not
|
||
built: attention-masking de-weighting as an alternative to stripping (it
|
||
would preserve the noise tokens' presence without letting them dominate);
|
||
unclosed `<system-reminder>` tags and un-fenced diff hunks are left in place
|
||
(only closed-tag spans, fenced blocks, and marker-prefixed lines are
|
||
stripped); and a code-grounded instruction whose pasted snippet is the
|
||
subject can still land on a near-category — measured 3/4 on a 4-task
|
||
grounded set, the residual miss being description similarity
|
||
("refactor this helper" + code → `debugging`), not noise dominance.
|
||
|
||
## Gaming mode, and the backoff that used to do nothing
|
||
|
||
**The classifier circuit breaker did not break the circuit.**
|
||
`_last_classifier_failure` was written by `_record_failure()` and read by
|
||
nothing — `_classify_cascade` gated only its *cloud* step, and on a different
|
||
timestamp. So the router re-dialled a known-dead local classifier on every
|
||
request. A stopped Ollama refuses immediately and costs little; a **hung** one,
|
||
or a VPN-bound one that black-holes, costs the full 120s `timeout_seconds` per
|
||
request for as long as the outage lasts. `_classifier_backoff_active()` is now
|
||
the read, consulted **before** the client is constructed.
|
||
|
||
One detail there is load-bearing: recording the failure lives in `classify()`'s
|
||
exception handlers, NOT in `_classify_cascade`. The cascade is walked for
|
||
reasons other than a fresh failure, and if those re-stamped the clock, every
|
||
request during an outage would push the deadline forward and the local
|
||
classifier would never be re-probed while traffic flowed — a permanent outage
|
||
wearing a circuit breaker's clothes.
|
||
|
||
**`local_compute.enabled` (default true) is the outer gate over local
|
||
hardware.** Turn it off when you stop Ollama for a game and the router *skips*
|
||
every local call rather than discovering the outage one timeout at a time:
|
||
classifier, `/health` probe, local verification, local-vision fallback, and
|
||
local dispatch rows (dropped in `load_candidates`, so a local row is never
|
||
picked and then 503'd). `/v1/models` stops listing local rows, and an explicit
|
||
pin gets a 503 naming the flag instead of NeuralWatt's unknown-model 400.
|
||
|
||
ONE flag the code reads, **not** a macro writing five keys — a macro is hard to
|
||
undo cleanly, drifts the moment a sixth call site appears, and leaves nobody
|
||
able to answer "why isn't the classifier running?" from one place.
|
||
`verification.local_llm_enabled`, `local_vision.enabled` and
|
||
`local_energy.enabled` keep their own meanings; this ANDs over them.
|
||
|
||
**It REFUSES to engage without `classifier.cloud_fallback`** — 409 on the
|
||
runtime knob, a validation error at config load. Skipping the local classifier
|
||
does not make classification remote; without a cloud classifier it stops
|
||
classifying, and every request falls through to a static guess recorded as
|
||
`general_chat`, a fully scored category indistinguishable from a real
|
||
classification afterwards. A refusal, not a warning, because a warning is what
|
||
nobody reads while their game is loading. Nothing auto-writes the block.
|
||
|
||
Cascade steps 1 and 2 still run **ahead** of the cloud call: a stale session
|
||
classification is free and was a real classification of that same session, so
|
||
paying to re-derive an answer already held is spending money for nothing.
|
||
|
||
**A latent substring bug fell out of the profiles work.** SQLite stores
|
||
`eligible_categories` as a comma-joined string, and
|
||
`task_category not in "<a>,<b>"` is a SUBSTRING test — so a row eligible only
|
||
for `file_summarization` also admitted `summarization`. Latent on main (the
|
||
category-less probe short-circuits before the compare) and live the moment
|
||
anything probes per category. `routing.parse_eligible_categories` is now the
|
||
single parser `dispatcher.load_candidates` and `admin.py`'s probe both use.
|
||
|
||
## Known open questions
|
||
|
||
- Answered: cost and eco stay separate axes — grid intensity spans 13.6x
|
||
across the catalog, so they rank models differently.
|
||
- Answered: the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh were
|
||
`carbon_source: static_fallback` — a substituted constant, not a
|
||
measurement. They are now excluded from eco rather than trusted. Still
|
||
worth asking NeuralWatt why the fallback keeps the original `grid_id`,
|
||
since that is what made it look like a real regional difference.
|
||
- Three models still fail a split-half stability check at 7 samples. Is the
|
||
instability real (variable serving conditions) or an artifact of when the
|
||
sweep ran? Re-sweeping at a different hour would tell.
|
||
- Answered, and the question no longer parses: tier-1 composites used to sit
|
||
within 0.009 of each other because min-max normalization compressed them.
|
||
There is no composite any more — ranking is quality first, cost as the
|
||
tiebreak inside `quality_tolerance` — so nothing normalizes and nothing
|
||
compresses.
|
||
- Answered: the eval set exists (`evals/tasks.yaml`, 43 tasks, four scoring
|
||
kinds) and `tests/test_task_set.py` keeps it honest. The benchmark-sourced
|
||
rows now split the coding categories, but two tasks are still flat at 1.00
|
||
(`debug_affine_coprime` and most of the BFCL set) and either need hardening
|
||
again or should be conceded as non-discriminating. **Try samples
|
||
before hardening.** `docs_writing` looked flat at the top too, and six more
|
||
passes spread it 0.66-0.97 without touching a task; two samples per model is
|
||
not enough to tell a saturated task from an unsampled one.
|
||
- How much context-assembly (RAG-style retrieval) belongs in the classifier
|
||
step vs. a separate pre-step? Leaning decoupled, undecided.
|
||
- Should `eco_score` use real-time grid carbon intensity per request or a
|
||
stable per-model average? Currently the latter, from the reference sweep.
|
||
`grid_carbon_intensity` and `grid_id` are logged per observation, so this
|
||
stays answerable from data without a re-run.
|
||
|
||
## Config is strict: an unknown key is an error
|
||
|
||
Pydantic ignores extra keys by default, which means a typo or a misplaced
|
||
setting loads cleanly, does nothing, and still looks configured. Every config
|
||
model now inherits `StrictModel` (`extra="forbid"`), so both of these fail at
|
||
load rather than silently:
|
||
|
||
```
|
||
verification.max_input_chars # right key, wrong section
|
||
routing.min_tool_proficency # sic
|
||
```
|
||
|
||
This is not hypothetical. `max_input_chars` shipped into the `verification:`
|
||
block instead of `classifier:` and was accepted and discarded — it happened to
|
||
match the code default, so behaviour was correct and the file was a lie.
|
||
Editing it would have done nothing.
|
||
|
||
The corollary worth keeping: **every knob belongs in `config.yaml`, not only
|
||
in a Pydantic default.** A default the file never mentions is invisible to
|
||
anyone tuning it. `classifier.outcome_attribution_window_seconds` was removed
|
||
in the same pass — it was declared, never read, and shadowed the
|
||
`verification` one that actually is.
|
||
|
||
## Setup
|
||
|
||
Full install steps (venv, deps, config, first run) in [README ## Installation](README.md#installation). Host-local deployment values go in `config/config.local.yaml` (gitignored overlay) — `classifier.model`/`base_url`, `local_energy.*`, host-specific URLs. General defaults in `config/config.yaml` stay shareable. `objective.plan_kwh_per_period` in the README config table. Model tags (`num_ctx`) + `verification.model` same-tag note in [docs/local-models.md](docs/local-models.md). Requirements are pinned — bump deliberately (README).
|
||
|
||
## Run as a service
|
||
|
||
`deploy/` holds the dispatcher's systemd **user** unit plus a timer and a
|
||
oneshot service each for the poller, the seed sweep, the backup, the offsite
|
||
sync and the feedback fold, and one drop-in for a *system* Ollama — see
|
||
`deploy/README.md` for install and operation. Every timer there is enabled on
|
||
install except `llm-router-feedback.timer`, which is not, on purpose. In short:
|
||
|
||
```bash
|
||
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
|
||
cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/
|
||
systemctl --user daemon-reload
|
||
systemctl --user enable --now llm-router.service llm-router-poller.timer
|
||
```
|
||
|
||
The dispatcher binds `127.0.0.1:8080`. **The poller timer is load-bearing, not
|
||
housekeeping** — but not for the reason this section used to give, and the
|
||
correction matters because it inverts which failure to watch for.
|
||
|
||
`mark_stale` runs only *inside* `poller.main()`, and `main()` returns early on
|
||
a `RequestException` — **before** `upsert` and **before** `mark_stale`. So a
|
||
stopped timer or a provider outage marks nothing: the catalog freezes at
|
||
last-known-good and the router keeps routing on prices that may be weeks old.
|
||
The failure is **silent and open**, not loud and closed. Nothing surfaces it,
|
||
because a frozen row still reads `availability = 'active'`.
|
||
|
||
The path that *can* empty the candidate set is narrower and is not governed by
|
||
the timer at all. `fetch_neuralwatt` reads `payload.get("data", [])` with no
|
||
floor on row count, so a 200 response carrying an empty or truncated `data`
|
||
array — a partial provider outage, a schema change, an auth path degrading to
|
||
an empty list — clears `raise_for_status()`, upserts nothing, and then lets
|
||
`mark_stale` run anyway. Three days of that and every row is stale and
|
||
`exclude_stale: true` leaves zero candidates for everything.
|
||
|
||
`stale_after_days: 3` only sets the length of that fuse; it does not arm or
|
||
disarm it. Against a 2-hourly poll it is 36 successful polls of margin, and
|
||
recovery is automatic — `upsert` writes `availability = excluded.availability`,
|
||
so one good poll flips every stale row back to active. The fix is a sanity
|
||
floor on the fetch, not a larger number.
|
||
|
||
**A second path empties the candidate set, and it bit on 2026-09-01.** Admin
|
||
availability overrides are not governed by the poller at all. Deprecating the
|
||
seven expensive models through `/admin` collapsed tier 3's context ceiling from
|
||
782,324 to **94,196** — while tiers 1 and 2 stayed at 782,324 — so every tier-3
|
||
request above 94k returned 422 with nothing warning anywhere. It surfaced ~19
|
||
hours later as an agent failing mid-task on an opaque error.
|
||
|
||
**The obvious check for this is wrong, and the reason is worth remembering.**
|
||
Tempting: warn when a higher tier's context ceiling sits below a lower tier's.
|
||
But `ceiling(T)` is the max `effective_context_window` over models with
|
||
`tier >= T`, and tier is a capability *floor*, so the eligible set shrinks
|
||
monotonically as T rises — `ceiling(1) >= ceiling(2) >= ceiling(3)` is a
|
||
theorem, true of every catalog. Such a warning fires always and means nothing.
|
||
What actually failed is that a tier's ceiling dropped below what that tier is
|
||
*asked* to serve, which is only knowable from traffic: compare `ceiling(T)`
|
||
against the observed `required_context_tokens` for decisions classified at tier
|
||
T. That is silent on all three tiers today and fires on the outage state
|
||
(94,196 vs an observed max of 268,168). See
|
||
`plans/catalog-staleness-and-poller-failure-modes.md` §4.4.
|
||
|
||
**This recurred on 2026-09-04 through a dimension the detector did not model,
|
||
and both halves of the fix are now in `metrics.py`.** Admin deprecations took
|
||
out `kimi-k3*` — the only vision-capable rows with enough context — so a
|
||
242,486-token image request 422'd while every existing check stayed silent,
|
||
because the *all-models* tier-1 ceiling was still 782,324. The vision-capable
|
||
ceiling had collapsed to 192,500.
|
||
|
||
- **Predictive:** `capability_ceilings` / `capability_demand_warnings` compute
|
||
`vision` and `json_mode` sub-ceilings and compare each against demand
|
||
actually observed for requests carrying images / requesting JSON. Two extra
|
||
series, not a bucket per capability combination.
|
||
- **Reactive:** `rejection_warnings` watches `route_decisions` for rows with
|
||
`selected_model IS NULL`. This is the more valuable half and the simpler
|
||
one — it catches the *next* dimension nobody predicted, at the cost of
|
||
firing after the first failure rather than before.
|
||
|
||
Two details in the reactive detector are load-bearing and easy to undo by
|
||
accident. It groups by `(task_tier, digit-normalized reason)` using the
|
||
**structured column**, because normalizing digits alone merges `tier >= 1`
|
||
and `tier >= 3` rejections into one group and hides whether the broadest or
|
||
the frontier candidate set went empty. And the signal is **novelty OR rate**,
|
||
never mere presence: measured on the live DB, routine rejections run ~3/hr
|
||
while the 2026-09-04 incident was only n=2 — *below* the noise floor — so no
|
||
single count threshold can both catch it and stay quiet. A group absent from
|
||
the 24h baseline warns at n≥2; a familiar group warns at the configured
|
||
count. Zero rejections warn about nothing: a genuinely impossible request
|
||
SHOULD 422.
|
||
|
||
The service holds a billable API key and has **no auth of its own**. Loopback
|
||
bind is the only thing standing between the open internet and your allowance;
|
||
add auth before widening `--host`.
|
||
|
||
The same applies to an Ollama shared over a VPN — it has no auth either, so
|
||
`deploy/ollama-over-vpn.conf` binds it to the VPN address rather than
|
||
`0.0.0.0`, which would publish it on whatever network the client happens to
|
||
be on.
|
||
|
||
### When the router goes unreachable, start at docs/incidents.md
|
||
|
||
Eight incidents so far, nearly all sharing one shape: a change that looked local
|
||
to the router silently degraded the agent depending on it, and none announced
|
||
itself as a router problem. **`docs/incidents.md` carries the full write-ups plus
|
||
a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
|
||
#8 is the exception worth knowing before an unattended agent run: the router
|
||
worked perfectly while agent workers looped for hours with no progress, and no
|
||
check noticed, because every check watched spend or availability rather than
|
||
whether changes landed (`plans/no-progress-detection.md`).
|
||
|
||
Two conventions from those incidents that bind every session, and so stay here:
|
||
|
||
- **8080 is production, always.** It is baked into `opencode.json`, the systemd
|
||
unit, every curl example here, and the admin frontend's own fetches. A
|
||
throwaway instance (manual iteration, Playwright smoke tests, anything that is
|
||
not "use the real router") binds **8081**. Never send a kill signal to a
|
||
process matched by name or port rather than by a PID you started yourself --
|
||
`Restart=always` will fight you, and on this repo it may be your own model
|
||
access.
|
||
- **Never point `config/config.yaml` at test fixtures.** It is the file the live
|
||
service reads. Pass a different config file, monkeypatch `cfg.database.path`
|
||
in-process, or use a temp copy.
|
||
|
||
Recovery for an unreachable-but-`active` service is
|
||
`systemctl --user restart llm-router.service` -- a hung process was never in a
|
||
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
|
||
|
||
The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`),
|
||
separate from the dispatcher, so a stuck router cannot silence the thing meant
|
||
to notice it is stuck. Install and enable it with `systemctl --user enable
|
||
--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a
|
||
placeholder home path and must be `sed`-repointed to the real one first), or
|
||
run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md)
|
||
for the signals it fires on, the alert lifecycle, and the known limits.
|
||
|
||
## Pointing a coding agent at it
|
||
|
||
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
|
||
opencode, an SDK, plain curl. Repo-local `opencode.json` is already wired up,
|
||
so running `opencode` from a clone of this repo routes by default. For global
|
||
use, merge `provider.llm-router` into `~/.config/opencode/opencode.json`.
|
||
|
||
| model name | behavior |
|
||
|---|---|
|
||
| `auto` | router picks; flex rows excluded so nothing is held during peak |
|
||
| `auto:batch` | router picks; flex rows admitted, for overnight/async work |
|
||
| any real model id | dispatched as asked, still logged |
|
||
|
||
Streaming is proxied chunk by chunk rather than buffered, so tokens still
|
||
render as they arrive. NeuralWatt emits its energy and cost blocks as SSE
|
||
**comment** lines (`: energy {...}`) before `data: [DONE]` — ordinary clients
|
||
ignore comments, so the stream passes through untouched while the router
|
||
reads the telemetry on the way past. Without that, streamed calls would log
|
||
no energy at all, which is most of the point of this project.
|
||
|
||
## Try it
|
||
|
||
Copy-pasteable `curl` examples (`/route`, `/dispatch`, `/v1` models + chat) and
|
||
the classifier-skip overrides (`task_category`, `task_tier`,
|
||
`required_context_tokens`) live in [README ## Usage](README.md#usage). Inspect
|
||
what a dispatch cost/burned with the `sqlite3` query in
|
||
[docs/operations.md](docs/operations.md).
|