Files
6krrt/CLAUDE.md
adlee-was-taken dfd19d6fef docs: revert stray CLAUDE.md line, correct classifier coverage wording
Reverts 2d145d6, which appended "CLIREF.md updated to CLAUDE.md for current
project state mapping" to the end of CLAUDE.md. It is a stray sentence
describing a rename that never happened.

Correction for b774d55: its subject and body say "CLIREF.md". No such file
exists in this repository; the 17 lines that commit describes (the North Star 1
classifier-coverage paragraphs) are in CLAUDE.md. History is not rewritten.

Also rewords the first of those paragraphs: "partially covered" understated
it. Every classifier scalar is now a control, card-backed, or a recorded
excuse, and fallback_category / max_input_chars are persisted-only.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
2026-10-05 01:43:40 -04:00

1413 lines
84 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Local LLM Model Router — project brief
`README.md` is the concise front door. Deep-dive module-by-module reference
docs live in **`docs/`**.
`design/local-llm-model-router.md` holds the architecture and rationale,
including parts still unbuilt. This file is the working state + immediate
next steps, and is the one to trust on what is currently true.
## NORTH STAR GUIDELINES
Four rules that outrank local cleverness. Each exists because it was broken
first and the breakage was expensive to find. When a change conflicts with one
of these, the change is wrong — not the rule.
### 1. Every config knob is reachable from the admin page
Every scalar knob in `config.yaml` gets an admin control — runtime, persisted,
or both as its mechanism warrants — **or** a recorded decision saying why it
must not. The absence of a control has to be a decision someone made, never an
oversight nobody noticed.
`objective.credit_attenuation.enabled` is the model exception: deliberately
absent, because enabling it must be a config edit plus a restart. That is a
recorded choice, not a gap.
The classifier section is now fully accounted for under the gate. Seven
classifier scalars (`context_framing`, `cooldown_seconds`, `fallback_tier`,
`fallback_category`, `max_input_chars`, `degraded_warn_min`,
`degraded_warn_threshold`) have a runtime control, a persisted control, or both
(`fallback_category` and `max_input_chars` are persisted-only). Every other
classifier scalar is either card-backed (`_CARD_BACKED_PATHS` in
`src/admin.py`: mode, cloud primary/fallback, encoder, and local-decision
fields) or excused in `DELIBERATELY_NOT_IN_ADMIN` (timeout, temperature,
`max_output_tokens`, `encoder.tier_from_features`).
Outside-gate knobs are still pending. Sections not yet reached by the portal
(database, dispatch providers/settings, local dispatch models, profiles,
local energy, tiers/tiering/proficiency/context, and deployment wiring inside
classifier/verification/local vision) are tracked in
`plans/deferred-knobs.md`. The first control added under any of those sections
will drag every scalar under it into scope at once, per the coverage test's
clause 1.
**Why:** Wave 2 shipped `incumbent_cache_pricing` and
`incumbent_challenger_cache_rate` with no control at all. Nobody decided that;
it just never came up. The dial's entire purpose is tuning from neutral to full
*without reverting code* — and hand-editing a tracked file and restarting is a
loop nobody walks, so the knob's rationale evaporated on contact with reality.
Enforced by `tests/test_admin_knob_coverage.py`, which fails naming the knob.
Honour it; do not add a `DELIBERATELY_NOT_IN_ADMIN` entry to silence it unless
the reason is true.
### 2. Fix classifier latency by improving the classifier, never by reusing a
stale decision
Classification latency is a real problem and the answer is a faster or better
classifier — a smaller model, a warmer process, a cheaper backend. The answer
is **never** to let one classification stand in for later, different work.
A stale label does not merely add noise. It silently redirects money: the whole
point of this router is sending each task to the model that fits *that task*,
and a replayed label routes the next task to whatever fitted the last one.
**Why:** the session cache reached **96.6% of all classifications**. One real
classification drove **107 consecutive turns** across 14 minutes on 2026-09-16.
It was introduced to avoid paying a 2-5s round trip per turn, which was a
reasonable trade in isolation and became the dominant path without anyone
choosing that.
Two costs, and the second is worse than the first:
- **Routing** decides on what the session was doing when it started, not what
this turn is. A mid-session pivot — docs, then bug fixes — routes the bug
fixes on the docs label.
- **Proficiency** is trained on those labels, because `POST /outcome`
attributes to `(model, task_category)`. That is part of how
`deepseek-v4-flash` came to hold a saturated 1.000 on `file_summarization`
while failing 50% of them on real traffic.
It also made a working classifier look broken: replaying a handful of session
labels across hundreds of turns produced "pages and pages of `diff_checking`",
which read as a classifier stuck on one label when the underlying
classifications were a reasonable mix.
If latency forces a cache, that is a **measured, time-boxed concession with an
expiry condition written next to it** — not a default.
### 3. Check worktrees and other agents' work before touching any file
Before editing, run `git worktree list` and check for other agents or sessions
working in the same tree. Do not assume a worktree is yours. When parallel work
is unavoidable, isolate it — a separate worktree per agent — and commit by
explicit path, never `git add -A`.
**Why:** two agents were once launched into the *same* worktree while each was
told the other was elsewhere. It survived only because the files were disjoint
and one of them committed by explicit path rather than sweeping the tree. The
second agent's test count was also measured against a tree containing the
first's uncommitted work, so its "verified green" was not a clean signal.
Related: a `router.db` inside a worktree is a **stale copy**, not the live
database. The live one is `/home/alee/Sources/6krrt/router.db`. Check mtimes
before measuring anything, and open it read-only:
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
not accept `-uri`.
### 4. Waste is surfaced in the admin portal, and stopping it is one click
When the router can see money being wasted, it shows the operator where they
already look, with the evidence and the lever to stop it side by side. A
detector that only writes a log line, or a fix that needs a config edit, a
restart or a long table scan, has not met this rule.
"Waste" here means **spend with no concrete change landing**: an agent session
looping, re-reading, or retrying the same failure, and a model that keeps
producing such sessions. It does **not** mean steady spend. A healthy agent run
can burn for hours, and a spend-rate alarm cannot tell the two apart.
In practice:
- **Stalled sessions are visible** in the portal with their evidence: turns and
$ since the last landed change, and the top repeated target. They also reach
the operator when nobody is watching (desktop alert).
- **A model that keeps producing them is visible** as a per-model rollup, with
**Block model** beside the evidence. The block records its reason, shows in a
Blocked list, and is one click to undo.
- **Automatic responses come after the visible one,** never instead of it, and
every automatic action shows up in the same place.
**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on
2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec
68 times its length, and a model confabulated truncation that was not there.
Every existing check stayed green. It was caught only by a human, or a Claude
session, reading opencode's session store by hand. Pulling the model took a trip
through a dropdown whose vocabulary is catalog `deprecated`. The detection
existed nowhere, and the lever existed only for someone who already knew.
## What this is
A router that uses a local model (served via Ollama) to classify incoming
coding/documentation tasks — category, tier, required context size — and
dispatch each task to the best-fit open-weight model on **Neuralwatt Cloud**,
under a per-request cost ceiling, tiebroken by price, and ranked by
category-level expected pass rate.
Every measurement in this file was taken on one deployment against one
provider account. They are recorded because the reasoning is worth more than
the conclusion, but treat them as observations with a date on them, not as
constants — the catalog, prices, grid intensity and pool load all move. When
a number here decides something, re-run the measurement before trusting it.
Neuralwatt remains the primary provider. OpenRouter is back as an opt-in
provider gated by the `provider_model_allowlist` table — a provider with
`require_allowlist: true` is ignored unless its model_id is explicitly
allowlisted, and any previously-upserted row that drops off the list is
marked `deprecated` on the next poll. The `provider` column and the
`(model_id, provider)` primary key are unchanged, so this adds no migration.
## Stack
- Python (chosen over Rust — this is I/O-bound against provider APIs, not
CPU-bound; iteration speed on the scoring/weighting logic matters more than
raw execution speed at this scale)
- SQLite for the decision table
- Ollama for local classification, via any OpenAI-compatible endpoint —
`localhost:11434/v1`, or an Ollama on another machine across a VPN
(`classifier.base_url`)
- FastAPI for the dispatcher service
## Billing is per-kWh, not per-token — and neither is what scoring uses
**Measured against the live API, 2026-08-11.** Neuralwatt bills a flat
**$8.00 per kWh** and the catalog's `input_per_million` /
`output_per_million` prices are not what this account is charged. Confirmed
across five models; `cost_usd / energy_kwh` came back 8.00 every time:
| model | list $/1M out | completion tokens | billed USD | kWh | $/kWh |
|---|---|---|---|---|---|
| deepseek-v4-flash | 0.28 | 600 | 5.00e-06 | 5.88e-07 | 8.50\* |
| gemma-4-31b | 0.42 | 540 | 3.80e-04 | 4.75e-05 | 8.00 |
| qwen3.6-35b-fast | 1.15 | 600 | 2.53e-04 | 3.16e-05 | 8.01 |
| kimi-k2.7-code-fast | 4.00 | 600 | 2.25e-03 | 2.81e-04 | 8.00 |
| kimi-k3-fast | 15.00 | 539 | 2.17e-04 | 2.71e-05 | 8.00 |
\* rounding — billed cost is quantized to ~1e-06.
The precise rule, validated against all 65 samples of the reference sweep
(61/65 within 2%; the 4 outliers are microdollar rounding, not misses):
```
cost_usd = min( $8.00/kWh x energy_kwh , 3 x list token price )
```
The ceiling bound in only 2 of 65 samples, both `deepseek-v4-flash` energy
spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x
$0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349.
**List price ranks models backwards.** Not approximately — invertedly:
| | list $/1M | actually billed | gCO2eq |
|---|---|---|---|
| `deepseek-v4-flash` | **0.28** | 9.80e-05 | 8.90e-04 |
| `gemma-4-31b` | 0.42 | **5.02e-05** | **2.32e-04** |
deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon.
**But cost and eco are NOT the same axis** — the tempting simplification, and
it is wrong. Cost tracks energy, but carbon is energy x the serving region's
grid intensity, and models run in different regions:
| grid | gCO2/kWh | models |
|---|---|---|
| `FI` | ~49-50 | most of the catalog; varies by time |
| `FI` (reported) | 475 | `glm-5.2-fast`, `glm-5.2-flex` |
| `US-MIDA-PJM` | ~442 | the `kimi-k3` family |
A 13.6x spread, so the two axes disagree: `glm-5.2-fast` is the 2nd cheapest
model and only the 6th cleanest; `kimi-k3-flex` draws 3.7x *less* energy than
`kimi-k2.7-code` while emitting 3.6x *more* carbon. Weighting them separately
is load-bearing, and `tests/test_routing.py` pins it.
**Superseded — cost no longer comes from the sweep at all.** `cost` was the
*median* measured USD over the reference sweep. That was measured to be WRONG
for real traffic, because the reference workload is the wrong shape.
The sweep sends a 400-token prompt with a 400-token completion. Real agent
traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache
hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the
figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts:
| workload | winner |
|---|---|
| reference sweep (400/400) | `glm-5.2-fast`, 3.2x cheaper |
| realistic (70k prompt, short answer) | **`deepseek-v4-flash`, 5.0x cheaper** |
Same two models, opposite answer. `glm-5.2-fast` sits at attribution 0.006 on
a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts
and badly on real ones. `deepseek-v4-flash` barely moves (0.21 -> 0.25).
So `routing.estimated_cost` prices each request from **catalog token prices,
scaled to that request's actual shape** (prompt size, assumed completion
length, `assumed_cache_rate`). List price is not what gets billed, but billing
is capped at 3x list, so it tracks the real ordering and bounds it — and on
the one case that was checked live it agrees with the measurement in direction
and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no
sweep, and refreshes whenever the poller runs.
`objective.plan_kwh_per_period` is a planning figure only: per-request traffic
is **never** refused for exceeding it — it **gates nothing**. Overage is billed
against the account's credit balance (`allowance_remaining_usd` from the
provider). The `/admin/api/snapshot` endpoint reports balance, estimated burn
rate, and projected runway **per provider** inside `quota.accounts[]` — a
list of per-provider billing shapes (`metered_plan`, `prepaid_credit`,
`self_hosted`, or `unmetered`) with `plan`, `pool`, `burn`, and `credit`
blocks as appropriate. `quota.spend` aggregates provider spend and a
list-price estimate. The old flat keys and `by_provider`/`total_balance_usd`
shape were removed; every consumer was updated in the same change, so there
are no deprecated aliases.
Three signals said `deepseek-v4-flash` — catalog token price (7.8x cheaper),
NeuralWatt's own published per-request energy (~10x lower), and a live 70k
measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the
workload you actually run.
`eco` still comes from the sweep's median gCO2eq, and is still not an
objective. `flex_cost_multiplier` is gone: a flex row's measured cost already
is its flex cost.
**Open, and worth knowing:** NeuralWatt's model cards publish *gross* energy
(~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is
gross x attribution. GLM burns roughly 7x more actual electricity per request
and charges ~5x less, because far more tenants share its GPUs. Anything built
on `eco` inherits that inversion — the attributed carbon figure answers "what
is my share", not "what was burned".
## Energy attribution: signal that looks like noise
Billed energy decomposes exactly:
```
energy_kwh = avg_power_watts x duration_seconds x attribution_ratio
```
`attribution_ratio` is the request's share of a shared multi-tenant GPU pool.
Up close it looks like pure noise — eight rapid identical calls to one model
spanned 20x in billed energy, correlating **+0.997** with the ratio while
power and duration held steady. Two sweeps of the same 13 models with the
same prompt disagreed by up to 36x.
Scoring on the pre-attribution product (`power x duration`) was tried, and it
is **wrong**. Across the sweep:
| | spread |
|---|---|
| median attribution, **between** models | **750x** |
| typical spread **within** one model | **1.8x** |
The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving
concurrency, a stable per-model property, not weather. A model whose GPUs
carry far more concurrent requests genuinely costs less per request, and
that is most of the real cost difference in the catalog: `deepseek-v4-flash`
bills ~1000x under its share of pool gross. Stripping attribution discards a
750x real signal to suppress a 1.8x one.
So scoring reads the attributed figures, and the **median** absorbs what
noise remains. A split-half check on the 7-sample sweep (median of first
three vs last four) shows that working:
- **10 of 13 models agree within 1.4x** — stable enough to route on
- **3 do not**: `kimi-k2.7-code-fast` (29x), `kimi-k3` (14x),
`glm-5.2-flex` (2.2x). Those need more samples before their position is
trustworthy.
`dispatcher.gross_energy_kwh` remains as a diagnostic on the identity, not a
scoring input.
### Attribution drifts across hours, so sampling must too
Within about 30 minutes the billed figures reproduce (0.3-1.1x on a
spot-check). Across hours they do not: between two sweeps,
`deepseek-v4-flash` moved roughly 50x and `qwen3.6-35b` about 7x the other
way — enough to **invert their cost ranking**. Attribution tracks pool load,
and pool load tracks time of day.
More samples inside one sweep does not fix this; it measures one moment more
precisely. Coverage across time does. `load_candidates` already takes the
median over ALL `seed_reference` rows, so repeated sweeps accumulate into a
median-across-time for free — hence `llm-router-seed.timer`, which runs a
small sweep every 6 hours.
Until several sweeps have accumulated, treat the eco ordering as provisional.
A single sweep's ranking is one sample of a moving quantity.
**And none have accumulated since 6e729ad.** That commit moved
`log_observation`'s trailing arguments to keyword-only without updating
`seed_energy.py`, so every timer run since spent one billed completion and then
died on `TypeError` — which is not a `RequestException`, so the per-sample
`except` did not catch it. Fixed, and the sweep now has an offline end-to-end
test, but the accumulation this section describes starts from the next run
rather than from months of history.
## What's built and working
- `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md).
- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below.
- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md).
- `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md).
- `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md).
- `tiering.py` / `tier.py` — pure tier resolver + DB pass. Why tier on `reasoning_default_enabled`, cheapness-not-ceiling, `tier1_context_max`: [routing#tiering](docs/routing.md#tiering).
- `routing.py` — pure hard filters + ranking, plus the request-side capability gates (fail-closed asymmetry). [routing](docs/routing.md).
- `routing.py` `rank_candidates` incumbency: prices the session's last chat model at its measured cache rate and every challenger at the `objective.incumbent_challenger_cache_rate` dial, behind a load-bearing `min(dial, incumbent_rate)` clamp; gated off by default (`incumbent_cache_pricing: false`), tunable from off to full in config. [routing#incumbency-and-cache-pricing](docs/routing.md#incumbency-and-cache-pricing).
- `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates.
- `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below).
- `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md).
- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded.
- `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md). `--dry-run` no longer just describes the fold, it **projects** it: `feedback_preview.py` copies the DB into memory, runs the real `add_outcome` against the copy, and reports the per-row before/after, opening the source `mode=ro` so a preview cannot write to what it is previewing. `deploy/llm-router-feedback.{service,timer}` gives this loop the timer it never had — and ships **not enabled**, because the fold is irreversible and the first one against an accumulated backlog is an operator decision. See `deploy/README.md`.
- `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md).
- `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md).
- `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable.
- `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md).
- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. [api](docs/api.md).
- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. The decision table leads with a `time` column and carries `profile` plus an `E` flag for exploratory picks; the quota panel now shows per-account billing shapes (`metered_plan`, `prepaid_credit`, etc.) with plan/pool/burn/credit blocks, spend aggregates, and an alarm line. [architecture](docs/architecture.md).
- `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted.
- `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.
- `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md).
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor.
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
- `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md).
- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md).
- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md).
- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md).
- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md).
- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md).
- Agent guardrails — `scripts/verify_commit.py` (is this commit good?),
`scripts/oc_dispatch_audit.py` (audit an orchestrator session's dispatches),
and the opencode plugin `deploy/opencode-plugin/guardrails.js` whose
`tool.execute.before` hook blocks rule-breaking tool calls.
[agent-guardrails](docs/agent-guardrails.md).
## #45 — OpenRouter is an opt-in allowlist provider
OpenRouter used to be ingested whole, then removed, and is now back — but only
as an opt-in provider. The base config sets `openrouter.require_allowlist: true`
and ships a short seed allowlist. Models on that seed list are upserted and
kept active; anything else in the OpenRouter catalog is filtered out before
upsert and any previously-active OpenRouter row that is not on the list is
marked `deprecated` on the next poll.
This is deliberately different from the old ingest-everything behavior. The
previous approach once pulled in a non-chat model (`lyria/...`) that returned
HTTP 404 on dispatch because the endpoint expected chat completions. The router
had paid for the classification, selected the model, and then failed on the
provider call. Allowlist-gating prevents that class of failure by default: if a
model id has not been reviewed and explicitly added, the router acts as if it
does not exist.
The seed allowlist is short and has firm exclusions. It does NOT include:
- `x-ai/*` (Grok)
- `openai/*`
- `anthropic/*`
Those exclusions are non-negotiable. They are not "currently excluded" or
planned for future inclusion; they are deliberately absent from the seed list.
Adding one requires editing both the seed allowlist and this file.
The admin portal exposes the allowlist under `/admin`: the providers page shows
which providers require one, and each provider row links to an allowlist editor
where entries can be added or removed. The underlying four API endpoints are
`GET /admin/api/providers/{name}/allowlist`,
`POST /admin/api/providers/{name}/allowlist`,
`DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers
page itself surfaces `require_allowlist` with a link to the allowlist editor.
Neuralwatt remains the primary, ungated provider. The allowlist behavior only
fires for providers with `require_allowlist: true`.
## Proficiency: category now changes routing
`proficiency_score` is the ONLY category-dependent term in the ranking, so
until this table had data, `task_category` could not change a decision at
all — the classifier computed it, the router paid ~10s for it, and then it
made no difference. It does now. Two categories were added for local dispatch:
`file_summarization` and `diff_checking`; see [evaluation](docs/evaluation.md). The score is also no longer a raw benchmark
level: it has been converted into an **expected pass rate on real traffic**,
calibrated against 1,059 client-reported outcomes and shrunk with a
20-pseudo-observation prior so thin data does not dominate.
`proficiency` now holds 141 rows:
| source | count | meaning |
|---|---|---|
| `outcome_blended` | 33 | Fresh per-model outcome evidence |
| `outcome_prior` | 76 | Trafficked-sibling rows inheriting the peer-rate prior |
| `self_eval_thin` | 32 | Cold categories (summarization, translation); benchmark preserved verbatim |
A further 33 (model, category) pairs have direct outcome samples. The biggest
evidence gains went to `deepseek-v4-flash` (notably `coding_general`,
`coding_refactor`, and `general_chat`), `kimi-k2.7-code` (across 7
categories), and `qwen3.6-35b`.
Sweeping 9 categories x 3 tiers currently returns **5 distinct winners** at
both 50k and 120k of context.
| context | winners over 27 decisions |
|---|---|
| 50k | `qwen3.6-35b` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) |
| 120k | `kimi-k2.7-code` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) |
**This spread is recent, and how it got here is the useful part.** For a long
time all 27 decisions returned ONE model, and that was the correct answer at
the time rather than a bug: with cost and eco both populated, `qwen3.6-35b`
was Pareto-dominant — cheapest AND cleanest in the routable set, while
scoring within `quality_tolerance` of the best. No defensible weighting picks
anything else out of that.
Two corrections widened it, and neither was a tuning change:
- **Cost stopped being a benchmark average.** It is now priced per request
from catalog prices scaled to the request's shape, so the ranking depends
on the workload instead of on a 400-token reference sweep that no real
traffic resembles.
- **Tier stopped being inferred from price.** `deepseek-v4-flash` was pinned
to tier 1 for being cheap, which excluded it from every tier-2 request
regardless of what any score said.
A third shift is under way: the outcome backlog has been spent, so the score
now reflects real pass/fail reports rather than the benchmark alone. That
changes the numbers; it does not change the rule. Quality is still the
objective and cost is still the tiebreak within `quality_tolerance`.
Note what changes between the two rows above: only the leader, and only
because of the hard context filter. That is the filter working, not the
scoring disagreeing with itself.
**If you see one model win everything again, check for dominance before
reaching for config.** One winner is a legitimate outcome. The levers, if a
genuinely different balance is wanted, are `objective.quality_tolerance`
(how large an expected-success-rate gap must be before it outranks a cost
saving) or `objective.max_energy_per_request` (a hard ceiling). There is no
weight to tune.
### What the task set actually found
**The benchmark could not discriminate these models on coding.** Every row
scored exactly 1.00 on `coding_general`, `coding_refactor` and `debugging` —
and that is after the tasks were deliberately hardened with touching
intervals, full semver, present-but-falsy defaults, late-binding closures and
a binary search that infinite-loops. Every model in this catalog is simply
good at that class of problem, so cost decides coding routes, which is the
right outcome.
**Real traffic broke one of those ties, which the benchmark never could.**
`coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29
samples folded in by `feedback.py` from an actual agent session, and crossed
`self_eval_min_samples` on the way, so it reads `self_eval` rather than
`self_eval_thin`. That is the intended shape of this system — the 43-task
benchmark establishes a floor, and your own traffic is what refines it.
`coding_refactor` and `debugging` are still flat at 1.00, awaiting the same
treatment.
**Benchmark-sourced hardening landed, and it broke both remaining ties.** The
task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after
the `score_exact` literal-eval fix), and three Exercism refactor/debug pairs.
The CRUXEval-O rows split `coding_general` into a 0-1 mix across models — five
of six now fail at least one model — and the Exercism refactor rows moved
`coding_refactor` off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00,
affine 0.56-1.00. The debug pairs split `debugging` the same way (bowling
0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted:
`debug_affine_coprime` and most of the BFCL tasks sat flat at 1.00, so they do
not discriminate.
**A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2
samples per model the category read 0.70-1.00 with a model at the ceiling, and
the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more
benchmark passes moved every score and left NOTHING at 1.00:
| model | n=2 | n=11-14 |
|---|---|---|
| `kimi-k3` | 0.85 | **0.973** |
| `kimi-k2.7-code` | 0.85 | 0.886 |
| `deepseek-v4-flash` | 0.80 | 0.864 |
| `kimi-k3-fast` | **1.00** | 0.864 |
| `qwen3.6-35b` | 0.85 | 0.800 |
| `gemma-4-31b` | 0.85 | 0.786 |
The winner moved to `kimi-k2.7-code`, **3.2x cheaper** at 50k of context
($0.0441 -> $0.0136), with no config change — `kimi-k3` scores higher but sits
inside `quality_tolerance`, so cost breaks the tie. `deepseek-v4-flash`
($0.0024) misses the band by 0.009, which is the kind of margin the tolerance
exists to describe rather than a verdict.
**The whole spread rests on one rubric line, though.** `docs_function` is
effectively saturated — 1.00 on nine of every ten samples — and nearly every
`docs_gotcha` deduction is the same omission: the model documents that order is
preserved, that the first occurrence is kept, and what `key` does, then never
says elements must be hashable. That is real discrimination, since it is a real
property of the function, but one sentence is deciding a category. Treat this
ordering as thinner than n=14 makes it look.
**The self-judging guard costs sample density, and it shows up here.** Most
models reached n=14; `kimi-k3` and `kimi-k3-fast` reached only 11, because
those two are the ones diverted to the alternate judge `qwen3.6-35b`, which
returns unparseable JSON more often than `kimi-k3` does. The guard is still
right — a model grading its own family is worse than a thinner sample — but
the alternate judges should be picked for parseability, not just for being
someone else.
The reflex when a category looks flat is to reach for `quality_tolerance`.
Neither tie broken so far was broken that way: `coding_general` opened up when
`feedback.py` folded in real traffic, and `docs_writing` opened up on six more
benchmark passes. Both were samples, not settings. `coding_refactor` and
`debugging` are still flat at 1.00 on 2-3 samples each — which is now a state
this project has mistaken for a measurement once.
Current spread by category, widest first:
| category | spread |
|---|---|
| `tool_use_agentic` | 0.33 - 1.00 |
| `summarization` | 0.60 - 1.00 |
| `reasoning_math` | 0.67 - 1.00 |
| `docs_writing` | 0.66 - 0.97 |
| `general_chat` | 0.80 - 1.00 |
| `translation` | 0.85 - 1.00 |
| `coding_general` | 0.86 - 1.00 |
| `coding_refactor`, `debugging` | flat at 1.00 |
**What does discriminate is tool use, arithmetic traps, and prose.**
`deepseek-v4-flash` scores 1.00 on all three coding categories yet **0.33 on
`tool_use_agentic`** and 0.67 on `reasoning_math`. Verified live, not an
artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes
away?" — both times supplied — it calls *two* tools rather than subtracting.
It over-reaches for tools, which is exactly the failure mode that matters in
an agent loop. The router now avoids it for those categories while still
picking it for coding.
Most rows still read `source='self_eval_thin'` (118 of 132): real
measurement, but below `self_eval_min_samples` at 2-3 tasks per category per
run. The 14 that have crossed it are all `docs_writing`, from the six extra
passes above. Two paths thicken it, and they are complementary — re-run
`eval_proficiency.py` to accumulate benchmark samples, or just use the router
and let `feedback.py` fold in real outcomes. Both fold into a running mean
rather than replacing, so samples add up across runs.
### A score is only as fresh as the row it was copied to
Proficiency is a property of the weights, not the queue, so the eval harness
scores one row per family and `propagate_to_variants` copies the result onto
the serving variants — `kimi-k3-flex` gets `kimi-k3`'s number, because no
benchmark rates a `-flex` row separately.
That copy used to happen **exactly once per variant, ever.** The guard skipped
any row with `self_eval_samples > 0`, meaning "measured directly, do not
overwrite" — but inheritance copies the sample count too, so after the first
propagation an inherited row was indistinguishable from a measured one and was
never refreshed again. `kimi-k3-flex` sat at 0.85/n=2 while `kimi-k3` moved to
0.973/n=11.
`proficiency.inherited_from` records the provenance that was missing, and the
**migration** was the delicate half, not the fix: `ADD COLUMN` gives every
existing row NULL, which reads as "measured here", so shipping the guard alone
would have permanently frozen the exact rows it exists to unfreeze. The
backfill infers provenance from the harness's own selection rule rather than
guessing — `eval_identities` only ever evaluates standard rows plus flex rows
with **no** standard equivalent, so a flex row that has one was never a
candidate for direct evaluation, whatever its sample count claims. Everything
else keeps NULL, which fails safe: NULL means "do not overwrite", so no real
measurement can be lost to a wrong guess.
Confirmed on the live database, and on the catalog's one genuine exception —
`glm-5.2` is canary, so `glm-5.2-flex` is the routable row the harness scores
directly, and its NULL is correct.
### Harness bugs this shook out
Three separate defects, each of which scored the rig rather than the model,
and each caught by reading per-task detail rather than the summary:
- **Token budget.** `max_tokens` was shared between a reasoning model's trace
and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and
returned an EMPTY content field, scoring 0.00 on tasks it can plainly do.
Now 24000, clamped per model (gemma-4-31b caps at 16384), and
`finish_reason: length` skips the sample instead of scoring it.
- **One leading space.** kimi-k2.7-code returns `" def f(...)"`, which becomes
IndentationError once the harness prepends its imports — 0.00 across all
nine coding tasks for a model with "code" in its name.
- **Judge failures scored as model failures.** 44% of judge calls returned
unparseable output (the judge is itself a reasoning model and leaks its
thinking despite `response_format`). Each was recorded as 0.0. Now the JSON
is extracted from surrounding prose and an unusable reply yields no sample.
`tests/test_task_set.py` exists so this stops happening: it implements a
reference solution for every `code` task and asserts it passes every check,
recomputes every `exact` answer (one by brute force), and confirms each
refactor target already passes its own checks while each debugging target
fails. It immediately caught a check where the expected value was simply
wrong — which would have docked every model on a task and been
indistinguishable from genuine difficulty.
## Tool competence is read from the request, not guessed at
Neither local classifier can identify agentic work. Asked to label six
unambiguous tool-use prompts ("read the config then update the manifest",
"run the tests and fix what fails"), `qwen3.5` got 2/6 and `mistral-nemo`
1/6 — and `mistral-nemo`'s misses collapse to `general_chat`, which is also
the configured `fallback_category`, so qwen3.5's crashes land in the same
place.
That mattered because `tool_use_agentic` has the widest proficiency spread in
the table (0.33-1.00) and `deepseek-v4-flash` — the current winner on coding —
sits at the bottom of it.
**The fix was not a better classifier.** Whether tools are on the table is
stated in the request: every agent client sends a `tools` array, and
`chat_completions` never looked at it. Reading it is exact and free.
It is applied as a **hard filter**, not a category override, and the
distinction is load-bearing. The question is not "is this task agentic" but
"can this model be trusted with tools that exist". The recorded failure is
precisely the second one: `deepseek-v4-flash` was given a *non*-agentic prompt
("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times
supplied) and called two tools rather than subtracting. A model that
over-reaches is a hazard on every request where tools are available, whatever
a classifier would have labelled the task.
So `routing.min_tool_proficiency` drops any candidate whose measured
`tool_use_agentic` score is below it, but only when the request carries tools:
| request | winner on `coding_general` @ 50k |
|---|---|
| no tools | `deepseek-v4-flash` ($0.0024) |
| tools present | `qwen3.6-35b` ($0.0041) |
Verified live through `/v1/chat/completions` with identical bodies differing
only by the `tools` array. The cost of safety here is 1.7x on that route,
paid only where tools exist.
**It is currently set to `null`, i.e. OFF**, deliberately and pending
experiment. opencode sends `tools` on essentially every request, so with the
filter on, `deepseek-v4-flash` is excluded from ordinary agent traffic and its
~7x cost advantage goes unused; with it off, that advantage applies and a
model measured at 0.33 on tool use handles requests where tools are on the
table. Which is right is an empirical question and the benchmark cannot
answer it — the 0.33 comes from 3 tasks.
What settles it is `POST /outcome`: run with the filter off, let real pass/fail
reports accumulate, and compare `deepseek-v4-flash`'s `tool_use_agentic`
proficiency before and after. That is the one signal here that knows whether
the work actually worked, and `feedback.py` folds client outcomes in both
directions, so success counts too.
**Update 2026-09-15: that accumulation path is closed.** The classifier no
longer emits `tool_use_agentic` — per-turn classification collapsed onto it
(31 of 31 consecutive live turns, all routed to the slowest model in the
catalog) — so no new outcome attributes to the category and its scores are
frozen at today's values. Accepted, not fixed; the reasoning and the rejected
alternatives are in [docs/routing.md](docs/routing.md), "The classifier's
labels are not the proficiency scoring axis". The filter experiment now either
finds a different signal or reads frozen data.
0.5 sits in the empty band between the only two values the catalog holds
(0.33 and 1.00), so it is not fitted to either. A model with **no** measured
tool score is unproven rather than proven bad and is not dropped — the same
rule as the tier-1 context gate. Config load refuses a
`routing.tool_use_category` that is not a real category, because a name
matching nothing yields NULL for every row and NULL means "do not
disqualify": the filter would silently stop filtering.
## Tier is an iteration budget, not just a floor
A tier used to mean only "do not route below this". It now also buys
corrective attempts after a verification failure:
| tier | batch | interactive |
|---|---|---|
| 1 | 0 retries | 0 |
| 2 | 1 | 1 |
| 3 | 2 | 1 |
Interactive is capped below its tier because every retry doubles
time-to-answer, and in interactive use latency **is** a quality loss.
Retries are matched to the failure, since the causes differ:
- **truncated** — raise the token budget on the same model; a different one
would run out too. If there is no cap to raise, the model's own output
ceiling is the wall, so escalate to a candidate that can emit more.
- **malformed** — more tokens will not make unparseable output parse, so
escalate to the next-ranked candidate.
- **ok / unverifiable** — buy nothing. Retrying `unverifiable` would burn
quota across the majority of prose traffic for no signal.
`escalation.preemptive_on_low_confidence` is now **off by default**. Bumping
the tier because the classifier was unsure pays frontier prices before
anything has gone wrong; spending after a check has actually failed is better
on both mandates — the cheap attempt usually succeeds, and when it fails you
have evidence rather than a hunch.
## The only ground truth: `POST /outcome`
Everything else the router records is a proxy. Structural checks know whether
code *parses*. The local checker guesses whether prose *looks* right. Neither
knows whether the answer did the job — the client does, because it ran the
tests.
```bash
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
```
Two things make this the highest-value signal available:
- **It is the only quality signal that survives streaming.** A retry cannot
reach a streamed response because the bytes are already gone; a report
arrives afterwards and works either way. Every agent client streams.
- **Its successes count.** `feedback.py` folds `ok: true` and `ok: false`
into `proficiency.outcome_score` via `add_outcome()`. Structural and
`local_llm` verdicts are now diagnostics only; they no longer move the
score, because a structural failure is not a verified client outcome.
Client outcomes are what calibrate the routing score. The benchmark is a
prior; `POST /outcome` is the posterior.
`request_id` is written back to `route_decisions.request_id` on both the
streamed and buffered paths, so the report joins cleanly to the decision that
produced it. An unknown `request_id` returns 404 rather than being quietly
accepted — a client whose reports go nowhere should find out.
## Verification: what local compute is actually good for
Local inference is a poor substitute for cloud completions here — 7-82x the
energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s).
But it is very good at stopping a cloud completion from being wasted, and
completions are where all the money is: fitted on real traffic, a completion
token costs **201x** a prompt token.
| check | cost | what it catches |
|---|---|---|
| structural (`verification.py`) | **free** | truncation, malformed code/JSON/YAML, empty answers |
| local LLM (Ollama) | ~$6.9e-05, ~6s | refusals, wrong-question answers, incoherence |
Structural checks run on every response and never execute the code — they
parse it. The local LLM check runs only on answers above
`verification.min_completion_tokens`, in the background after the client has
its response, because it costs ~15% of a median 193-token answer and only
pays above ~600 tokens.
### An agent turn is not a prose answer
Both checkers had mirror halves of the same blind spot, and real traffic is
what found it. On the first genuine agent session (63 completions, a shipped
feature, 349 passing tests, clean mypy and ruff):
| checker | called it | actually |
|---|---|---|
| structural | `malformed: empty response` x29 | turns that ended in a tool call |
| local LLM | `cuts off mid-sentence` x8 of 9 | same turns, judged from the other side |
A turn that calls a tool has empty or half-finished text **by design**. Both
paths now take `has_tool_calls` and return `unverifiable`; in `verify_response`
that check outranks even `finish_reason == 'length'`, because stopping
mid-sentence at a call boundary is not a budget overrun. `worth_local_check`
declines outright, which also stops paying ~6s of local inference to
mis-grade a tool call.
Had `feedback.py` run before this, it would have applied ~12 false failures to
the models that had just shipped the feature. That is the **fourth** harness
bug in this project that would have scored the rig rather than the model, and
the first caught by real traffic instead of a synthetic test. The pre-fix rows
are kept but set `model_attributable = 0`, so the record survives without
steering routing.
**`client_capped` was also over-applied.** It marked *every* verdict
non-attributable whenever the client set `max_tokens` — and opencode always
does — so real failures were invisible to feedback. A client's token cap
explains a `truncated` verdict and nothing else; it is now scoped to exactly
that.
`feedback.py` folds observed failures into `proficiency`, so routing learns
from your traffic rather than only the 43-task benchmark. Only **failures**
are folded in: a structural 'ok' means the code parsed, not that it was
correct, and recording those as 1.0 would flatten every score toward the
ceiling. Failures the model did not cause — a client's own tight `max_tokens`
truncating the answer — are recorded but excluded.
## The classifier is the latency floor
Every routed request pays a full local classification round-trip before an
upstream token is requested — the classifier is the latency floor. Four settings keep it usable:
`max_retries=0` (SDK retry → 3x wall-clock), `max_output_tokens: 1024` (bounds reasoning), `max_input_chars: 8000` (doesn't need the document; head+tail clamp), `fallback_tier`/`fallback_category` (mid tier, not 502/503), plus `temperature: 0` (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not `0.0.0.0` (Ollama has no auth).
Full measurements: [docs/local-models.md](docs/local-models.md#the-classifier-is-the-latency-floor).
### A classifier failure no longer collapses to one fixed guess
`fallback_tier`/`fallback_category` above is now the LAST step, not the only
one. When the local classifier fails, `classify()` walks a cascade:
| # | step | cost |
|---|---|---|
| 1 | this session's cached classification, **staleness ignored** | free |
| 2 | this session's history in `route_decisions` | free |
| 3 | `classifier.cloud_fallback`, if configured | money |
| 4 | `fallback_tier` / `fallback_category` | free |
Steps 1 and 2 reuse a *real* classification rather than guessing. Step 2 is
what survives a restart, when the in-memory cache is gone but the session's
decisions are still on disk. Both filter degraded sources, so one outage's
guess cannot propagate through every later turn of a session and end up
looking like a measurement.
**Step 3 is absent by default and the whole feature costs nothing until it is
configured.** Three things bound it once it is:
- `classifier.cooldown_seconds` (30) is a **global** backoff, so a sustained
outage buys one cloud attempt per window across ALL requests, not one per
request;
- an account-level provider refusal suppresses step 3 for the same window —
an out-of-credit account makes that call a guaranteed wasted request;
- the `session_cache.put()` gate admits `classifier_cloud`, bounding a session
to one cloud call per staleness window instead of one per turn.
`session_stale` and `session_history` are deliberately NOT cached: borrowed
answers must not renew a staleness clock they never earned.
**The trap this had to avoid, and it is the reason to read this section.** A
degraded classification is recorded with `task_category = general_chat` —
which is a *fully scored* category (13 proficiency rows) as well as the
configured `fallback_category`. Without a guard, `POST /outcome` reports on
mislabelled outage traffic fold into `proficiency(model, general_chat)` and
drag real scores toward whatever happened to be flowing while the classifier
was down. So `report_outcome` marks outcomes of degraded-source decisions
`model_attributable = 0` — kept as a record, excluded from folding, the same
treatment the tool-call false-failures got. `feedback.py` is untouched; its
existing `AND model_attributable = 1` already does the work.
Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not:
`fallback`, `session_stale`, `session_history`. The asymmetry is deliberate —
a degraded source is good enough to route one visibly-flagged request, but a
proficiency score is consulted by every future request, so attribution must
not trust more than routing does. Unknown decision rows **fail open**; an
over-applied exclusion already starved feedback once (`client_capped`).
`/metrics` warns when the degraded share of the last 24h crosses
`classifier.degraded_warn_threshold` over at least `degraded_warn_min`
decisions. A survivable failure is exactly the kind that goes unnoticed for
weeks.
**Why a decline is now recorded, and what the numbers said (2026-10-04).**
Under `local_decision` on the 4b the classifier declines about a third of agent
turns on purpose: 769 of 2,557 (30%) that day, 8% the day before, against 0% for
`local_llm` and `local_encoder`. 22.5% were `session_history` replays and 7.6%
the static `general_chat` guess, and `session_history` has no age bound (p99 67
minutes, max 100). None of that could be tuned from the database, because
`route_decisions.confidence` is the override's hard-coded 1.0 on 99.9% of chat
rows. `classifier_confidence`, `classifier_coverage` and `classifier_reject` now
hold the classifier's own numbers and the reason (codes in
[data-model](docs/data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)),
and the degraded-share warning lists the reasons it saw. The warning's default
threshold (0.5) never fires at that rate, so the live overlay sets 0.2.
`degraded_warn_min` and `degraded_warn_threshold` have no admin control, and that
is a deferred gap, not a decision: the `classifier` section is outside the
knob-coverage gate because it has its own card, and adding any `classifier.*`
knob to the generic registries would pull every other classifier scalar into the
gate at once. Do it in one pass when classification is reworked.
### Which implementation is PRIMARY is now a config choice
`classifier.mode` in the live deployment is currently `local_encoder`, set in
`config.local.yaml` with `device: cpu` on the classifier host. That is the
mode answering real traffic right now. The explicit caveat is that a
CPU-vs-CUDA latency and confidence comparison on the live classifier host is
still pending; until that measurement exists, flipping the project default to
`local_encoder` is not decided.
**Turning that mode on for real found three real bugs in one afternoon
(2026-09-06), each a fresh instance of this project's own recurring lesson —
verify against the live system, not the plan.** First, the config-load
validator for `confidence_threshold` didn't exist yet: the admin UI saved a
raw `80` (meant as 80%) straight into the overlay with no conversion, which
would have made every real confidence score read as below-threshold on the
next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability
exceeds 1.0). Caught before the restart, not after — see `confidence_threshold`
in [local-models](docs/local-models.md) for the full incident and the fix.
Second, once that was corrected and the service actually restarted, it
crash-looped twice more before coming up clean: the shipped default model
(`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace
sometime after this project picked it (401 on an unauthenticated GET of its
own model page), and separately `HF_HOME`'s default cache path falls outside
this service's `ProtectHome=read-only` sandbox exception — both are now fixed
(switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo).
Third, and most consequential: the very first real classifications measured
only 5 of 9 test categories correct, because `classify_zero_shot` was passing
raw config identifiers like `tool_use_agentic` and `diff_checking` directly as
zero-shot candidate labels — HF's pipeline scores a label against a hypothesis
template ("This example is {}."), and an underscored code token is not a
sentence the model's NLI training ever saw. Mapping each category to a natural-
language description before scoring, plus `multi_label=True` (the pipeline's
single-label default forces every candidate to compete for the same
probability mass), brought that to 8 of 9 correct with confidence scores
0.77-0.999 on the hits — the one remaining miss scored below the configured
threshold and correctly fell through to the safe fallback rather than
mis-routing.
`classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`,
`local_decision`) picks what answers a classification request — a peer concept to the cascade above,
**not a replacement for it**. Whichever mode is primary, a failure still
walks the exact same cascade (stale session → session history →
`cloud_fallback` → the static guess), unmodified.
- **`cloud_llm`** makes a cloud model the primary attempt, not just the
cascade's post-failure backup. Either pin one (`classifier.cloud_primary`,
same shape as `cloud_fallback`) or set `cloud_primary_auto: true` to
resolve the cheapest currently-routable model **live** against the catalog
(`routing.cheapest_classifier_candidate`, priced for the classifier's own
short-prompt/short-completion shape — not the task's). A success here
records `source="classifier"`, deliberately the *same* string
`local_llm`'s success uses, not `"classifier_cloud"` — that string means
specifically "the cascade's backup step fired" and feeds the `/metrics`
degradation-share warning above as a degraded signal. An intentionally
configured primary succeeding is not degraded.
- **`local_encoder`** classifies with a small, non-generative zero-shot
model instead of an LLM — structurally immune to the runaway-reasoning
failure mode documented above, since there is no generation to run away.
Zero-shot rather than fine-tuned: this router never stores raw task text
anywhere, so there is no training corpus without a new, separate opt-in
capture feature (not built). Only produces `task_category`; `task_tier`
falls back to `fallback_tier` — a real limitation, not a bug. A
below-threshold confidence is treated as a failure and cascades exactly
like a local-LLM parse failure would.
- **`local_decision`** asks a small generative model (configured in
`classifier.decision`) to pick a category from a set of natural-language
descriptions, then optionally classifies tier and — when enabled in its
decision config — runs the same A/B confidence check on the chosen label.
Requires a `classifier.decision:` block in config (confidence threshold,
base URL, model name); it is an opt-in overlay mode, not the default, and
the config load will refuse to start without it when the mode is selected.
Only produces `task_category` and an optional `task_tier` when the
decision block's `tier_classification.enabled` is true. Below-threshold
confidence cascades like a parse failure would.
- **Neither is gated by `local_compute.enabled`** (gaming mode, below) the
way `local_llm` is: `cloud_llm` never touches local hardware, and
`local_encoder` is small enough to run on CPU, so neither competes for the
GPU gaming mode exists to free up.
Configured via `config.yaml` (global default) and overridable per-machine
in `config.local.yaml` — the existing overlay, not a new mechanism — or
through the admin portal's Classifier card, which reports the **live**
resolved primary for `cloud_primary_auto` rather than echoing the config
value (see [admin-portal](docs/admin-portal.md)).
The default in `config.yaml` is still `local_llm`. `local_encoder` is
intentionally an opt-in per-deployment choice rather than the repository
default until the pending CPU-vs-CUDA comparison on the live classifier host is
available. `local_decision` is also opt-in: it requires a `classifier.decision:`
block with its own base URL, model, and confidence thresholds — the config load
refuses to start without it when the mode is selected.
## Local dispatch model
A second local model can now be dispatched directly for specific categories.
`qwen2.5-coder-router:14b` is configured as a tier-1 local row with
`provider='ollama-local'`, gated by `models.eligible_categories`
(`file_summarization` and `diff_checking`). The poller refreshes the row each
run; `seed_local_dispatch_energy.py` derives its price from measured GPU draw
and the user's tariff. Routing treats a local row like any other candidate
once the category filter admits it, and the circuit breaker excludes it on a
local failure so the next request reroutes to cloud candidates.
Known limitations of the local dispatch branch right now:
- **No true streaming.** The response is shaped into an SSE stream, but the
local answer is generated before any bytes leave the router.
- **No verification rows.** Structural and local-LLM checks run but are not
written to `verifications` for local answers.
- **No within-request cloud failover into the upload path is gone.** An eligible
routed request degrades to the local dispatch model when the cloud account
refuses or is exhausted (a fallback, not a preference), so the local row is
no longer a dead-end before a client retry. The degraded answer is recorded
as `kind='local_dispatch_fallback'`.
- **Follow-ups are not special-cased.** A pinned or auto-routed follow-up to the
same local model works, but nothing caches the loaded model between turns.
Dormant under the default profile by design — see docs/routing.md § Local dispatch branch.
`POST /outcome` now attributes through the local energy ledger too. Local rows
include `request_id` and `session_dir` in `local_energy_observations`, so a
client report on a local answer resolves to the same `(model_id, provider,
task_category)` provider-agnostic record as a cloud one.
## Routing notes
Ranking is quality-first, cost as a tiebreak; cost is never allowed to override
a real quality gap. The optional `objective.credit_attenuation` block extends
that tiebreak without changing it: when the block is enabled, a per-provider
multiplier is applied to a candidate's comparison cost only, producing an
`effective_cost` that breaks ties. The multiplier is derived from the provider's
polled account balance (the `balance_url` path, such as OpenRouter), so a low
prepaid balance can nudge a near-tie toward a healthier provider. The logged
`est_cost_usd` and the decision history stay as raw catalog estimates. The
multiplier is 1.0 for providers whose balance comes from per-completion
`allowance_remaining_usd` telemetry (NeuralWatt), so normal overage readings do
not bias routing.
Two semantics matter when reading the numbers. `total_balance_usd` is a sum of
heterogeneous provider-reported readings: OpenRouter's prepaid credits plus
NeuralWatt's overage allowance, which normally reads near -$0.004. It can be
negative and it is not a single spendable figure. `credit_attenuation.enabled`
deliberately lives only in the config file; it is absent from the admin
persisted-config allowlist and from provider edits. Turning it on or off
requires editing `config/config.yaml` and `systemctl --user restart
llm-router.service`, because the dispatcher's `cfg` binds at import time.
## What's NOT built yet — pick up here
Built: session-directory attribution, the local energy ledger, local model
dispatch, admin profiles/proficiency/gaming-mode, and the configurable
classifier backend (`classifier.mode`, including its admin card — all listed
under "What's built and working" above).
**As of 2026-09-05, PRs #26-#36 landed** the gitignored config overlay, admin
profile CRUD writing to it, the provider literal cleanup, quota
balance/burn/runway, capability-aware ceiling and rejection warnings, the TUI
schema catch-up, the classifier fallback cascade, and the admin portal uplift
(proficiency page, profiles duplicate/coverage fix, gaming mode, the
classifier-backoff bug fix). `classifier.mode` (this document's own section
above) is a further, independent addition on top of that.
`plans/multi-provider-support.md` is PARKED on provider selection — Z.ai was
the recommendation and is no longer settled; the coupling surface in it is
measured and still valid.
Known follow-ups recorded but not specced, both small:
- The TUI decision table renders the literal `"None"` in the `ctx` cell when
`required_context_tokens` is absent — the same defect the `profile` cell was
written to avoid. See `.omo/notepads/tui-overhaul/issues.md`.
- A NULL `required_context_tokens` raises `TypeError` inside the
demand-ceiling comparison, and the column is nullable. Latent only: zero
such rows exist today, checked on the live DB.
The items below remain open.
1. **Fine-tuning `local_encoder` on real traffic.** Scoped, not built:
`classifier.training_capture.enabled` (opt-in, off by default — a
deliberate reversal of "never store task text", so it must be impossible
to enable by accident), a `classifier_training_samples` table gated the
same way `report_outcome` already filters proficiency (only rows whose
`classification_source` is in the attributable set), and a
`train_local_encoder.py` script matching `eval_proficiency.py`'s
conventions. `local_encoder.py` currently ships zero-shot only.
2. **Leaderboard priors are unfilled.** `leaderboards.yaml` ships empty on
purpose — inventing plausible-looking benchmark numbers would put
fabricated data straight into routing, the same failure as the provider's
`static_fallback` carbon constant this project already excludes. Until real
sourced figures go in, a newly listed NeuralWatt family has no prior and
relies entirely on self-eval accumulating. `python leaderboard.py --check`
lists what is missing.
3. **Sampling depth for three models — now `eco`-only.** 7 samples/model gives
split-half agreement within 1.4x for 10 of 13, but `kimi-k2.7-code-fast`
(29x), `kimi-k3` (14x) and `glm-5.2-flex` (2.2x) are still unsettled. This
no longer touches cost, which is priced per-request from the catalog, so it
only affects `eco` — which is not an objective. Low priority unless eco
comes back.
4. **Retry does not reach streaming.** The iteration budget (`iteration.py`)
retries after a failed check, but only on the non-streaming path — once
bytes have gone to the client there is nothing to take back. Buffering to
fix that would cost streaming itself, a worse trade for interactive work.
`POST /outcome` is the answer for streamed traffic: it arrives afterwards,
so it works identically either way.
5. **`local_encoder` noise isolation — built for the confirmed shapes; residuals below.**
The raw task's fenced code blocks, `Tool result:`-shaped lines, and closed
`<system-reminder>` spans are now stripped by `_isolate_task_text` (pure,
stdlib-only, deterministic) before the fit + embed pass — `classify_zero_shot`
runs isolate → fit → prefix, so cleaning happens first and a noisy task often
fits the token window outright. Measured on the real `BAAI/bge-large-en-v1.5`
(2026-09-19): 8 clean one-sentence tasks scored 8/8, but the same instructions
wrapped in that noise scored 2/8 with the truncation fix already in place —
and tail-biased windowing ALONE also scored 2/8 on long noisy pairs, so the
fit does not subsume isolation. Post-isolation: 8/8 on short and long noisy
pairs, clean-vs-noisy pair consistency 8/8 + 8/8; `'[code]'`/`'[elided]'`
placeholder tokens measured worse than pure removal (7/8, 6/8 on long pairs)
and were rejected; a 20%-ratio floor guard measured harmful (3/8 — it reverts
exactly the short noisy inputs isolation exists to fix) in favor of an
absolute 24-char floor that only catches near-all-code inputs. Offline
regression: a noise tripwire (same instruction bare vs wrapped must classify
identically) fails against pre-isolation code and passes after. Still not
built: attention-masking de-weighting as an alternative to stripping (it
would preserve the noise tokens' presence without letting them dominate);
unclosed `<system-reminder>` tags and un-fenced diff hunks are left in place
(only closed-tag spans, fenced blocks, and marker-prefixed lines are
stripped); and a code-grounded instruction whose pasted snippet is the
subject can still land on a near-category — measured 3/4 on a 4-task
grounded set, the residual miss being description similarity
("refactor this helper" + code → `debugging`), not noise dominance.
## Gaming mode, and the backoff that used to do nothing
**The classifier circuit breaker did not break the circuit.**
`_last_classifier_failure` was written by `_record_failure()` and read by
nothing — `_classify_cascade` gated only its *cloud* step, and on a different
timestamp. So the router re-dialled a known-dead local classifier on every
request. A stopped Ollama refuses immediately and costs little; a **hung** one,
or a VPN-bound one that black-holes, costs the full 120s `timeout_seconds` per
request for as long as the outage lasts. `_classifier_backoff_active()` is now
the read, consulted **before** the client is constructed.
One detail there is load-bearing: recording the failure lives in `classify()`'s
exception handlers, NOT in `_classify_cascade`. The cascade is walked for
reasons other than a fresh failure, and if those re-stamped the clock, every
request during an outage would push the deadline forward and the local
classifier would never be re-probed while traffic flowed — a permanent outage
wearing a circuit breaker's clothes.
**`local_compute.enabled` (default true) is the outer gate over local
hardware.** Turn it off when you stop Ollama for a game and the router *skips*
every local call rather than discovering the outage one timeout at a time:
classifier, `/health` probe, local verification, local-vision fallback, and
local dispatch rows (dropped in `load_candidates`, so a local row is never
picked and then 503'd). `/v1/models` stops listing local rows, and an explicit
pin gets a 503 naming the flag instead of NeuralWatt's unknown-model 400.
ONE flag the code reads, **not** a macro writing five keys — a macro is hard to
undo cleanly, drifts the moment a sixth call site appears, and leaves nobody
able to answer "why isn't the classifier running?" from one place.
`verification.local_llm_enabled`, `local_vision.enabled` and
`local_energy.enabled` keep their own meanings; this ANDs over them.
**It REFUSES to engage without `classifier.cloud_fallback`** — 409 on the
runtime knob, a validation error at config load. Skipping the local classifier
does not make classification remote; without a cloud classifier it stops
classifying, and every request falls through to a static guess recorded as
`general_chat`, a fully scored category indistinguishable from a real
classification afterwards. A refusal, not a warning, because a warning is what
nobody reads while their game is loading. Nothing auto-writes the block.
Cascade steps 1 and 2 still run **ahead** of the cloud call: a stale session
classification is free and was a real classification of that same session, so
paying to re-derive an answer already held is spending money for nothing.
**A latent substring bug fell out of the profiles work.** SQLite stores
`eligible_categories` as a comma-joined string, and
`task_category not in "<a>,<b>"` is a SUBSTRING test — so a row eligible only
for `file_summarization` also admitted `summarization`. Latent on main (the
category-less probe short-circuits before the compare) and live the moment
anything probes per category. `routing.parse_eligible_categories` is now the
single parser `dispatcher.load_candidates` and `admin.py`'s probe both use.
## Known open questions
- Answered: cost and eco stay separate axes — grid intensity spans 13.6x
across the catalog, so they rank models differently.
- Answered: the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh were
`carbon_source: static_fallback` — a substituted constant, not a
measurement. They are now excluded from eco rather than trusted. Still
worth asking NeuralWatt why the fallback keeps the original `grid_id`,
since that is what made it look like a real regional difference.
- Three models still fail a split-half stability check at 7 samples. Is the
instability real (variable serving conditions) or an artifact of when the
sweep ran? Re-sweeping at a different hour would tell.
- Answered, and the question no longer parses: tier-1 composites used to sit
within 0.009 of each other because min-max normalization compressed them.
There is no composite any more — ranking is quality first, cost as the
tiebreak inside `quality_tolerance` — so nothing normalizes and nothing
compresses.
- Answered: the eval set exists (`evals/tasks.yaml`, 43 tasks, four scoring
kinds) and `tests/test_task_set.py` keeps it honest. The benchmark-sourced
rows now split the coding categories, but two tasks are still flat at 1.00
(`debug_affine_coprime` and most of the BFCL set) and either need hardening
again or should be conceded as non-discriminating. **Try samples
before hardening.** `docs_writing` looked flat at the top too, and six more
passes spread it 0.66-0.97 without touching a task; two samples per model is
not enough to tell a saturated task from an unsampled one.
- How much context-assembly (RAG-style retrieval) belongs in the classifier
step vs. a separate pre-step? Leaning decoupled, undecided.
- Should `eco_score` use real-time grid carbon intensity per request or a
stable per-model average? Currently the latter, from the reference sweep.
`grid_carbon_intensity` and `grid_id` are logged per observation, so this
stays answerable from data without a re-run.
## Config is strict: an unknown key is an error
Pydantic ignores extra keys by default, which means a typo or a misplaced
setting loads cleanly, does nothing, and still looks configured. Every config
model now inherits `StrictModel` (`extra="forbid"`), so both of these fail at
load rather than silently:
```
verification.max_input_chars # right key, wrong section
routing.min_tool_proficency # sic
```
This is not hypothetical. `max_input_chars` shipped into the `verification:`
block instead of `classifier:` and was accepted and discarded — it happened to
match the code default, so behaviour was correct and the file was a lie.
Editing it would have done nothing.
The corollary worth keeping: **every knob belongs in `config.yaml`, not only
in a Pydantic default.** A default the file never mentions is invisible to
anyone tuning it. `classifier.outcome_attribution_window_seconds` was removed
in the same pass — it was declared, never read, and shadowed the
`verification` one that actually is.
## Setup
Full install steps (venv, deps, config, first run) in [README ## Installation](README.md#installation). Host-local deployment values go in `config/config.local.yaml` (gitignored overlay) — `classifier.model`/`base_url`, `local_energy.*`, host-specific URLs. General defaults in `config/config.yaml` stay shareable. `objective.plan_kwh_per_period` in the README config table. Model tags (`num_ctx`) + `verification.model` same-tag note in [docs/local-models.md](docs/local-models.md). Requirements are pinned — bump deliberately (README).
## Run as a service
`deploy/` holds the dispatcher's systemd **user** unit plus a timer and a
oneshot service each for the poller, the seed sweep, the backup, the offsite
sync and the feedback fold, and one drop-in for a *system* Ollama — see
`deploy/README.md` for install and operation. Every timer there is enabled on
install except `llm-router-feedback.timer`, which is not, on purpose. In short:
```bash
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer
```
The dispatcher binds `127.0.0.1:8080`. **The poller timer is load-bearing, not
housekeeping** — but not for the reason this section used to give, and the
correction matters because it inverts which failure to watch for.
`mark_stale` runs only *inside* `poller.main()`, and `main()` returns early on
a `RequestException` — **before** `upsert` and **before** `mark_stale`. So a
stopped timer or a provider outage marks nothing: the catalog freezes at
last-known-good and the router keeps routing on prices that may be weeks old.
The failure is **silent and open**, not loud and closed. Nothing surfaces it,
because a frozen row still reads `availability = 'active'`.
The path that *can* empty the candidate set is narrower and is not governed by
the timer at all. `fetch_neuralwatt` reads `payload.get("data", [])` with no
floor on row count, so a 200 response carrying an empty or truncated `data`
array — a partial provider outage, a schema change, an auth path degrading to
an empty list — clears `raise_for_status()`, upserts nothing, and then lets
`mark_stale` run anyway. Three days of that and every row is stale and
`exclude_stale: true` leaves zero candidates for everything.
`stale_after_days: 3` only sets the length of that fuse; it does not arm or
disarm it. Against a 2-hourly poll it is 36 successful polls of margin, and
recovery is automatic — `upsert` writes `availability = excluded.availability`,
so one good poll flips every stale row back to active. The fix is a sanity
floor on the fetch, not a larger number.
**A second path empties the candidate set, and it bit on 2026-09-01.** Admin
availability overrides are not governed by the poller at all. Deprecating the
seven expensive models through `/admin` collapsed tier 3's context ceiling from
782,324 to **94,196** — while tiers 1 and 2 stayed at 782,324 — so every tier-3
request above 94k returned 422 with nothing warning anywhere. It surfaced ~19
hours later as an agent failing mid-task on an opaque error.
**The obvious check for this is wrong, and the reason is worth remembering.**
Tempting: warn when a higher tier's context ceiling sits below a lower tier's.
But `ceiling(T)` is the max `effective_context_window` over models with
`tier >= T`, and tier is a capability *floor*, so the eligible set shrinks
monotonically as T rises — `ceiling(1) >= ceiling(2) >= ceiling(3)` is a
theorem, true of every catalog. Such a warning fires always and means nothing.
What actually failed is that a tier's ceiling dropped below what that tier is
*asked* to serve, which is only knowable from traffic: compare `ceiling(T)`
against the observed `required_context_tokens` for decisions classified at tier
T. That is silent on all three tiers today and fires on the outage state
(94,196 vs an observed max of 268,168). See
`plans/catalog-staleness-and-poller-failure-modes.md` §4.4.
**This recurred on 2026-09-04 through a dimension the detector did not model,
and both halves of the fix are now in `metrics.py`.** Admin deprecations took
out `kimi-k3*` — the only vision-capable rows with enough context — so a
242,486-token image request 422'd while every existing check stayed silent,
because the *all-models* tier-1 ceiling was still 782,324. The vision-capable
ceiling had collapsed to 192,500.
- **Predictive:** `capability_ceilings` / `capability_demand_warnings` compute
`vision` and `json_mode` sub-ceilings and compare each against demand
actually observed for requests carrying images / requesting JSON. Two extra
series, not a bucket per capability combination.
- **Reactive:** `rejection_warnings` watches `route_decisions` for rows with
`selected_model IS NULL`. This is the more valuable half and the simpler
one — it catches the *next* dimension nobody predicted, at the cost of
firing after the first failure rather than before.
Two details in the reactive detector are load-bearing and easy to undo by
accident. It groups by `(task_tier, digit-normalized reason)` using the
**structured column**, because normalizing digits alone merges `tier >= 1`
and `tier >= 3` rejections into one group and hides whether the broadest or
the frontier candidate set went empty. And the signal is **novelty OR rate**,
never mere presence: measured on the live DB, routine rejections run ~3/hr
while the 2026-09-04 incident was only n=2 — *below* the noise floor — so no
single count threshold can both catch it and stay quiet. A group absent from
the 24h baseline warns at n≥2; a familiar group warns at the configured
count. Zero rejections warn about nothing: a genuinely impossible request
SHOULD 422.
The service holds a billable API key and has **no auth of its own**. Loopback
bind is the only thing standing between the open internet and your allowance;
add auth before widening `--host`.
The same applies to an Ollama shared over a VPN — it has no auth either, so
`deploy/ollama-over-vpn.conf` binds it to the VPN address rather than
`0.0.0.0`, which would publish it on whatever network the client happens to
be on.
### When the router goes unreachable, start at docs/incidents.md
Eight incidents so far, nearly all sharing one shape: a change that looked local
to the router silently degraded the agent depending on it, and none announced
itself as a router problem. **`docs/incidents.md` carries the full write-ups plus
a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
#8 is the exception worth knowing before an unattended agent run: the router
worked perfectly while agent workers looped for hours with no progress, and no
check noticed, because every check watched spend or availability rather than
whether changes landed (`plans/no-progress-detection.md`).
Two conventions from those incidents that bind every session, and so stay here:
- **8080 is production, always.** It is baked into `opencode.json`, the systemd
unit, every curl example here, and the admin frontend's own fetches. A
throwaway instance (manual iteration, Playwright smoke tests, anything that is
not "use the real router") binds **8081**. Never send a kill signal to a
process matched by name or port rather than by a PID you started yourself --
`Restart=always` will fight you, and on this repo it may be your own model
access.
- **Never point `config/config.yaml` at test fixtures.** It is the file the live
service reads. Pass a different config file, monkeypatch `cfg.database.path`
in-process, or use a temp copy.
Recovery for an unreachable-but-`active` service is
`systemctl --user restart llm-router.service` -- a hung process was never in a
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`),
separate from the dispatcher, so a stuck router cannot silence the thing meant
to notice it is stuck. Install and enable it with `systemctl --user enable
--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a
placeholder home path and must be `sed`-repointed to the real one first), or
run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md)
for the signals it fires on, the alert lifecycle, and the known limits.
## Pointing a coding agent at it
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
opencode, an SDK, plain curl. Repo-local `opencode.json` is already wired up,
so running `opencode` from a clone of this repo routes by default. For global
use, merge `provider.llm-router` into `~/.config/opencode/opencode.json`.
| model name | behavior |
|---|---|
| `auto` | router picks; flex rows excluded so nothing is held during peak |
| `auto:batch` | router picks; flex rows admitted, for overnight/async work |
| any real model id | dispatched as asked, still logged |
Streaming is proxied chunk by chunk rather than buffered, so tokens still
render as they arrive. NeuralWatt emits its energy and cost blocks as SSE
**comment** lines (`: energy {...}`) before `data: [DONE]` — ordinary clients
ignore comments, so the stream passes through untouched while the router
reads the telemetry on the way past. Without that, streamed calls would log
no energy at all, which is most of the point of this project.
## Try it
Copy-pasteable `curl` examples (`/route`, `/dispatch`, `/v1` models + chat) and
the classifier-skip overrides (`task_category`, `task_tier`,
`required_context_tokens`) live in [README ## Usage](README.md#usage). Inspect
what a dispatch cost/burned with the `sqlite3` query in
[docs/operations.md](docs/operations.md).