diff --git a/CLAUDE.md b/CLAUDE.md index 51cd028..2aa5d66 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -20,9 +20,12 @@ the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it. -Neuralwatt is the only provider. OpenRouter was removed — the `provider` -column and the `(model_id, provider)` primary key stay so a second provider -can be added later without a migration. +Neuralwatt remains the primary provider. OpenRouter is back as an opt-in +provider gated by the `provider_model_allowlist` table — a provider with +`require_allowlist: true` is ignored unless its model_id is explicitly +allowlisted, and any previously-upserted row that drops off the list is +marked `deprecated` on the next poll. The `provider` column and the +`(model_id, provider)` primary key are unchanged, so this adds no migration. ## Stack @@ -207,7 +210,7 @@ rather than from months of history. ## What's built and working - `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md). -- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). +- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below. - `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md). - `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md). - `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md). @@ -216,9 +219,9 @@ rather than from months of history. - `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates. - `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below). - `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md). +- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded. - `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md). - `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md). -- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md). - `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md). - `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable. - `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md). @@ -227,10 +230,50 @@ rather than from months of history. - `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted. - `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted. - `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md). -- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). +- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor. - `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md). +- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below. +- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted. - `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md). +## #45 — OpenRouter is an opt-in allowlist provider + +OpenRouter used to be ingested whole, then removed, and is now back — but only +as an opt-in provider. The base config sets `openrouter.require_allowlist: true` +and ships a short seed allowlist. Models on that seed list are upserted and +kept active; anything else in the OpenRouter catalog is filtered out before +upsert and any previously-active OpenRouter row that is not on the list is +marked `deprecated` on the next poll. + +This is deliberately different from the old ingest-everything behavior. The +previous approach once pulled in a non-chat model (`lyria/...`) that returned +HTTP 404 on dispatch because the endpoint expected chat completions. The router +had paid for the classification, selected the model, and then failed on the +provider call. Allowlist-gating prevents that class of failure by default: if a +model id has not been reviewed and explicitly added, the router acts as if it +does not exist. + +The seed allowlist is short and has firm exclusions. It does NOT include: + +- `x-ai/*` (Grok) +- `openai/*` +- `anthropic/*` + +Those exclusions are non-negotiable. They are not "currently excluded" or +planned for future inclusion; they are deliberately absent from the seed list. +Adding one requires editing both the seed allowlist and this file. + +The admin portal exposes the allowlist under `/admin`: the providers page shows +which providers require one, and each provider row links to an allowlist editor +where entries can be added or removed. The underlying four API endpoints are +`GET /admin/api/providers/{name}/allowlist`, +`POST /admin/api/providers/{name}/allowlist`, +`DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers +page itself surfaces `require_allowlist` with a link to the allowlist editor. + +Neuralwatt remains the primary, ungated provider. The allowlist behavior only +fires for providers with `require_allowlist: true`. + ## Proficiency: category now changes routing `proficiency_score` is the ONLY category-dependent term in the ranking, so @@ -695,6 +738,42 @@ weeks. ### Which implementation is PRIMARY is now a config choice +`classifier.mode` in the live deployment is currently `local_encoder`, set in +`config.local.yaml` with `device: cuda` on the classifier host. That is the +mode answering real traffic right now. The explicit caveat is that a +CPU-vs-CUDA latency and confidence comparison on the live classifier host is +still pending; until that measurement exists, flipping the project default to +`local_encoder` is not decided. + +**Turning that mode on for real found three real bugs in one afternoon +(2026-09-06), each a fresh instance of this project's own recurring lesson — +verify against the live system, not the plan.** First, the config-load +validator for `confidence_threshold` didn't exist yet: the admin UI saved a +raw `80` (meant as 80%) straight into the overlay with no conversion, which +would have made every real confidence score read as below-threshold on the +next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability +exceeds 1.0). Caught before the restart, not after — see `confidence_threshold` +in [local-models](docs/local-models.md) for the full incident and the fix. +Second, once that was corrected and the service actually restarted, it +crash-looped twice more before coming up clean: the shipped default model +(`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace +sometime after this project picked it (401 on an unauthenticated GET of its +own model page), and separately `HF_HOME`'s default cache path falls outside +this service's `ProtectHome=read-only` sandbox exception — both are now fixed +(switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo). +Third, and most consequential: the very first real classifications measured +only 5 of 9 test categories correct, because `classify_zero_shot` was passing +raw config identifiers like `tool_use_agentic` and `diff_checking` directly as +zero-shot candidate labels — HF's pipeline scores a label against a hypothesis +template ("This example is {}."), and an underscored code token is not a +sentence the model's NLI training ever saw. Mapping each category to a natural- +language description before scoring, plus `multi_label=True` (the pipeline's +single-label default forces every candidate to compete for the same +probability mass), brought that to 8 of 9 correct with confidence scores +0.77-0.999 on the hits — the one remaining miss scored below the configured +threshold and correctly fell through to the safe fallback rather than +mis-routing. + `classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks what answers a classification request — a peer concept to the cascade above, **not a replacement for it**. Whichever mode is primary, a failure still @@ -732,6 +811,11 @@ through the admin portal's Classifier card, which reports the **live** resolved primary for `cloud_primary_auto` rather than echoing the config value (see [admin-portal](docs/admin-portal.md)). +The default in `config.yaml` is still `local_llm`. `local_encoder` is +intentionally an opt-in per-deployment choice rather than the repository +default until the pending CPU-vs-CUDA comparison on the live classifier host is +available. + ## Local dispatch model A second local model can now be dispatched directly for specific categories. diff --git a/docs/admin-portal.md b/docs/admin-portal.md index 1ac56f1..def0557 100644 --- a/docs/admin-portal.md +++ b/docs/admin-portal.md @@ -9,13 +9,20 @@ endpoints it is reachable only from `127.0.0.1`. The portal is a **six-page, glass dark-mode** dashboard built on Tabler/Bootstrap with a few lighter inline SVG icons (see `plans/admin-design-standards.md` for the visual language and `plans/admin-work-framework.md` for the method -used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for -its future tables and `admin/frontend/` for the browser UI. +used to evolve it). The six pages are Dashboard, Models, Profiles, Decisions, +Controls, and Providers. It is implemented in `admin.py`, with +`config/admin_schema.sql` for its future tables and `admin/frontend/` for the +browser UI. **Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip (burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict mix and category breakdown, five history mini-charts, and a recent-decisions -table. +table. Clicking the chip opens a detail modal with per-provider balances and +runway; each line is labeled by its billing shape, so `"telemetry"` balances +read as "overage allowance" and `"polled"` balances read as "credits". Tiny +values keep their sign: a negative balance smaller than half a cent renders as +`~$0.00 (slight overage)`, a positive tiny balance renders as `<$0.01`, and an +exact zero balance renders as `$0.00`. **Pinch savings** — the dashboard card shows pinch's effect: share pruned (as a percent), tokens saved, median saved, and 30d dollars saved. When pinch is @@ -30,7 +37,9 @@ for the full model. **Models** — `GET /admin/models` is the model-availability table. Each row shows the serving class, tier, status, and an override dropdown (`active` / -`deprecated` / `stale`) that writes through to the routing hard filters. +`deprecated` / `stale`) that writes through to the routing hard filters. By +default the table hides deprecated and stale rows; enable the **Show deprecated +/ stale** toggle to include them.
@@ -127,6 +136,21 @@ can't safely represent: