Files
6krrt/README.md
adlee-was-taken 1a6354deea refactor(session-cache): rename staleness_minutes to staleness_seconds
Change the unit of session_cache.staleness from minutes to seconds so it
can express finer-grained (sub-minute) staleness windows. This is a
straight rename, not an additive/compat knob — no deprecated alias, per
the project's convention of updating every consumer in the same change.

New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was
120 minutes). Default: 1200 seconds (was 20 minutes). The validator's
reasoning is unit-independent and carries over: the floor is deliberately
> 0 because 0 would make session_cache.get() miss every turn while
put() still writes and the classifier-failure cascade's stale_read
ignores staleness; the ceiling reasoning (unbounded window = never-expiring
cache, 7200s still >> 840s real max run) also carries over in seconds.

Every consumer updated in the same commit:
- src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX
= 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated
- src/dispatcher.py: drop the * 60 conversion (field is native seconds)
- src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST,
  _CONFIG_GET_ORDER, _runtime_state, error message template
- admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS
- config/config.yaml: staleness_seconds: 1200
- tests: test_admin_runtime/config/frontend/knob_coverage, plus stale
  comment in test_chat_completions
- docs: admin-portal.md, evaluation.md, README.md

config.local.yaml is gitignored and will be migrated separately.
2026-09-18 00:45:08 -04:00

484 lines
27 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<h1 align="center">
<img src="assets/6krrt-logo.svg" width="120" alt="6krrt logo"><br>
6krrt — Local LLM Model Router
</h1>
A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs.
It classifies each incoming task with a local Ollama model — category, tier,
required context — and dispatches to the highest-expected-success open-weight
model on an OpenAI-compatible provider, under a per-request cost ceiling and
tiebroken by price. Point `model: "auto"` at `POST /v1/chat/completions` for
automatic selection; pin a real model id and it dispatches as asked, still
logged.
Any OpenAI-compatible carrier should work. **NeuralWatt is the reference
provider** — it powers every measurement in this README — and uniquely exposes
per-request **power/energy telemetry** that the router ingests for cost
accounting; other carriers that report cost only in their usage response are
supported but carry no energy telemetry. Configure additional providers and
their API keys under `dispatch_providers:` in `config/config.yaml`.
**Numbers in this README are measurements, not specifications.** They come from
one deployment against one provider account, and the catalog, prices, grid
intensity and pool load all move. They are here because the reasoning behind a
design choice is worth more than the choice, and re-running the measurement is
how you check whether it still holds for you.
## Features
- **Check routing before spending anything.** `POST /route` classifies, ranks,
and returns the selected model with no provider call and no cost.
- **Price per request from the actual shape of the traffic.** Cost is estimated
from catalog token prices scaled to the prompt size, an assumed completion
length, and an assumed cache rate — not a single fixed benchmark.
- **Dispatch selected tasks to a local model.** When the classifier puts a task
in `file_summarization` or `diff_checking`, the router will consider the local
model (`qwen2.5-coder-router:14b`) only when it scores within the quality
tolerance of the cloud leader. (Dormant under the default profile — local
dispatch does not routinely win.)
- **Fall back to local vision when no cloud row supports images.** If no
vision-capable catalog candidate survives the hard filters, the router proxies
the request to a local Ollama vision model instead of returning 422.
- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard
that follows `/events/decisions` as decisions are recorded, with no polling
delay.
- **Verify before learning.** Every routed response is structurally parsed in
the background; larger prose answers get an async local-LLM spot-check, and
failures fold back into per-model proficiency through `feedback.py`.
## Documentation
- [Routing internals](docs/routing.md) — quality-first ranking, local vision fallback, candidate ranking
- [Data model](docs/data-model.md) — SQLite schema: models, proficiency, energy, verifications, route_decisions
- [Response verification](docs/verification.md) — structural + local-LLM checks, feedback loop
- [API reference](docs/api.md) — endpoints, virtual models, streaming, logging headers
- [Client setup](docs/clients.md) — pointing opencode / OpenAI-compatible clients at the router
- [Architecture](docs/architecture.md) — module-by-module map
- [Operations](docs/operations.md) — logfmt, journalctl, tracing requests
- [Incidents](docs/incidents.md) — how the router has broken, with a symptom → one-line-check table
- [Evaluation & classifier](docs/evaluation.md) — self-eval harness, classifier reliability notes
- [Admin portal](docs/admin-portal.md) — browser dashboard, controls, decision log
- [Local model sizing](docs/local-models.md) — fitting Ollama models into 24GB VRAM
- [Context pruning (pinch)](docs/pinch.md) — pre-dispatch conversation trimming once a session grows past a token budget
- [Local config overlay](docs/config-local-overlay.md) — per-machine `config.local.yaml` overrides (gitignored)
## Requirements
- API key(s) for the provider(s) you route to — a NeuralWatt key by default,
plus any others under `dispatch_providers:` (e.g. an OpenRouter key). The
reference deployment uses NeuralWatt, whose energy telemetry feeds cost
accounting.
- Python 3.10+ (the test suite is verified on 3.10 and 3.14).
- An Ollama reachable from wherever this runs, with a classifier model pulled.
- For local dispatch: `qwen2.5-coder-router:14b` created from
`qwen2.5-coder:14b` with `num_ctx 32768` (see [docs/local-models.md](docs/local-models.md)).
Nothing else is assumed about the host — routing itself is SQLite and arithmetic.
## Installation
### Local (`venv`)
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
cp config/config.local.yaml.example config/config.local.yaml # optional: deployment-specific overlay (gitignored)
PYTHONPATH=src python -m poller # populate the catalog
PYTHONPATH=src python -m tier # resolve tiers
PYTHONPATH=src python -m config # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
Then edit `config/config.yaml` for your own setup — at minimum:
| Key | Why |
|---|---|
| `classifier.model` | must match a model `ollama list` reports |
| `classifier.base_url` | where that Ollama actually is |
| `objective.plan_kwh_per_period` | your plan's quota; `/health` reports burn against it |
| `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
| `session_cache.enabled` | off by default; caches category/tier per session for `staleness_seconds` to skip repeat classifier round-trips on long agent sessions |
`PYTHONPATH=src python -m seed_energy` is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path.
### As a systemd service
Five user services (with companion timers) cover continuous dispatch, catalog
polling, periodic energy reseeding, and database backups/offsite snapshots. See
`deploy/README.md` for full instructions.
**The poller timer is load-bearing, not optional** — but it fails *silently*,
not loudly. `mark_stale` runs only inside a poll run that got past the fetch, so
a stopped timer or a provider outage marks nothing: the catalog freezes at
last-known-good and the router keeps routing on prices that may be weeks old,
with every row still reading `active`. Watch for staleness; don't expect it to
announce itself. (`freshness.stale_after_days` is 3 with `exclude_stale: true`,
which against a 2-hourly poll is 36 polls of margin.)
### Where Ollama lives
```bash
ollama pull mistral-nemo:12b # or whatever you set as classifier.model
```
It does not have to be on the machine running the router; the box with the
GPU usually isn't the laptop. To use one across a VPN, point **both**
endpoints at it:
```yaml
classifier:
base_url: "http://<vpn-ip>:11434/v1"
verification:
base_url: "http://<vpn-ip>:11434" # same host, so `model` can stay null
```
and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds
`127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address
rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the
port can run inference and enumerate your models.
Both endpoints move together because the verifier speaks Ollama's *native*
API and cannot follow the classifier to a cloud provider. Config load refuses
the case where they are on different hosts and `verification.model` is null,
because that combination fails silently.
## Usage
### Route without spending anything
Use `POST /route` when you want to see what the router would pick for a task before paying for a provider call.
```bash
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view into service objects."}'
```
- Runs the local classifier to determine category, tier, and required context.
- Applies hard filters and ranks the surviving candidates.
- Returns the selected model, estimated cost, and estimated proficiency.
- **No provider call is made and no quota is consumed.**
### Skip the classifier when you already know the shape
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`, and `required_context_tokens` overrides.
```bash
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}'
```
- The classifier is not called, so latency is just the routing pass.
- `route_decisions.classification_source` is logged as `override`.
### Actually dispatch and log energy
`POST /dispatch` does the same routing work as `/route` and then calls the selected provider, streams the response if requested, and records the request as an `energy_observations` row.
```bash
curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
-d '{"task":"What is a Python context manager?"}'
```
- Provider response is proxied back, including streaming chunks.
- Energy, carbon, cost, and duration are scraped from SSE comments and logged.
### Report whether it actually worked
Every other quality signal is a proxy — structural checks know code *parses*,
the local LLM check guesses prose *looks* right. Only the client knows whether
the answer did the job, because it ran the tests.
```bash
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
```
- The only signal that survives streaming — a retry can't reach bytes already
sent, but a report arrives afterward and works either way.
- Successes count too: unlike the automatic checks, `feedback.py` folds a
reported `ok: true` in as a real success, not just failures.
- An unknown `request_id` returns `404` rather than being silently accepted.
### Point any OpenAI-compatible client at it
The `/v1` endpoints speak the OpenAI completions and models API.
```bash
curl -s localhost:8080/v1/models
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'
```
Ask for a **virtual router model** and the router picks a candidate, subject
to the profile you specify. The general form is `auto:<profile>`; asking for
just `auto` is shorthand for `auto:default`.
```bash
# Normal interactive routing (same as "auto")
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":...}'
# Restrict to local dispatch models (ollama-local provider only)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:locality","messages":...}'
# Restrict to models priced at ≤ $0.50/1M completion tokens
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:onlycheaps","messages":...}'
# Restrict to frontier tier (tier 3)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:bigboybritches","messages":...}'
# Overnight/async work via flex rows
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:batch","messages":...}'
```
| Profile | What it does |
|---|---|
| `auto` / `auto:default` | Normal quality-first routing; `-flex` rows excluded |
| `auto:batch` | Admits `-flex` rows that may be held during peak hours |
| `auto:locality` | Routes only to the `ollama-local` provider (local dispatch) |
| `auto:onlycheaps` | Limits to models at or below $0.50 per 1M completion tokens |
| `auto:bigboybritches` | Routes only to tier-3 (frontier) models |
Profiles are **candidate-set filters** — they narrow which models the router
may pick from, but do not change the ranking objective (quality-first, cost as
tiebreak). An unknown profile raises HTTP 422.
Operators can define custom profiles under `profiles:` in
`config/config.yaml`. Each entry accepts `provider`, `min_tier`/`max_tier`,
`max_cost_per_1m_completion`, `latency_tolerance`, and `allowed_model_ids`
as rewrite rules on top of the default profile.
- Ask for any real model id and the router dispatches directly, still logged.
- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines.
### Ask an image question
Cloud vision is not universal in the catalog. Requests carrying `image_url` parts
are routed only to vision-capable catalog rows. If none survives, the router can
fall back to a local vision model instead of returning 422.
```bash
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}'
```
- Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined.
- Configure the fallback in `config/config.yaml` under `local_vision:`.
### Force JSON output
Use `response_format` when you need structured output. The router treats this as
a hard capability requirement and only admits rows that declare `supports_json_mode = 1`.
```bash
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}'
```
- `NULL` fails closed: an unknown flag means the capability cannot be confirmed.
- When streaming is not used, the structural verification verdict appears in the `X-Router-Verification` header.
### Watch it live
Three read-only ways to observe the router without spending quota:
```bash
# Aggregate JSON health/usage summary
curl -s localhost:8080/metrics | python -m json.tool
# Live SSE stream of routing decisions
curl -s localhost:8080/events/decisions
# Terminal dashboard with live routing feed
PYTHONPATH=src python -m tui
```
- **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency.
- **`GET /events/decisions`** is a Server-Sent Events stream; replays recent decisions, then streams new ones.
- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels.
### Compare against trivial baselines
`baseline_report.py` replays recent `route_decisions` against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend.
```bash
PYTHONPATH=src python baseline_report.py --since 2026-08-01
PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv
```
A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice.
### Admin web portal
For the full admin portal walkthrough, see [docs/admin-portal.md](docs/admin-portal.md).
### Probe routing without spending
`router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree.
```bash
PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
```
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
- Use `--category`, `--tier`, and `--context` to override the classifier deterministically.
## At a Glance
| Dimension | Detail |
|---|---|
| **Cost model** | Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog. |
| **Latency** | Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response. |
| **Quality** | Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior. |
| **Verification** | Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via `feedback.py`. |
| **Fault tolerance** | Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent `timeout_seconds` silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate. |
| **API surface** | OpenAI-compatible `/v1` endpoints, streaming chunk proxy with SSE telemetry scraping. |
## Architecture
```
┌─────────────────────┐
incoming task ───▶│ Local Classifier │ Ollama (mistral-nemo:12b)
│ - task_category │ ~4s warm, ~120s cold cap
│ - task_tier │ temperature: 0, max 1024 tokens
│ - required_context │ max_retries: 0 (silent 3× cap guard)
│ - confidence │
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Classifier Fallback │ tier 2 / general_chat on failure
│ (graceful degrade) │ — not a 502/503
└──────────┬───────────┘
▼
┌─────────────────┐
│ Escalation │ low-confidence tier bump
│ (optional) │ threshold: 0.6 confidence
└────────┬─────────┘
▼
┌─────────────────────┐
│ Hard Filters │ context window ≥ required
│ (routing.py) │ tier ≥ required
│ │ freshness (active, not stale)
│ │ access_level allowed
│ │ latency_class compatible
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Quality-first select │ max proficiency, cheapest
│ (routing.py) │ among equals, under a
│ │ per-request kWh ceiling
└──────────┬───────────┘
▼
┌──────────────────────┐
│ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al.
│ (FastAPI API) │──▶ OpenAI-compatible /v1
│ │ Streaming SSE + energy scrape
└───────┬──────────────┘
▼
┌─────────────────────────┐
│ Verification Pipeline │
│ Layer 1: Structural │ ast.parse, json.load, yaml.safe_load
│ Layer 2: Local LLM │ async, ~6s, >600 token gate
└──────┬──────────────────┘
│
▼
┌─────────────────────┐
│ feedback.py │ failures → proficiency updates
│ (on-demand agent) │ idempotent, attributed-only
└─────────────────────┘
```
Not pictured: a model that 5xxs is passively excluded from Hard Filters by
`circuit_breaker.py` until its cooldown clears, and a verification failure can
loop back into Dispatcher for a tier-budgeted retry before falling through to
`feedback.py` — see [docs/verification.md](docs/verification.md#retry-budget--iterationpy).
For the detailed module-by-module map, see [docs/architecture.md](docs/architecture.md).
## Tech Stack
| Layer | Technology |
|---|---|
| **Language** | Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed) |
| **Framework** | FastAPI + uvicorn |
| **Database** | SQLite (`router.db`) — decision table, energy observations, proficiency, verifications |
| **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN |
| **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works |
| **Cloud Providers** | OpenAI-compatible carriers via `dispatch_providers:` — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works |
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
| **Testing** | `pytest` — 1923 tests across ~90 files, all offline |
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
Dependencies are pinned in `requirements.txt` — recreate the venv with those exact versions to avoid silent drift. **Bump deliberately**: the old `>=` ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use `>=`.
## Testing
```bash
python -m pytest # 1923 tests
python -m pytest --cov # with coverage
```
No test calls a provider or a local model — the pure modules take rows and
config as arguments, so the suite runs offline on a clean checkout.
`pyproject.toml` puts the repo root on `sys.path` for pytest — the modules live
at the root rather than in a package, so `pytest` (console script) and
`python -m pytest` would otherwise disagree about whether `import config`
resolves.
Representative test files (~90 files total; the list below is a useful subset, not exhaustive):
| Test file | What it covers |
|---|---|
| `test_scoring.py` | `normalize_inverted`, `composite_score`, None-handling |
| `test_routing.py` | Hard filters, candidate selection & ranking across all 9 categories × 3 tiers |
| `test_tiering.py` | Tier resolver precedence: override → reasoning → cost + context window → mid |
| `test_apply_tiering.py` | DB tiering pass, sanity guards |
| `test_poller_parsing.py` | Serving class, base model, access level parsing |
| `test_proficiency.py` | Blending rule, accumulation |
| `test_load_candidates.py` | `load_candidates`: cost/eco/proficiency join |
| `test_eval_scoring.py` | Code, exact, tool, and judge scoring |
| `test_task_set.py` | Reference solutions validating every `code` task's checks & `exact` answers |
| `test_verification.py` | Structural checks on every language, verifier edge cases, local-LLM gating |
| `test_feedback.py` | Failure identification, idempotency, attribution filtering |
| `test_iteration.py` | Retry budget per tier, and matching the retry to the failure kind |
| `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal |
| `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split |
| `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior |
| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
| `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
| `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
| `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy |
| `test_classifier_input.py` | Classifier framing: `_previous_context` scope, `_classifier_user_content` framing |
| `test_route_decisions.py` | `route_decisions` table, inline-create helper, config gate |
| `test_admin_*.py` | Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family) |
| `test_local_dispatch*.py` | Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep |
| `test_multi_provider*.py` | Multi-provider dispatch + poller; per-provider balance & attenuation |
| `test_circuit_breaker.py` | Passive circuit breaker: cooldown + backoff on model 5xx |
| `test_session_cache.py` | Per-session category/tier caching to skip repeat classifier round-trips |
| `test_outcome_attribution.py` | Outcome attribution & ambiguity refusal (streaming answer path) |
## Known Limitations & Open Items
- **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. See [CLAUDE.md](CLAUDE.md) "What's NOT built yet" for the live list.
- **Three models unsettled** — split-half stability varies too much for some model positions.
- **Retry does not reach streaming** — corrective attempts work only on the non-streaming path; `POST /outcome` is the streamed answer.
- **Local dispatch has its own limits** — no true streaming (answer is buffered), no `verifications` rows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see [docs/routing.md](docs/routing.md)).
- **No auth** — the service holds a billable API key with no authentication of its own; loopback is the only guard.
- **Local energy is metered when enabled** — `local_energy:` in `config/config.yaml` gates classifier/verifier/vision draw on a separate `local_energy_observations` table (off by default, refuses `enabled: true` without a tariff rate). See [data-model.md](docs/data-model.md).
- **Context assembly (RAG)** is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.