
6krrt — Local LLM Model Router
A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs.
It classifies each incoming task with a local Ollama model — category, tier,
required context — and dispatches to the highest-expected-success open-weight
model on an OpenAI-compatible provider, under a per-request cost ceiling and
tiebroken by price. Point `model: "auto"` at `POST /v1/chat/completions` for
automatic selection; pin a real model id and it dispatches as asked, still
logged.
Any OpenAI-compatible carrier should work. **NeuralWatt is the reference
provider** — it powers every measurement in this README — and uniquely exposes
per-request **power/energy telemetry** that the router ingests for cost
accounting; other carriers that report cost only in their usage response are
supported but carry no energy telemetry. Configure additional providers and
their API keys under `dispatch_providers:` in `config/config.yaml`.
**Numbers in this README are measurements, not specifications.** They come from
one deployment against one provider account, and the catalog, prices, grid
intensity and pool load all move. They are here because the reasoning behind a
design choice is worth more than the choice, and re-running the measurement is
how you check whether it still holds for you.
## Features
- **Check routing before spending anything.** `POST /route` classifies, ranks,
and returns the selected model with no provider call and no cost.
- **Price per request from the actual shape of the traffic.** Cost is estimated
from catalog token prices scaled to the prompt size, an assumed completion
length, and an assumed cache rate — not a single fixed benchmark.
- **Dispatch selected tasks to a local model.** When the classifier puts a task
in `file_summarization` or `diff_checking`, the router will consider the local
model (`qwen2.5-coder-router:14b`) only when it scores within the quality
tolerance of the cloud leader. (Dormant under the default profile — local
dispatch does not routinely win.)
- **Fall back to local vision when no cloud row supports images.** If no
vision-capable catalog candidate survives the hard filters, the router proxies
the request to a local Ollama vision model instead of returning 422.
- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard
that follows `/events/decisions` as decisions are recorded, with no polling
delay.
- **Verify before learning.** Every routed response is structurally parsed in
the background; larger prose answers get an async local-LLM spot-check, and
failures fold back into per-model proficiency through `feedback.py`.
## Documentation
- [Routing internals](docs/routing.md) — quality-first ranking, local vision fallback, candidate ranking
- [Data model](docs/data-model.md) — SQLite schema: models, proficiency, energy, verifications, route_decisions
- [Response verification](docs/verification.md) — structural + local-LLM checks, feedback loop
- [API reference](docs/api.md) — endpoints, virtual models, streaming, logging headers
- [Client setup](docs/clients.md) — pointing opencode / OpenAI-compatible clients at the router
- [Architecture](docs/architecture.md) — module-by-module map
- [Operations](docs/operations.md) — logfmt, journalctl, tracing requests
- [Incidents](docs/incidents.md) — how the router has broken, with a symptom → one-line-check table
- [Evaluation & classifier](docs/evaluation.md) — self-eval harness, classifier reliability notes
- [Admin portal](docs/admin-portal.md) — browser dashboard, controls, decision log
- [Local model sizing](docs/local-models.md) — fitting Ollama models into 24GB VRAM
- [Context pruning (pinch)](docs/pinch.md) — pre-dispatch conversation trimming once a session grows past a token budget
- [Local config overlay](docs/config-local-overlay.md) — per-machine `config.local.yaml` overrides (gitignored)
## Requirements
- API key(s) for the provider(s) you route to — a NeuralWatt key by default,
plus any others under `dispatch_providers:` (e.g. an OpenRouter key). The
reference deployment uses NeuralWatt, whose energy telemetry feeds cost
accounting.
- Python 3.10+ (the test suite is verified on 3.10 and 3.14).
- An Ollama reachable from wherever this runs, with a classifier model pulled.
- For local dispatch: `qwen2.5-coder-router:14b` created from
`qwen2.5-coder:14b` with `num_ctx 32768` (see [docs/local-models.md](docs/local-models.md)).
Nothing else is assumed about the host — routing itself is SQLite and arithmetic.
## Installation
### Local (`venv`)
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
cp config/config.local.yaml.example config/config.local.yaml # optional: deployment-specific overlay (gitignored)
PYTHONPATH=src python -m poller # populate the catalog
PYTHONPATH=src python -m tier # resolve tiers
PYTHONPATH=src python -m config # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
Then edit `config/config.yaml` for your own setup — at minimum:
| Key | Why |
|---|---|
| `classifier.model` | must match a model `ollama list` reports |
| `classifier.base_url` | where that Ollama actually is |
| `objective.plan_kwh_per_period` | your plan's quota; `/health` reports burn against it |
| `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
| `session_cache.enabled` | off by default; caches category/tier per session for `staleness_seconds` to skip repeat classifier round-trips on long agent sessions |
`PYTHONPATH=src python -m seed_energy` is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path.
### As a systemd service
Five user services (with companion timers) cover continuous dispatch, catalog
polling, periodic energy reseeding, and database backups/offsite snapshots. See
`deploy/README.md` for full instructions.
**The poller timer is load-bearing, not optional** — but it fails *silently*,
not loudly. `mark_stale` runs only inside a poll run that got past the fetch, so
a stopped timer or a provider outage marks nothing: the catalog freezes at
last-known-good and the router keeps routing on prices that may be weeks old,
with every row still reading `active`. Watch for staleness; don't expect it to
announce itself. (`freshness.stale_after_days` is 3 with `exclude_stale: true`,
which against a 2-hourly poll is 36 polls of margin.)
### Where Ollama lives
```bash
ollama pull mistral-nemo:12b # or whatever you set as classifier.model
```
It does not have to be on the machine running the router; the box with the
GPU usually isn't the laptop. To use one across a VPN, point **both**
endpoints at it:
```yaml
classifier:
base_url: "http://:11434/v1"
verification:
base_url: "http://:11434" # same host, so `model` can stay null
```
and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds
`127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address
rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the
port can run inference and enumerate your models.
Both endpoints move together because the verifier speaks Ollama's *native*
API and cannot follow the classifier to a cloud provider. Config load refuses
the case where they are on different hosts and `verification.model` is null,
because that combination fails silently.
## Usage
### Route without spending anything
Use `POST /route` when you want to see what the router would pick for a task before paying for a provider call.
```bash
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view into service objects."}'
```
- Runs the local classifier to determine category, tier, and required context.
- Applies hard filters and ranks the surviving candidates.
- Returns the selected model, estimated cost, and estimated proficiency.
- **No provider call is made and no quota is consumed.**
### Skip the classifier when you already know the shape
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`, and `required_context_tokens` overrides.
```bash
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}'
```
- The classifier is not called, so latency is just the routing pass.
- `route_decisions.classification_source` is logged as `override`.
### Actually dispatch and log energy
`POST /dispatch` does the same routing work as `/route` and then calls the selected provider, streams the response if requested, and records the request as an `energy_observations` row.
```bash
curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
-d '{"task":"What is a Python context manager?"}'
```
- Provider response is proxied back, including streaming chunks.
- Energy, carbon, cost, and duration are scraped from SSE comments and logged.
### Report whether it actually worked
Every other quality signal is a proxy — structural checks know code *parses*,
the local LLM check guesses prose *looks* right. Only the client knows whether
the answer did the job, because it ran the tests.
```bash
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
```
- The only signal that survives streaming — a retry can't reach bytes already
sent, but a report arrives afterward and works either way.
- Successes count too: unlike the automatic checks, `feedback.py` folds a
reported `ok: true` in as a real success, not just failures.
- An unknown `request_id` returns `404` rather than being silently accepted.
### Point any OpenAI-compatible client at it
The `/v1` endpoints speak the OpenAI completions and models API.
```bash
curl -s localhost:8080/v1/models
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'
```
Ask for a **virtual router model** and the router picks a candidate, subject
to the profile you specify. The general form is `auto:`; asking for
just `auto` is shorthand for `auto:default`.
```bash
# Normal interactive routing (same as "auto")
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":...}'
# Restrict to local dispatch models (ollama-local provider only)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:locality","messages":...}'
# Restrict to models priced at ≤ $0.50/1M completion tokens
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:onlycheaps","messages":...}'
# Restrict to frontier tier (tier 3)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:bigboybritches","messages":...}'
# Overnight/async work via flex rows
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:batch","messages":...}'
```
| Profile | What it does |
|---|---|
| `auto` / `auto:default` | Normal quality-first routing; `-flex` rows excluded |
| `auto:batch` | Admits `-flex` rows that may be held during peak hours |
| `auto:locality` | Routes only to the `ollama-local` provider (local dispatch) |
| `auto:onlycheaps` | Limits to models at or below $0.50 per 1M completion tokens |
| `auto:bigboybritches` | Routes only to tier-3 (frontier) models |
Profiles are **candidate-set filters** — they narrow which models the router
may pick from, but do not change the ranking objective (quality-first, cost as
tiebreak). An unknown profile raises HTTP 422.
Operators can define custom profiles under `profiles:` in
`config/config.yaml`. Each entry accepts `provider`, `min_tier`/`max_tier`,
`max_cost_per_1m_completion`, `latency_tolerance`, and `allowed_model_ids`
as rewrite rules on top of the default profile.
- Ask for any real model id and the router dispatches directly, still logged.
- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines.
### Ask an image question
Cloud vision is not universal in the catalog. Requests carrying `image_url` parts
are routed only to vision-capable catalog rows. If none survives, the router can
fall back to a local vision model instead of returning 422.
```bash
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}'
```
- Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined.
- Configure the fallback in `config/config.yaml` under `local_vision:`.
### Force JSON output
Use `response_format` when you need structured output. The router treats this as
a hard capability requirement and only admits rows that declare `supports_json_mode = 1`.
```bash
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}'
```
- `NULL` fails closed: an unknown flag means the capability cannot be confirmed.
- When streaming is not used, the structural verification verdict appears in the `X-Router-Verification` header.
### Watch it live
Three read-only ways to observe the router without spending quota:
```bash
# Aggregate JSON health/usage summary
curl -s localhost:8080/metrics | python -m json.tool
# Live SSE stream of routing decisions
curl -s localhost:8080/events/decisions
# Terminal dashboard with live routing feed
PYTHONPATH=src python -m tui
```
- **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency.
- **`GET /events/decisions`** is a Server-Sent Events stream; replays recent decisions, then streams new ones.
- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels.
### Compare against trivial baselines
`baseline_report.py` replays recent `route_decisions` against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend.
```bash
PYTHONPATH=src python baseline_report.py --since 2026-08-01
PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv
```
A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice.
### Admin web portal
For the full admin portal walkthrough, see [docs/admin-portal.md](docs/admin-portal.md).
### Probe routing without spending
`router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree.
```bash
PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
```
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
- Use `--category`, `--tier`, and `--context` to override the classifier deterministically.
## At a Glance
| Dimension | Detail |
|---|---|
| **Cost model** | Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog. |
| **Latency** | Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response. |
| **Quality** | Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior. |
| **Verification** | Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via `feedback.py`. |
| **Fault tolerance** | Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent `timeout_seconds` silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate. |
| **API surface** | OpenAI-compatible `/v1` endpoints, streaming chunk proxy with SSE telemetry scraping. |
## Architecture
```
┌─────────────────────┐
incoming task ───▶│ Local Classifier │ Ollama (mistral-nemo:12b)
│ - task_category │ ~4s warm, ~120s cold cap
│ - task_tier │ temperature: 0, max 1024 tokens
│ - required_context │ max_retries: 0 (silent 3× cap guard)
│ - confidence │
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Classifier Fallback │ tier 2 / general_chat on failure
│ (graceful degrade) │ — not a 502/503
└──────────┬───────────┘
▼
┌─────────────────┐
│ Escalation │ low-confidence tier bump
│ (optional) │ threshold: 0.6 confidence
└────────┬─────────┘
▼
┌─────────────────────┐
│ Hard Filters │ context window ≥ required
│ (routing.py) │ tier ≥ required
│ │ freshness (active, not stale)
│ │ access_level allowed
│ │ latency_class compatible
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Quality-first select │ max proficiency, cheapest
│ (routing.py) │ among equals, under a
│ │ per-request kWh ceiling
└──────────┬───────────┘
▼
┌──────────────────────┐
│ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al.
│ (FastAPI API) │──▶ OpenAI-compatible /v1
│ │ Streaming SSE + energy scrape
└───────┬──────────────┘
▼
┌─────────────────────────┐
│ Verification Pipeline │
│ Layer 1: Structural │ ast.parse, json.load, yaml.safe_load
│ Layer 2: Local LLM │ async, ~6s, >600 token gate
└──────┬──────────────────┘
│
▼
┌─────────────────────┐
│ feedback.py │ failures → proficiency updates
│ (on-demand agent) │ idempotent, attributed-only
└─────────────────────┘
```
Not pictured: a model that 5xxs is passively excluded from Hard Filters by
`circuit_breaker.py` until its cooldown clears, and a verification failure can
loop back into Dispatcher for a tier-budgeted retry before falling through to
`feedback.py` — see [docs/verification.md](docs/verification.md#retry-budget--iterationpy).
For the detailed module-by-module map, see [docs/architecture.md](docs/architecture.md).
## Tech Stack
| Layer | Technology |
|---|---|
| **Language** | Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed) |
| **Framework** | FastAPI + uvicorn |
| **Database** | SQLite (`router.db`) — decision table, energy observations, proficiency, verifications |
| **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN |
| **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works |
| **Cloud Providers** | OpenAI-compatible carriers via `dispatch_providers:` — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works |
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
| **Testing** | `pytest` — 1923 tests across ~90 files, all offline |
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
Dependencies are pinned in `requirements.txt` — recreate the venv with those exact versions to avoid silent drift. **Bump deliberately**: the old `>=` ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use `>=`.
## Testing
```bash
python -m pytest # 1923 tests
python -m pytest --cov # with coverage
```
No test calls a provider or a local model — the pure modules take rows and
config as arguments, so the suite runs offline on a clean checkout.
`pyproject.toml` puts the repo root on `sys.path` for pytest — the modules live
at the root rather than in a package, so `pytest` (console script) and
`python -m pytest` would otherwise disagree about whether `import config`
resolves.
Representative test files (~90 files total; the list below is a useful subset, not exhaustive):
| Test file | What it covers |
|---|---|
| `test_scoring.py` | `normalize_inverted`, `composite_score`, None-handling |
| `test_routing.py` | Hard filters, candidate selection & ranking across all 9 categories × 3 tiers |
| `test_tiering.py` | Tier resolver precedence: override → reasoning → cost + context window → mid |
| `test_apply_tiering.py` | DB tiering pass, sanity guards |
| `test_poller_parsing.py` | Serving class, base model, access level parsing |
| `test_proficiency.py` | Blending rule, accumulation |
| `test_load_candidates.py` | `load_candidates`: cost/eco/proficiency join |
| `test_eval_scoring.py` | Code, exact, tool, and judge scoring |
| `test_task_set.py` | Reference solutions validating every `code` task's checks & `exact` answers |
| `test_verification.py` | Structural checks on every language, verifier edge cases, local-LLM gating |
| `test_feedback.py` | Failure identification, idempotency, attribution filtering |
| `test_iteration.py` | Retry budget per tier, and matching the retry to the failure kind |
| `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal |
| `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split |
| `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior |
| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
| `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
| `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
| `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy |
| `test_classifier_input.py` | Classifier framing: `_previous_context` scope, `_classifier_user_content` framing |
| `test_route_decisions.py` | `route_decisions` table, inline-create helper, config gate |
| `test_admin_*.py` | Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family) |
| `test_local_dispatch*.py` | Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep |
| `test_multi_provider*.py` | Multi-provider dispatch + poller; per-provider balance & attenuation |
| `test_circuit_breaker.py` | Passive circuit breaker: cooldown + backoff on model 5xx |
| `test_session_cache.py` | Per-session category/tier caching to skip repeat classifier round-trips |
| `test_outcome_attribution.py` | Outcome attribution & ambiguity refusal (streaming answer path) |
## Known Limitations & Open Items
- **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. See [CLAUDE.md](CLAUDE.md) "What's NOT built yet" for the live list.
- **Three models unsettled** — split-half stability varies too much for some model positions.
- **Retry does not reach streaming** — corrective attempts work only on the non-streaming path; `POST /outcome` is the streamed answer.
- **Local dispatch has its own limits** — no true streaming (answer is buffered), no `verifications` rows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see [docs/routing.md](docs/routing.md)).
- **No auth** — the service holds a billable API key with no authentication of its own; loopback is the only guard.
- **Local energy is metered when enabled** — `local_energy:` in `config/config.yaml` gates classifier/verifier/vision draw on a separate `local_energy_observations` table (off by default, refuses `enabled: true` without a tariff rate). See [data-model.md](docs/data-model.md).
- **Context assembly (RAG)** is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.