Both fields are read nowhere outside config.py; item E is a design sketch only. An admin toggle would control nothing real until that feature is implemented. The existing gap in test_admin_knob_coverage.py (the whole classifier.* section sits outside the generic registries) is pre-existing and not addressed here.

6krrt — Local LLM Model Router
A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs.
It classifies each incoming task with a local Ollama model — category, tier,
required context — and dispatches to the highest-expected-success open-weight
model on an OpenAI-compatible provider, under a per-request cost ceiling and
tiebroken by price. Point model: "auto" at POST /v1/chat/completions for
automatic selection; pin a real model id and it dispatches as asked, still
logged.
Any OpenAI-compatible carrier should work. NeuralWatt is the reference
provider — it powers every measurement in this README — and uniquely exposes
per-request power/energy telemetry that the router ingests for cost
accounting; other carriers that report cost only in their usage response are
supported but carry no energy telemetry. Configure additional providers and
their API keys under dispatch_providers: in config/config.yaml.
Numbers in this README are measurements, not specifications. They come from one deployment against one provider account, and the catalog, prices, grid intensity and pool load all move. They are here because the reasoning behind a design choice is worth more than the choice, and re-running the measurement is how you check whether it still holds for you.
Features
- Check routing before spending anything.
POST /routeclassifies, ranks, and returns the selected model with no provider call and no cost. - Price per request from the actual shape of the traffic. Cost is estimated from catalog token prices scaled to the prompt size, an assumed completion length, and an assumed cache rate — not a single fixed benchmark.
- Dispatch selected tasks to a local model. When the classifier puts a task
in
file_summarizationordiff_checking, the router will consider the local model (qwen2.5-coder-router:14b) only when it scores within the quality tolerance of the cloud leader. (Dormant under the default profile — local dispatch does not routinely win.) - Fall back to local vision when no cloud row supports images. If no vision-capable catalog candidate survives the hard filters, the router proxies the request to a local Ollama vision model instead of returning 422.
- Watch decisions arrive live.
PYTHONPATH=src python -m tuiopens a terminal dashboard that follows/events/decisionsas decisions are recorded, with no polling delay. - Verify before learning. Every routed response is structurally parsed in
the background; larger prose answers get an async local-LLM spot-check, and
failures fold back into per-model proficiency through
feedback.py.
Documentation
- Routing internals — quality-first ranking, local vision fallback, candidate ranking
- Data model — SQLite schema: models, proficiency, energy, verifications, route_decisions
- Response verification — structural + local-LLM checks, feedback loop
- API reference — endpoints, virtual models, streaming, logging headers
- Client setup — pointing opencode / OpenAI-compatible clients at the router
- Architecture — module-by-module map
- Operations — logfmt, journalctl, tracing requests
- Incidents — how the router has broken, with a symptom → one-line-check table
- Evaluation & classifier — self-eval harness, classifier reliability notes
- Admin portal — browser dashboard, controls, decision log
- Local model sizing — fitting Ollama models into 24GB VRAM
- Context pruning (pinch) — pre-dispatch conversation trimming once a session grows past a token budget
- Local config overlay — per-machine
config.local.yamloverrides (gitignored)
Requirements
- API key(s) for the provider(s) you route to — a NeuralWatt key by default,
plus any others under
dispatch_providers:(e.g. an OpenRouter key). The reference deployment uses NeuralWatt, whose energy telemetry feeds cost accounting. - Python 3.10+ (the test suite is verified on 3.10 and 3.14).
- An Ollama reachable from wherever this runs, with a classifier model pulled.
- For local dispatch:
qwen2.5-coder-router:14bcreated fromqwen2.5-coder:14bwithnum_ctx 32768(see docs/local-models.md).
Nothing else is assumed about the host — routing itself is SQLite and arithmetic.
Installation
Local (venv)
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
cp config/config.local.yaml.example config/config.local.yaml # optional: deployment-specific overlay (gitignored)
PYTHONPATH=src python -m poller # populate the catalog
PYTHONPATH=src python -m tier # resolve tiers
PYTHONPATH=src python -m config # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
Then edit config/config.yaml for your own setup — at minimum:
| Key | Why |
|---|---|
classifier.model |
must match a model ollama list reports |
classifier.base_url |
where that Ollama actually is |
objective.plan_kwh_per_period |
your plan's quota; /health reports burn against it |
objective.assumed_cache_rate |
0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
session_cache.enabled |
off by default; caches category/tier per session for staleness_seconds to skip repeat classifier round-trips on long agent sessions |
PYTHONPATH=src python -m seed_energy is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path.
As a systemd service
Five user services (with companion timers) cover continuous dispatch, catalog
polling, periodic energy reseeding, and database backups/offsite snapshots. See
deploy/README.md for full instructions.
The poller timer is load-bearing, not optional — but it fails silently,
not loudly. mark_stale runs only inside a poll run that got past the fetch, so
a stopped timer or a provider outage marks nothing: the catalog freezes at
last-known-good and the router keeps routing on prices that may be weeks old,
with every row still reading active. Watch for staleness; don't expect it to
announce itself. (freshness.stale_after_days is 3 with exclude_stale: true,
which against a 2-hourly poll is 36 polls of margin.)
Where Ollama lives
ollama pull mistral-nemo:12b # or whatever you set as classifier.model
It does not have to be on the machine running the router; the box with the GPU usually isn't the laptop. To use one across a VPN, point both endpoints at it:
classifier:
base_url: "http://<vpn-ip>:11434/v1"
verification:
base_url: "http://<vpn-ip>:11434" # same host, so `model` can stay null
and apply deploy/ollama-over-vpn.conf on the serving host — Ollama binds
127.0.0.1 by default and will otherwise refuse. Bind it to the VPN address
rather than 0.0.0.0: Ollama has no authentication, so anything reaching the
port can run inference and enumerate your models.
Both endpoints move together because the verifier speaks Ollama's native
API and cannot follow the classifier to a cloud provider. Config load refuses
the case where they are on different hosts and verification.model is null,
because that combination fails silently.
Usage
Route without spending anything
Use POST /route when you want to see what the router would pick for a task before paying for a provider call.
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view into service objects."}'
- Runs the local classifier to determine category, tier, and required context.
- Applies hard filters and ranks the surviving candidates.
- Returns the selected model, estimated cost, and estimated proficiency.
- No provider call is made and no quota is consumed.
Skip the classifier when you already know the shape
Input to /route and /dispatch can include task_category, task_tier, and required_context_tokens overrides.
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
-d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}'
- The classifier is not called, so latency is just the routing pass.
route_decisions.classification_sourceis logged asoverride.
Actually dispatch and log energy
POST /dispatch does the same routing work as /route and then calls the selected provider, streams the response if requested, and records the request as an energy_observations row.
curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
-d '{"task":"What is a Python context manager?"}'
- Provider response is proxied back, including streaming chunks.
- Energy, carbon, cost, and duration are scraped from SSE comments and logged.
Report whether it actually worked
Every other quality signal is a proxy — structural checks know code parses, the local LLM check guesses prose looks right. Only the client knows whether the answer did the job, because it ran the tests.
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
- The only signal that survives streaming — a retry can't reach bytes already sent, but a report arrives afterward and works either way.
- Successes count too: unlike the automatic checks,
feedback.pyfolds a reportedok: truein as a real success, not just failures. - An unknown
request_idreturns404rather than being silently accepted.
Point any OpenAI-compatible client at it
The /v1 endpoints speak the OpenAI completions and models API.
curl -s localhost:8080/v1/models
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'
Ask for a virtual router model and the router picks a candidate, subject
to the profile you specify. The general form is auto:<profile>; asking for
just auto is shorthand for auto:default.
# Normal interactive routing (same as "auto")
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":...}'
# Restrict to local dispatch models (ollama-local provider only)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:locality","messages":...}'
# Restrict to models priced at ≤ $0.50/1M completion tokens
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:onlycheaps","messages":...}'
# Restrict to frontier tier (tier 3)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:bigboybritches","messages":...}'
# Overnight/async work via flex rows
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto:batch","messages":...}'
| Profile | What it does |
|---|---|
auto / auto:default |
Normal quality-first routing; -flex rows excluded |
auto:batch |
Admits -flex rows that may be held during peak hours |
auto:locality |
Routes only to the ollama-local provider (local dispatch) |
auto:onlycheaps |
Limits to models at or below $0.50 per 1M completion tokens |
auto:bigboybritches |
Routes only to tier-3 (frontier) models |
Profiles are candidate-set filters — they narrow which models the router may pick from, but do not change the ranking objective (quality-first, cost as tiebreak). An unknown profile raises HTTP 422.
Operators can define custom profiles under profiles: in
config/config.yaml. Each entry accepts provider, min_tier/max_tier,
max_cost_per_1m_completion, latency_tolerance, and allowed_model_ids
as rewrite rules on top of the default profile.
- Ask for any real model id and the router dispatches directly, still logged.
- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines.
Ask an image question
Cloud vision is not universal in the catalog. Requests carrying image_url parts
are routed only to vision-capable catalog rows. If none survives, the router can
fall back to a local vision model instead of returning 422.
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}'
- Only inline
data:URIs are accepted; remotehttp(s)image URLs are declined. - Configure the fallback in
config/config.yamlunderlocal_vision:.
Force JSON output
Use response_format when you need structured output. The router treats this as
a hard capability requirement and only admits rows that declare supports_json_mode = 1.
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}'
NULLfails closed: an unknown flag means the capability cannot be confirmed.- When streaming is not used, the structural verification verdict appears in the
X-Router-Verificationheader.
Watch it live
Three read-only ways to observe the router without spending quota:
# Aggregate JSON health/usage summary
curl -s localhost:8080/metrics | python -m json.tool
# Live SSE stream of routing decisions
curl -s localhost:8080/events/decisions
# Terminal dashboard with live routing feed
PYTHONPATH=src python -m tui
GET /metricsreturns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency.GET /events/decisionsis a Server-Sent Events stream; replays recent decisions, then streams new ones.PYTHONPATH=src python -m tuiis a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels.
Compare against trivial baselines
baseline_report.py replays recent route_decisions against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend.
PYTHONPATH=src python baseline_report.py --since 2026-08-01
PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv
A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice.
Admin web portal
For the full admin portal walkthrough, see docs/admin-portal.md.
Probe routing without spending
router_cli.py is a one-shot shell probe that POSTs to /route once and prints the full decision tree.
PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
- Use
--category,--tier, and--contextto override the classifier deterministically.
At a Glance
| Dimension | Detail |
|---|---|
| Cost model | Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog. |
| Latency | Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response. |
| Quality | Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior. |
| Verification | Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via feedback.py. |
| Fault tolerance | Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent timeout_seconds silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate. |
| API surface | OpenAI-compatible /v1 endpoints, streaming chunk proxy with SSE telemetry scraping. |
Architecture
┌─────────────────────┐
incoming task ───▶│ Local Classifier │ Ollama (mistral-nemo:12b)
│ - task_category │ ~4s warm, ~120s cold cap
│ - task_tier │ temperature: 0, max 1024 tokens
│ - required_context │ max_retries: 0 (silent 3× cap guard)
│ - confidence │
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Classifier Fallback │ tier 2 / general_chat on failure
│ (graceful degrade) │ — not a 502/503
└──────────┬───────────┘
▼
┌─────────────────┐
│ Escalation │ low-confidence tier bump
│ (optional) │ threshold: 0.6 confidence
└────────┬─────────┘
▼
┌─────────────────────┐
│ Hard Filters │ context window ≥ required
│ (routing.py) │ tier ≥ required
│ │ freshness (active, not stale)
│ │ access_level allowed
│ │ latency_class compatible
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Quality-first select │ max proficiency, cheapest
│ (routing.py) │ among equals, under a
│ │ per-request kWh ceiling
└──────────┬───────────┘
▼
┌──────────────────────┐
│ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al.
│ (FastAPI API) │──▶ OpenAI-compatible /v1
│ │ Streaming SSE + energy scrape
└───────┬──────────────┘
▼
┌─────────────────────────┐
│ Verification Pipeline │
│ Layer 1: Structural │ ast.parse, json.load, yaml.safe_load
│ Layer 2: Local LLM │ async, ~6s, >600 token gate
└──────┬──────────────────┘
│
▼
┌─────────────────────┐
│ feedback.py │ failures → proficiency updates
│ (on-demand agent) │ idempotent, attributed-only
└─────────────────────┘
Not pictured: a model that 5xxs is passively excluded from Hard Filters by
circuit_breaker.py until its cooldown clears, and a verification failure can
loop back into Dispatcher for a tier-budgeted retry before falling through to
feedback.py — see docs/verification.md.
For the detailed module-by-module map, see docs/architecture.md.
Tech Stack
| Layer | Technology |
|---|---|
| Language | Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed) |
| Framework | FastAPI + uvicorn |
| Database | SQLite (router.db) — decision table, energy observations, proficiency, verifications |
| Local Classification | Ollama, OpenAI-compatible — localhost:11434/v1 or an Ollama across your VPN |
| Local Model | classifier.model — mistral-nemo:12b by default; any Ollama model works |
| Cloud Providers | OpenAI-compatible carriers via dispatch_providers: — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works |
| Config | config/config.yaml loaded & validated by Pydantic (src/config.py) |
| OpenAI Client | openai==3.0.0 (official SDK) |
| HTTP | requests for poller, httpx (via openai/uvicorn) |
| Testing | pytest — 1923 tests across ~90 files, all offline |
| Config Files | config/config.yaml, config/leaderboards.yaml, evals/tasks.yaml |
| Deployment | systemd user units (.service + .timer files in deploy/) |
| Integration | opencode.json in the repo routes through it by default; any OpenAI-compatible client works |
Dependencies are pinned in requirements.txt — recreate the venv with those exact versions to avoid silent drift. Bump deliberately: the old >= ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use >=.
Testing
python -m pytest # 1923 tests
python -m pytest --cov # with coverage
No test calls a provider or a local model — the pure modules take rows and config as arguments, so the suite runs offline on a clean checkout.
pyproject.toml puts the repo root on sys.path for pytest — the modules live
at the root rather than in a package, so pytest (console script) and
python -m pytest would otherwise disagree about whether import config
resolves.
Representative test files (~90 files total; the list below is a useful subset, not exhaustive):
| Test file | What it covers |
|---|---|
test_scoring.py |
normalize_inverted, composite_score, None-handling |
test_routing.py |
Hard filters, candidate selection & ranking across all 9 categories × 3 tiers |
test_tiering.py |
Tier resolver precedence: override → reasoning → cost + context window → mid |
test_apply_tiering.py |
DB tiering pass, sanity guards |
test_poller_parsing.py |
Serving class, base model, access level parsing |
test_proficiency.py |
Blending rule, accumulation |
test_load_candidates.py |
load_candidates: cost/eco/proficiency join |
test_eval_scoring.py |
Code, exact, tool, and judge scoring |
test_task_set.py |
Reference solutions validating every code task's checks & exact answers |
test_verification.py |
Structural checks on every language, verifier edge cases, local-LLM gating |
test_feedback.py |
Failure identification, idempotency, attribution filtering |
test_iteration.py |
Retry budget per tier, and matching the retry to the failure kind |
test_session_identity.py |
Outcome attribution: session matching, ambiguity refusal |
test_config_endpoints.py |
Classifier and verifier are separately addressable; guards on the split |
test_metrics_endpoint.py |
/metrics endpoint, SSE /events/decisions headers + replay/stream behavior |
test_admin_frontend.py |
Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
test_events.py |
Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
test_tui.py |
TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
test_context_prune.py |
Context pruning: image_url handling, structured content, recency guards, stats accuracy |
test_classifier_input.py |
Classifier framing: _previous_context scope, _classifier_user_content framing |
test_route_decisions.py |
route_decisions table, inline-create helper, config gate |
test_admin_*.py |
Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family) |
test_local_dispatch*.py |
Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep |
test_multi_provider*.py |
Multi-provider dispatch + poller; per-provider balance & attenuation |
test_circuit_breaker.py |
Passive circuit breaker: cooldown + backoff on model 5xx |
test_session_cache.py |
Per-session category/tier caching to skip repeat classifier round-trips |
test_outcome_attribution.py |
Outcome attribution & ambiguity refusal (streaming answer path) |
Known Limitations & Open Items
- Leaderboard priors unfilled —
leaderboards.yamlships empty. See CLAUDE.md "What's NOT built yet" for the live list. - Three models unsettled — split-half stability varies too much for some model positions.
- Retry does not reach streaming — corrective attempts work only on the non-streaming path;
POST /outcomeis the streamed answer. - Local dispatch has its own limits — no true streaming (answer is buffered), no
verificationsrows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see docs/routing.md). - No auth — the service holds a billable API key with no authentication of its own; loopback is the only guard.
- Local energy is metered when enabled —
local_energy:inconfig/config.yamlgates classifier/verifier/vision draw on a separatelocal_energy_observationstable (off by default, refusesenabled: truewithout a tariff rate). See data-model.md. - Context assembly (RAG) is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.