2026-10-05 21:32:12 -04:00
2026-10-05 21:32:12 -04:00

6krrt logo
6krrt — Local LLM Model Router

A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs. It classifies each incoming task with a local Ollama model — category, tier, required context — and dispatches to the highest-expected-success open-weight model on an OpenAI-compatible provider, under a per-request cost ceiling and tiebroken by price. Point model: "auto" at POST /v1/chat/completions for automatic selection; pin a real model id and it dispatches as asked, still logged.

Any OpenAI-compatible carrier should work. NeuralWatt is the reference provider — it powers every measurement in this README — and uniquely exposes per-request power/energy telemetry that the router ingests for cost accounting; other carriers that report cost only in their usage response are supported but carry no energy telemetry. Configure additional providers and their API keys under dispatch_providers: in config/config.yaml.

Numbers in this README are measurements, not specifications. They come from one deployment against one provider account, and the catalog, prices, grid intensity and pool load all move. They are here because the reasoning behind a design choice is worth more than the choice, and re-running the measurement is how you check whether it still holds for you.

Features

  • Check routing before spending anything. POST /route classifies, ranks, and returns the selected model with no provider call and no cost.
  • Price per request from the actual shape of the traffic. Cost is estimated from catalog token prices scaled to the prompt size, an assumed completion length, and an assumed cache rate — not a single fixed benchmark.
  • Dispatch selected tasks to a local model. When the classifier puts a task in file_summarization or diff_checking, the router will consider the local model (qwen2.5-coder-router:14b) only when it scores within the quality tolerance of the cloud leader. (Dormant under the default profile — local dispatch does not routinely win.)
  • Fall back to local vision when no cloud row supports images. If no vision-capable catalog candidate survives the hard filters, the router proxies the request to a local Ollama vision model instead of returning 422.
  • Watch decisions arrive live. PYTHONPATH=src python -m tui opens a terminal dashboard that follows /events/decisions as decisions are recorded, with no polling delay.
  • Verify before learning. Every routed response is structurally parsed in the background; larger prose answers get an async local-LLM spot-check, and failures fold back into per-model proficiency through feedback.py.

Documentation

Requirements

  • API key(s) for the provider(s) you route to — a NeuralWatt key by default, plus any others under dispatch_providers: (e.g. an OpenRouter key). The reference deployment uses NeuralWatt, whose energy telemetry feeds cost accounting.
  • Python 3.10+ (the test suite is verified on 3.10 and 3.14).
  • An Ollama reachable from wherever this runs, with a classifier model pulled.
  • For local dispatch: qwen2.5-coder-router:14b created from qwen2.5-coder:14b with num_ctx 32768 (see docs/local-models.md).

Nothing else is assumed about the host — routing itself is SQLite and arithmetic.

Installation

Local (venv)

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env                          # fill in NEURALWATT_API_KEY (.env stays at repo root)
cp config/config.local.yaml.example config/config.local.yaml   # optional: deployment-specific overlay (gitignored)
PYTHONPATH=src python -m poller       # populate the catalog
PYTHONPATH=src python -m tier         # resolve tiers
PYTHONPATH=src python -m config       # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081

Then edit config/config.yaml for your own setup — at minimum:

Key Why
classifier.model must match a model ollama list reports
classifier.base_url where that Ollama actually is
objective.plan_kwh_per_period your plan's quota; /health reports burn against it
objective.assumed_cache_rate 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures
session_cache.enabled off by default; caches category/tier per session for staleness_seconds to skip repeat classifier round-trips on long agent sessions

PYTHONPATH=src python -m seed_energy is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path.

As a systemd service

Five user services (with companion timers) cover continuous dispatch, catalog polling, periodic energy reseeding, and database backups/offsite snapshots. See deploy/README.md for full instructions.

The poller timer is load-bearing, not optional — but it fails silently, not loudly. mark_stale runs only inside a poll run that got past the fetch, so a stopped timer or a provider outage marks nothing: the catalog freezes at last-known-good and the router keeps routing on prices that may be weeks old, with every row still reading active. Watch for staleness; don't expect it to announce itself. (freshness.stale_after_days is 3 with exclude_stale: true, which against a 2-hourly poll is 36 polls of margin.)

Where Ollama lives

ollama pull mistral-nemo:12b   # or whatever you set as classifier.model

It does not have to be on the machine running the router; the box with the GPU usually isn't the laptop. To use one across a VPN, point both endpoints at it:

classifier:
  base_url: "http://<vpn-ip>:11434/v1"
verification:
  base_url: "http://<vpn-ip>:11434"    # same host, so `model` can stay null

and apply deploy/ollama-over-vpn.conf on the serving host — Ollama binds 127.0.0.1 by default and will otherwise refuse. Bind it to the VPN address rather than 0.0.0.0: Ollama has no authentication, so anything reaching the port can run inference and enumerate your models.

Both endpoints move together because the verifier speaks Ollama's native API and cannot follow the classifier to a cloud provider. Config load refuses the case where they are on different hosts and verification.model is null, because that combination fails silently.

Usage

Route without spending anything

Use POST /route when you want to see what the router would pick for a task before paying for a provider call.

curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
  -d '{"task":"Refactor this 800-line Django view into service objects."}'
  • Runs the local classifier to determine category, tier, and required context.
  • Applies hard filters and ranks the surviving candidates.
  • Returns the selected model, estimated cost, and estimated proficiency.
  • No provider call is made and no quota is consumed.

Skip the classifier when you already know the shape

Input to /route and /dispatch can include task_category, task_tier, and required_context_tokens overrides.

curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
  -d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}'
  • The classifier is not called, so latency is just the routing pass.
  • route_decisions.classification_source is logged as override.

Actually dispatch and log energy

POST /dispatch does the same routing work as /route and then calls the selected provider, streams the response if requested, and records the request as an energy_observations row.

curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
  -d '{"task":"What is a Python context manager?"}'
  • Provider response is proxied back, including streaming chunks.
  • Energy, carbon, cost, and duration are scraped from SSE comments and logged.

Report whether it actually worked

Every other quality signal is a proxy — structural checks know code parses, the local LLM check guesses prose looks right. Only the client knows whether the answer did the job, because it ran the tests.

# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
  • The only signal that survives streaming — a retry can't reach bytes already sent, but a report arrives afterward and works either way.
  • Successes count too: unlike the automatic checks, feedback.py folds a reported ok: true in as a real success, not just failures.
  • An unknown request_id returns 404 rather than being silently accepted.

Point any OpenAI-compatible client at it

The /v1 endpoints speak the OpenAI completions and models API.

curl -s localhost:8080/v1/models

curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'

Ask for a virtual router model and the router picks a candidate, subject to the profile you specify. The general form is auto:<profile>; asking for just auto is shorthand for auto:default.

# Normal interactive routing (same as "auto")
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":...}'

# Restrict to local dispatch models (ollama-local provider only)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto:locality","messages":...}'

# Restrict to models priced at ≤ $0.50/1M completion tokens
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto:onlycheaps","messages":...}'

# Restrict to frontier tier (tier 3)
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto:bigboybritches","messages":...}'

# Overnight/async work via flex rows
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto:batch","messages":...}'
Profile What it does
auto / auto:default Normal quality-first routing; -flex rows excluded
auto:batch Admits -flex rows that may be held during peak hours
auto:locality Routes only to the ollama-local provider (local dispatch)
auto:onlycheaps Limits to models at or below $0.50 per 1M completion tokens
auto:bigboybritches Routes only to tier-3 (frontier) models

Profiles are candidate-set filters — they narrow which models the router may pick from, but do not change the ranking objective (quality-first, cost as tiebreak). An unknown profile raises HTTP 422.

Operators can define custom profiles under profiles: in config/config.yaml. Each entry accepts provider, min_tier/max_tier, max_cost_per_1m_completion, latency_tolerance, and allowed_model_ids as rewrite rules on top of the default profile.

  • Ask for any real model id and the router dispatches directly, still logged.
  • Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines.

Ask an image question

Cloud vision is not universal in the catalog. Requests carrying image_url parts are routed only to vision-capable catalog rows. If none survives, the router can fall back to a local vision model instead of returning 422.

curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}'
  • Only inline data: URIs are accepted; remote http(s) image URLs are declined.
  • Configure the fallback in config/config.yaml under local_vision:.

Force JSON output

Use response_format when you need structured output. The router treats this as a hard capability requirement and only admits rows that declare supports_json_mode = 1.

curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}'
  • NULL fails closed: an unknown flag means the capability cannot be confirmed.
  • When streaming is not used, the structural verification verdict appears in the X-Router-Verification header.

Watch it live

Three read-only ways to observe the router without spending quota:

# Aggregate JSON health/usage summary
curl -s localhost:8080/metrics | python -m json.tool

# Live SSE stream of routing decisions
curl -s localhost:8080/events/decisions

# Terminal dashboard with live routing feed
PYTHONPATH=src python -m tui
  • GET /metrics returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency.
  • GET /events/decisions is a Server-Sent Events stream; replays recent decisions, then streams new ones.
  • PYTHONPATH=src python -m tui is a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels.

Compare against trivial baselines

baseline_report.py replays recent route_decisions against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend.

PYTHONPATH=src python baseline_report.py --since 2026-08-01
PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv

A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice.

Admin web portal

For the full admin portal walkthrough, see docs/admin-portal.md.

Probe routing without spending

router_cli.py is a one-shot shell probe that POSTs to /route once and prints the full decision tree.

PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
  • Prints the selected model, candidates, estimated cost, and estimated proficiency.
  • Use --category, --tier, and --context to override the classifier deterministically.

At a Glance

Dimension Detail
Cost model Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog.
Latency Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response.
Quality Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior.
Verification Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via feedback.py.
Fault tolerance Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent timeout_seconds silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate.
API surface OpenAI-compatible /v1 endpoints, streaming chunk proxy with SSE telemetry scraping.

Architecture

                     ┌─────────────────────┐
   incoming task ───▶│  Local Classifier    │  Ollama (mistral-nemo:12b)
                     │  - task_category     │  ~4s warm, ~120s cold cap
                     │  - task_tier         │  temperature: 0, max 1024 tokens
                     │  - required_context  │  max_retries: 0 (silent 3× cap guard)
                     │  - confidence        │
                     └──────────┬───────────┘
                                ▼
                     ┌─────────────────────┐
                     │ Classifier Fallback  │  tier 2 / general_chat on failure
                     │ (graceful degrade)   │  — not a 502/503
                     └──────────┬───────────┘
                                ▼
                        ┌─────────────────┐
                        │ Escalation       │  low-confidence tier bump
                        │ (optional)       │  threshold: 0.6 confidence
                        └────────┬─────────┘
                                ▼
                     ┌─────────────────────┐
                     │ Hard Filters         │  context window ≥ required
                     │ (routing.py)         │  tier ≥ required
                     │                      │  freshness (active, not stale)
                     │                      │  access_level allowed
                     │                      │  latency_class compatible
                     └──────────┬───────────┘
                                ▼
                     ┌─────────────────────┐
                     │ Quality-first select │  max proficiency, cheapest
                     │ (routing.py)         │  among equals, under a
                     │                      │  per-request kWh ceiling
                     └──────────┬───────────┘
                                ▼
                       ┌──────────────────────┐
                       │ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al.
                       │ (FastAPI API)         │──▶ OpenAI-compatible /v1
                       │                       │    Streaming SSE + energy scrape
                       └───────┬──────────────┘
                              ▼
                ┌─────────────────────────┐
                │  Verification Pipeline   │
                │  Layer 1: Structural     │  ast.parse, json.load, yaml.safe_load
                │  Layer 2: Local LLM      │  async, ~6s, >600 token gate
                └──────┬──────────────────┘
                       │
                       ▼
                ┌─────────────────────┐
                │  feedback.py        │  failures → proficiency updates
                │  (on-demand agent)   │  idempotent, attributed-only
                └─────────────────────┘

Not pictured: a model that 5xxs is passively excluded from Hard Filters by circuit_breaker.py until its cooldown clears, and a verification failure can loop back into Dispatcher for a tier-budgeted retry before falling through to feedback.py — see docs/verification.md.

For the detailed module-by-module map, see docs/architecture.md.

Tech Stack

Layer Technology
Language Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed)
Framework FastAPI + uvicorn
Database SQLite (router.db) — decision table, energy observations, proficiency, verifications
Local Classification Ollama, OpenAI-compatible — localhost:11434/v1 or an Ollama across your VPN
Local Model classifier.model — mistral-nemo:12b by default; any Ollama model works
Cloud Providers OpenAI-compatible carriers via dispatch_providers: — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works
Config config/config.yaml loaded & validated by Pydantic (src/config.py)
OpenAI Client openai==3.0.0 (official SDK)
HTTP requests for poller, httpx (via openai/uvicorn)
Testing pytest — 1923 tests across ~90 files, all offline
Config Files config/config.yaml, config/leaderboards.yaml, evals/tasks.yaml
Deployment systemd user units (.service + .timer files in deploy/)
Integration opencode.json in the repo routes through it by default; any OpenAI-compatible client works

Dependencies are pinned in requirements.txt — recreate the venv with those exact versions to avoid silent drift. Bump deliberately: the old >= ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use >=.

Testing

python -m pytest        # 1923 tests
python -m pytest --cov  # with coverage

No test calls a provider or a local model — the pure modules take rows and config as arguments, so the suite runs offline on a clean checkout.

pyproject.toml puts the repo root on sys.path for pytest — the modules live at the root rather than in a package, so pytest (console script) and python -m pytest would otherwise disagree about whether import config resolves.

Representative test files (~90 files total; the list below is a useful subset, not exhaustive):

Test file What it covers
test_scoring.py normalize_inverted, composite_score, None-handling
test_routing.py Hard filters, candidate selection & ranking across all 9 categories × 3 tiers
test_tiering.py Tier resolver precedence: override → reasoning → cost + context window → mid
test_apply_tiering.py DB tiering pass, sanity guards
test_poller_parsing.py Serving class, base model, access level parsing
test_proficiency.py Blending rule, accumulation
test_load_candidates.py load_candidates: cost/eco/proficiency join
test_eval_scoring.py Code, exact, tool, and judge scoring
test_task_set.py Reference solutions validating every code task's checks & exact answers
test_verification.py Structural checks on every language, verifier edge cases, local-LLM gating
test_feedback.py Failure identification, idempotency, attribution filtering
test_iteration.py Retry budget per tier, and matching the retry to the failure kind
test_session_identity.py Outcome attribution: session matching, ambiguity refusal
test_config_endpoints.py Classifier and verifier are separately addressable; guards on the split
test_metrics_endpoint.py /metrics endpoint, SSE /events/decisions headers + replay/stream behavior
test_admin_frontend.py Admin portal serves the four HTML pages with their expected markers and Chart.js asset
test_events.py Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction
test_tui.py TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls
test_context_prune.py Context pruning: image_url handling, structured content, recency guards, stats accuracy
test_classifier_input.py Classifier framing: _previous_context scope, _classifier_user_content framing
test_route_decisions.py route_decisions table, inline-create helper, config gate
test_admin_*.py Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family)
test_local_dispatch*.py Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep
test_multi_provider*.py Multi-provider dispatch + poller; per-provider balance & attenuation
test_circuit_breaker.py Passive circuit breaker: cooldown + backoff on model 5xx
test_session_cache.py Per-session category/tier caching to skip repeat classifier round-trips
test_outcome_attribution.py Outcome attribution & ambiguity refusal (streaming answer path)

Known Limitations & Open Items

  • Leaderboard priors unfilled — leaderboards.yaml ships empty. See CLAUDE.md "What's NOT built yet" for the live list.
  • Three models unsettled — split-half stability varies too much for some model positions.
  • Retry does not reach streaming — corrective attempts work only on the non-streaming path; POST /outcome is the streamed answer.
  • Local dispatch has its own limits — no true streaming (answer is buffered), no verifications rows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see docs/routing.md).
  • No auth — the service holds a billable API key with no authentication of its own; loopback is the only guard.
  • Local energy is metered when enabled — local_energy: in config/config.yaml gates classifier/verifier/vision draw on a separate local_energy_observations table (off by default, refuses enabled: true without a tariff rate). See data-model.md.
  • Context assembly (RAG) is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.
Description
Local LLM router I built to use with Opencode and NeuralWatt
Readme 42 MiB
Languages
Python 80%
HTML 12.7%
JavaScript 6.6%
Shell 0.5%
CSS 0.2%