6krrt logo
6krrt — Local LLM Model Router

A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs. It classifies each incoming task with a local Ollama model — category, tier, required context — and dispatches to the highest-expected-success open-weight model on an OpenAI-compatible provider, under a per-request cost ceiling and tiebroken by price. Point `model: "auto"` at `POST /v1/chat/completions` for automatic selection; pin a real model id and it dispatches as asked, still logged. Any OpenAI-compatible carrier should work. **NeuralWatt is the reference provider** — it powers every measurement in this README — and uniquely exposes per-request **power/energy telemetry** that the router ingests for cost accounting; other carriers that report cost only in their usage response are supported but carry no energy telemetry. Configure additional providers and their API keys under `dispatch_providers:` in `config/config.yaml`. **Numbers in this README are measurements, not specifications.** They come from one deployment against one provider account, and the catalog, prices, grid intensity and pool load all move. They are here because the reasoning behind a design choice is worth more than the choice, and re-running the measurement is how you check whether it still holds for you. ## Features - **Check routing before spending anything.** `POST /route` classifies, ranks, and returns the selected model with no provider call and no cost. - **Price per request from the actual shape of the traffic.** Cost is estimated from catalog token prices scaled to the prompt size, an assumed completion length, and an assumed cache rate — not a single fixed benchmark. - **Dispatch selected tasks to a local model.** When the classifier puts a task in `file_summarization` or `diff_checking`, the router will consider the local model (`qwen2.5-coder-router:14b`) only when it scores within the quality tolerance of the cloud leader. (Dormant under the default profile — local dispatch does not routinely win.) - **Fall back to local vision when no cloud row supports images.** If no vision-capable catalog candidate survives the hard filters, the router proxies the request to a local Ollama vision model instead of returning 422. - **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard that follows `/events/decisions` as decisions are recorded, with no polling delay. - **Verify before learning.** Every routed response is structurally parsed in the background; larger prose answers get an async local-LLM spot-check, and failures fold back into per-model proficiency through `feedback.py`. ## Documentation - [Routing internals](docs/routing.md) — quality-first ranking, local vision fallback, candidate ranking - [Data model](docs/data-model.md) — SQLite schema: models, proficiency, energy, verifications, route_decisions - [Response verification](docs/verification.md) — structural + local-LLM checks, feedback loop - [API reference](docs/api.md) — endpoints, virtual models, streaming, logging headers - [Client setup](docs/clients.md) — pointing opencode / OpenAI-compatible clients at the router - [Architecture](docs/architecture.md) — module-by-module map - [Operations](docs/operations.md) — logfmt, journalctl, tracing requests - [Incidents](docs/incidents.md) — how the router has broken, with a symptom → one-line-check table - [Evaluation & classifier](docs/evaluation.md) — self-eval harness, classifier reliability notes - [Admin portal](docs/admin-portal.md) — browser dashboard, controls, decision log - [Local model sizing](docs/local-models.md) — fitting Ollama models into 24GB VRAM - [Context pruning (pinch)](docs/pinch.md) — pre-dispatch conversation trimming once a session grows past a token budget - [Local config overlay](docs/config-local-overlay.md) — per-machine `config.local.yaml` overrides (gitignored) ## Requirements - API key(s) for the provider(s) you route to — a NeuralWatt key by default, plus any others under `dispatch_providers:` (e.g. an OpenRouter key). The reference deployment uses NeuralWatt, whose energy telemetry feeds cost accounting. - Python 3.10+ (the test suite is verified on 3.10 and 3.14). - An Ollama reachable from wherever this runs, with a classifier model pulled. - For local dispatch: `qwen2.5-coder-router:14b` created from `qwen2.5-coder:14b` with `num_ctx 32768` (see [docs/local-models.md](docs/local-models.md)). Nothing else is assumed about the host — routing itself is SQLite and arithmetic. ## Installation ### Local (`venv`) ```bash python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt sqlite3 router.db < config/schema.sql cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root) cp config/config.local.yaml.example config/config.local.yaml # optional: deployment-specific overlay (gitignored) PYTHONPATH=src python -m poller # populate the catalog PYTHONPATH=src python -m tier # resolve tiers PYTHONPATH=src python -m config # sanity-check config loads PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081 ``` Then edit `config/config.yaml` for your own setup — at minimum: | Key | Why | |---|---| | `classifier.model` | must match a model `ollama list` reports | | `classifier.base_url` | where that Ollama actually is | | `objective.plan_kwh_per_period` | your plan's quota; `/health` reports burn against it | | `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures | | `session_cache.enabled` | off by default; caches category/tier per session for `staleness_seconds` to skip repeat classifier round-trips on long agent sessions | `PYTHONPATH=src python -m seed_energy` is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path. ### As a systemd service Five user services (with companion timers) cover continuous dispatch, catalog polling, periodic energy reseeding, and database backups/offsite snapshots. See `deploy/README.md` for full instructions. **The poller timer is load-bearing, not optional** — but it fails *silently*, not loudly. `mark_stale` runs only inside a poll run that got past the fetch, so a stopped timer or a provider outage marks nothing: the catalog freezes at last-known-good and the router keeps routing on prices that may be weeks old, with every row still reading `active`. Watch for staleness; don't expect it to announce itself. (`freshness.stale_after_days` is 3 with `exclude_stale: true`, which against a 2-hourly poll is 36 polls of margin.) ### Where Ollama lives ```bash ollama pull mistral-nemo:12b # or whatever you set as classifier.model ``` It does not have to be on the machine running the router; the box with the GPU usually isn't the laptop. To use one across a VPN, point **both** endpoints at it: ```yaml classifier: base_url: "http://:11434/v1" verification: base_url: "http://:11434" # same host, so `model` can stay null ``` and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds `127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the port can run inference and enumerate your models. Both endpoints move together because the verifier speaks Ollama's *native* API and cannot follow the classifier to a cloud provider. Config load refuses the case where they are on different hosts and `verification.model` is null, because that combination fails silently. ## Usage ### Route without spending anything Use `POST /route` when you want to see what the router would pick for a task before paying for a provider call. ```bash curl -s -X POST localhost:8080/route -H 'content-type: application/json' \ -d '{"task":"Refactor this 800-line Django view into service objects."}' ``` - Runs the local classifier to determine category, tier, and required context. - Applies hard filters and ranks the surviving candidates. - Returns the selected model, estimated cost, and estimated proficiency. - **No provider call is made and no quota is consumed.** ### Skip the classifier when you already know the shape Input to `/route` and `/dispatch` can include `task_category`, `task_tier`, and `required_context_tokens` overrides. ```bash curl -s -X POST localhost:8080/route -H 'content-type: application/json' \ -d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}' ``` - The classifier is not called, so latency is just the routing pass. - `route_decisions.classification_source` is logged as `override`. ### Actually dispatch and log energy `POST /dispatch` does the same routing work as `/route` and then calls the selected provider, streams the response if requested, and records the request as an `energy_observations` row. ```bash curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \ -d '{"task":"What is a Python context manager?"}' ``` - Provider response is proxied back, including streaming chunks. - Energy, carbon, cost, and duration are scraped from SSE comments and logged. ### Report whether it actually worked Every other quality signal is a proxy — structural checks know code *parses*, the local LLM check guesses prose *looks* right. Only the client knows whether the answer did the job, because it ran the tests. ```bash # id comes from the completion body, or any stream chunk curl -s localhost:8080/outcome -H 'content-type: application/json' \ -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}' ``` - The only signal that survives streaming — a retry can't reach bytes already sent, but a report arrives afterward and works either way. - Successes count too: unlike the automatic checks, `feedback.py` folds a reported `ok: true` in as a real success, not just failures. - An unknown `request_id` returns `404` rather than being silently accepted. ### Point any OpenAI-compatible client at it The `/v1` endpoints speak the OpenAI completions and models API. ```bash curl -s localhost:8080/v1/models curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}' ``` Ask for a **virtual router model** and the router picks a candidate, subject to the profile you specify. The general form is `auto:`; asking for just `auto` is shorthand for `auto:default`. ```bash # Normal interactive routing (same as "auto") curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto","messages":...}' # Restrict to local dispatch models (ollama-local provider only) curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto:locality","messages":...}' # Restrict to models priced at ≤ $0.50/1M completion tokens curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto:onlycheaps","messages":...}' # Restrict to frontier tier (tier 3) curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto:bigboybritches","messages":...}' # Overnight/async work via flex rows curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto:batch","messages":...}' ``` | Profile | What it does | |---|---| | `auto` / `auto:default` | Normal quality-first routing; `-flex` rows excluded | | `auto:batch` | Admits `-flex` rows that may be held during peak hours | | `auto:locality` | Routes only to the `ollama-local` provider (local dispatch) | | `auto:onlycheaps` | Limits to models at or below $0.50 per 1M completion tokens | | `auto:bigboybritches` | Routes only to tier-3 (frontier) models | Profiles are **candidate-set filters** — they narrow which models the router may pick from, but do not change the ranking objective (quality-first, cost as tiebreak). An unknown profile raises HTTP 422. Operators can define custom profiles under `profiles:` in `config/config.yaml`. Each entry accepts `provider`, `min_tier`/`max_tier`, `max_cost_per_1m_completion`, `latency_tolerance`, and `allowed_model_ids` as rewrite rules on top of the default profile. - Ask for any real model id and the router dispatches directly, still logged. - Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines. ### Ask an image question Cloud vision is not universal in the catalog. Requests carrying `image_url` parts are routed only to vision-capable catalog rows. If none survives, the router can fall back to a local vision model instead of returning 422. ```bash curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}' ``` - Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined. - Configure the fallback in `config/config.yaml` under `local_vision:`. ### Force JSON output Use `response_format` when you need structured output. The router treats this as a hard capability requirement and only admits rows that declare `supports_json_mode = 1`. ```bash curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ -d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}' ``` - `NULL` fails closed: an unknown flag means the capability cannot be confirmed. - When streaming is not used, the structural verification verdict appears in the `X-Router-Verification` header. ### Watch it live Three read-only ways to observe the router without spending quota: ```bash # Aggregate JSON health/usage summary curl -s localhost:8080/metrics | python -m json.tool # Live SSE stream of routing decisions curl -s localhost:8080/events/decisions # Terminal dashboard with live routing feed PYTHONPATH=src python -m tui ``` - **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency. - **`GET /events/decisions`** is a Server-Sent Events stream; replays recent decisions, then streams new ones. - **`PYTHONPATH=src python -m tui`** is a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels. ### Compare against trivial baselines `baseline_report.py` replays recent `route_decisions` against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend. ```bash PYTHONPATH=src python baseline_report.py --since 2026-08-01 PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv ``` A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice. ### Admin web portal For the full admin portal walkthrough, see [docs/admin-portal.md](docs/admin-portal.md). ### Probe routing without spending `router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree. ```bash PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects" PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2 ``` - Prints the selected model, candidates, estimated cost, and estimated proficiency. - Use `--category`, `--tier`, and `--context` to override the classifier deterministically. ## At a Glance | Dimension | Detail | |---|---| | **Cost model** | Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog. | | **Latency** | Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response. | | **Quality** | Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior. | | **Verification** | Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via `feedback.py`. | | **Fault tolerance** | Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent `timeout_seconds` silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate. | | **API surface** | OpenAI-compatible `/v1` endpoints, streaming chunk proxy with SSE telemetry scraping. | ## Architecture ``` ┌─────────────────────┐ incoming task ───▶│ Local Classifier │ Ollama (mistral-nemo:12b) │ - task_category │ ~4s warm, ~120s cold cap │ - task_tier │ temperature: 0, max 1024 tokens │ - required_context │ max_retries: 0 (silent 3× cap guard) │ - confidence │ └──────────┬───────────┘ ▼ ┌─────────────────────┐ │ Classifier Fallback │ tier 2 / general_chat on failure │ (graceful degrade) │ — not a 502/503 └──────────┬───────────┘ ▼ ┌─────────────────┐ │ Escalation │ low-confidence tier bump │ (optional) │ threshold: 0.6 confidence └────────┬─────────┘ ▼ ┌─────────────────────┐ │ Hard Filters │ context window ≥ required │ (routing.py) │ tier ≥ required │ │ freshness (active, not stale) │ │ access_level allowed │ │ latency_class compatible └──────────┬───────────┘ ▼ ┌─────────────────────┐ │ Quality-first select │ max proficiency, cheapest │ (routing.py) │ among equals, under a │ │ per-request kWh ceiling └──────────┬───────────┘ ▼ ┌──────────────────────┐ │ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al. │ (FastAPI API) │──▶ OpenAI-compatible /v1 │ │ Streaming SSE + energy scrape └───────┬──────────────┘ ▼ ┌─────────────────────────┐ │ Verification Pipeline │ │ Layer 1: Structural │ ast.parse, json.load, yaml.safe_load │ Layer 2: Local LLM │ async, ~6s, >600 token gate └──────┬──────────────────┘ │ ▼ ┌─────────────────────┐ │ feedback.py │ failures → proficiency updates │ (on-demand agent) │ idempotent, attributed-only └─────────────────────┘ ``` Not pictured: a model that 5xxs is passively excluded from Hard Filters by `circuit_breaker.py` until its cooldown clears, and a verification failure can loop back into Dispatcher for a tier-budgeted retry before falling through to `feedback.py` — see [docs/verification.md](docs/verification.md#retry-budget--iterationpy). For the detailed module-by-module map, see [docs/architecture.md](docs/architecture.md). ## Tech Stack | Layer | Technology | |---|---| | **Language** | Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed) | | **Framework** | FastAPI + uvicorn | | **Database** | SQLite (`router.db`) — decision table, energy observations, proficiency, verifications | | **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN | | **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works | | **Cloud Providers** | OpenAI-compatible carriers via `dispatch_providers:` — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works | | **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) | | **OpenAI Client** | `openai==3.0.0` (official SDK) | | **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) | | **Testing** | `pytest` — 1923 tests across ~90 files, all offline | | **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` | | **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) | | **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works | Dependencies are pinned in `requirements.txt` — recreate the venv with those exact versions to avoid silent drift. **Bump deliberately**: the old `>=` ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use `>=`. ## Testing ```bash python -m pytest # 1923 tests python -m pytest --cov # with coverage ``` No test calls a provider or a local model — the pure modules take rows and config as arguments, so the suite runs offline on a clean checkout. `pyproject.toml` puts the repo root on `sys.path` for pytest — the modules live at the root rather than in a package, so `pytest` (console script) and `python -m pytest` would otherwise disagree about whether `import config` resolves. Representative test files (~90 files total; the list below is a useful subset, not exhaustive): | Test file | What it covers | |---|---| | `test_scoring.py` | `normalize_inverted`, `composite_score`, None-handling | | `test_routing.py` | Hard filters, candidate selection & ranking across all 9 categories × 3 tiers | | `test_tiering.py` | Tier resolver precedence: override → reasoning → cost + context window → mid | | `test_apply_tiering.py` | DB tiering pass, sanity guards | | `test_poller_parsing.py` | Serving class, base model, access level parsing | | `test_proficiency.py` | Blending rule, accumulation | | `test_load_candidates.py` | `load_candidates`: cost/eco/proficiency join | | `test_eval_scoring.py` | Code, exact, tool, and judge scoring | | `test_task_set.py` | Reference solutions validating every `code` task's checks & `exact` answers | | `test_verification.py` | Structural checks on every language, verifier edge cases, local-LLM gating | | `test_feedback.py` | Failure identification, idempotency, attribution filtering | | `test_iteration.py` | Retry budget per tier, and matching the retry to the failure kind | | `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal | | `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split | | `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior | | `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset | | `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction | | `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls | | `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy | | `test_classifier_input.py` | Classifier framing: `_previous_context` scope, `_classifier_user_content` framing | | `test_route_decisions.py` | `route_decisions` table, inline-create helper, config gate | | `test_admin_*.py` | Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family) | | `test_local_dispatch*.py` | Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep | | `test_multi_provider*.py` | Multi-provider dispatch + poller; per-provider balance & attenuation | | `test_circuit_breaker.py` | Passive circuit breaker: cooldown + backoff on model 5xx | | `test_session_cache.py` | Per-session category/tier caching to skip repeat classifier round-trips | | `test_outcome_attribution.py` | Outcome attribution & ambiguity refusal (streaming answer path) | ## Known Limitations & Open Items - **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. See [CLAUDE.md](CLAUDE.md) "What's NOT built yet" for the live list. - **Three models unsettled** — split-half stability varies too much for some model positions. - **Retry does not reach streaming** — corrective attempts work only on the non-streaming path; `POST /outcome` is the streamed answer. - **Local dispatch has its own limits** — no true streaming (answer is buffered), no `verifications` rows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see [docs/routing.md](docs/routing.md)). - **No auth** — the service holds a billable API key with no authentication of its own; loopback is the only guard. - **Local energy is metered when enabled** — `local_energy:` in `config/config.yaml` gates classifier/verifier/vision draw on a separate `local_energy_observations` table (off by default, refuses `enabled: true` without a tariff rate). See [data-model.md](docs/data-model.md). - **Context assembly (RAG)** is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.