Change the unit of session_cache.staleness from minutes to seconds so it can express finer-grained (sub-minute) staleness windows. This is a straight rename, not an additive/compat knob — no deprecated alias, per the project's convention of updating every consumer in the same change. New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was 120 minutes). Default: 1200 seconds (was 20 minutes). The validator's reasoning is unit-independent and carries over: the floor is deliberately > 0 because 0 would make session_cache.get() miss every turn while put() still writes and the classifier-failure cascade's stale_read ignores staleness; the ceiling reasoning (unbounded window = never-expiring cache, 7200s still >> 840s real max run) also carries over in seconds. Every consumer updated in the same commit: - src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX = 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated - src/dispatcher.py: drop the * 60 conversion (field is native seconds) - src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST, _CONFIG_GET_ORDER, _runtime_state, error message template - admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS - config/config.yaml: staleness_seconds: 1200 - tests: test_admin_runtime/config/frontend/knob_coverage, plus stale comment in test_chat_completions - docs: admin-portal.md, evaluation.md, README.md config.local.yaml is gitignored and will be migrated separately.
484 lines
27 KiB
Markdown
484 lines
27 KiB
Markdown
<h1 align="center">
|
||
<img src="assets/6krrt-logo.svg" width="120" alt="6krrt logo"><br>
|
||
6krrt — Local LLM Model Router
|
||
</h1>
|
||
|
||
A router between a coding agent (or any OpenAI-compatible client) and cloud LLMs.
|
||
It classifies each incoming task with a local Ollama model — category, tier,
|
||
required context — and dispatches to the highest-expected-success open-weight
|
||
model on an OpenAI-compatible provider, under a per-request cost ceiling and
|
||
tiebroken by price. Point `model: "auto"` at `POST /v1/chat/completions` for
|
||
automatic selection; pin a real model id and it dispatches as asked, still
|
||
logged.
|
||
|
||
Any OpenAI-compatible carrier should work. **NeuralWatt is the reference
|
||
provider** — it powers every measurement in this README — and uniquely exposes
|
||
per-request **power/energy telemetry** that the router ingests for cost
|
||
accounting; other carriers that report cost only in their usage response are
|
||
supported but carry no energy telemetry. Configure additional providers and
|
||
their API keys under `dispatch_providers:` in `config/config.yaml`.
|
||
|
||
**Numbers in this README are measurements, not specifications.** They come from
|
||
one deployment against one provider account, and the catalog, prices, grid
|
||
intensity and pool load all move. They are here because the reasoning behind a
|
||
design choice is worth more than the choice, and re-running the measurement is
|
||
how you check whether it still holds for you.
|
||
|
||
## Features
|
||
|
||
- **Check routing before spending anything.** `POST /route` classifies, ranks,
|
||
and returns the selected model with no provider call and no cost.
|
||
- **Price per request from the actual shape of the traffic.** Cost is estimated
|
||
from catalog token prices scaled to the prompt size, an assumed completion
|
||
length, and an assumed cache rate — not a single fixed benchmark.
|
||
- **Dispatch selected tasks to a local model.** When the classifier puts a task
|
||
in `file_summarization` or `diff_checking`, the router will consider the local
|
||
model (`qwen2.5-coder-router:14b`) only when it scores within the quality
|
||
tolerance of the cloud leader. (Dormant under the default profile — local
|
||
dispatch does not routinely win.)
|
||
- **Fall back to local vision when no cloud row supports images.** If no
|
||
vision-capable catalog candidate survives the hard filters, the router proxies
|
||
the request to a local Ollama vision model instead of returning 422.
|
||
- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard
|
||
that follows `/events/decisions` as decisions are recorded, with no polling
|
||
delay.
|
||
- **Verify before learning.** Every routed response is structurally parsed in
|
||
the background; larger prose answers get an async local-LLM spot-check, and
|
||
failures fold back into per-model proficiency through `feedback.py`.
|
||
|
||
## Documentation
|
||
|
||
- [Routing internals](docs/routing.md) — quality-first ranking, local vision fallback, candidate ranking
|
||
- [Data model](docs/data-model.md) — SQLite schema: models, proficiency, energy, verifications, route_decisions
|
||
- [Response verification](docs/verification.md) — structural + local-LLM checks, feedback loop
|
||
- [API reference](docs/api.md) — endpoints, virtual models, streaming, logging headers
|
||
- [Client setup](docs/clients.md) — pointing opencode / OpenAI-compatible clients at the router
|
||
- [Architecture](docs/architecture.md) — module-by-module map
|
||
- [Operations](docs/operations.md) — logfmt, journalctl, tracing requests
|
||
- [Incidents](docs/incidents.md) — how the router has broken, with a symptom → one-line-check table
|
||
- [Evaluation & classifier](docs/evaluation.md) — self-eval harness, classifier reliability notes
|
||
- [Admin portal](docs/admin-portal.md) — browser dashboard, controls, decision log
|
||
- [Local model sizing](docs/local-models.md) — fitting Ollama models into 24GB VRAM
|
||
- [Context pruning (pinch)](docs/pinch.md) — pre-dispatch conversation trimming once a session grows past a token budget
|
||
- [Local config overlay](docs/config-local-overlay.md) — per-machine `config.local.yaml` overrides (gitignored)
|
||
|
||
## Requirements
|
||
|
||
- API key(s) for the provider(s) you route to — a NeuralWatt key by default,
|
||
plus any others under `dispatch_providers:` (e.g. an OpenRouter key). The
|
||
reference deployment uses NeuralWatt, whose energy telemetry feeds cost
|
||
accounting.
|
||
- Python 3.10+ (the test suite is verified on 3.10 and 3.14).
|
||
- An Ollama reachable from wherever this runs, with a classifier model pulled.
|
||
- For local dispatch: `qwen2.5-coder-router:14b` created from
|
||
`qwen2.5-coder:14b` with `num_ctx 32768` (see [docs/local-models.md](docs/local-models.md)).
|
||
|
||
Nothing else is assumed about the host — routing itself is SQLite and arithmetic.
|
||
|
||
## Installation
|
||
|
||
### Local (`venv`)
|
||
|
||
```bash
|
||
python -m venv .venv && source .venv/bin/activate
|
||
pip install -r requirements.txt
|
||
sqlite3 router.db < config/schema.sql
|
||
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
|
||
cp config/config.local.yaml.example config/config.local.yaml # optional: deployment-specific overlay (gitignored)
|
||
PYTHONPATH=src python -m poller # populate the catalog
|
||
PYTHONPATH=src python -m tier # resolve tiers
|
||
PYTHONPATH=src python -m config # sanity-check config loads
|
||
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
|
||
```
|
||
|
||
Then edit `config/config.yaml` for your own setup — at minimum:
|
||
|
||
| Key | Why |
|
||
|---|---|
|
||
| `classifier.model` | must match a model `ollama list` reports |
|
||
| `classifier.base_url` | where that Ollama actually is |
|
||
| `objective.plan_kwh_per_period` | your plan's quota; `/health` reports burn against it |
|
||
| `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
|
||
| `session_cache.enabled` | off by default; caches category/tier per session for `staleness_seconds` to skip repeat classifier round-trips on long agent sessions |
|
||
|
||
`PYTHONPATH=src python -m seed_energy` is optional — it populates energy data logged but not used by routing. It costs real money and quota, so it is not in the install path.
|
||
|
||
### As a systemd service
|
||
|
||
Five user services (with companion timers) cover continuous dispatch, catalog
|
||
polling, periodic energy reseeding, and database backups/offsite snapshots. See
|
||
`deploy/README.md` for full instructions.
|
||
|
||
**The poller timer is load-bearing, not optional** — but it fails *silently*,
|
||
not loudly. `mark_stale` runs only inside a poll run that got past the fetch, so
|
||
a stopped timer or a provider outage marks nothing: the catalog freezes at
|
||
last-known-good and the router keeps routing on prices that may be weeks old,
|
||
with every row still reading `active`. Watch for staleness; don't expect it to
|
||
announce itself. (`freshness.stale_after_days` is 3 with `exclude_stale: true`,
|
||
which against a 2-hourly poll is 36 polls of margin.)
|
||
|
||
### Where Ollama lives
|
||
|
||
```bash
|
||
ollama pull mistral-nemo:12b # or whatever you set as classifier.model
|
||
```
|
||
|
||
It does not have to be on the machine running the router; the box with the
|
||
GPU usually isn't the laptop. To use one across a VPN, point **both**
|
||
endpoints at it:
|
||
|
||
```yaml
|
||
classifier:
|
||
base_url: "http://<vpn-ip>:11434/v1"
|
||
verification:
|
||
base_url: "http://<vpn-ip>:11434" # same host, so `model` can stay null
|
||
```
|
||
|
||
and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds
|
||
`127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address
|
||
rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the
|
||
port can run inference and enumerate your models.
|
||
|
||
Both endpoints move together because the verifier speaks Ollama's *native*
|
||
API and cannot follow the classifier to a cloud provider. Config load refuses
|
||
the case where they are on different hosts and `verification.model` is null,
|
||
because that combination fails silently.
|
||
|
||
## Usage
|
||
|
||
### Route without spending anything
|
||
|
||
Use `POST /route` when you want to see what the router would pick for a task before paying for a provider call.
|
||
|
||
```bash
|
||
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
|
||
-d '{"task":"Refactor this 800-line Django view into service objects."}'
|
||
```
|
||
|
||
- Runs the local classifier to determine category, tier, and required context.
|
||
- Applies hard filters and ranks the surviving candidates.
|
||
- Returns the selected model, estimated cost, and estimated proficiency.
|
||
- **No provider call is made and no quota is consumed.**
|
||
|
||
### Skip the classifier when you already know the shape
|
||
|
||
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`, and `required_context_tokens` overrides.
|
||
|
||
```bash
|
||
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
|
||
-d '{"task":"Refactor this 800-line Django view","task_category":"coding_refactor","task_tier":2}'
|
||
```
|
||
|
||
- The classifier is not called, so latency is just the routing pass.
|
||
- `route_decisions.classification_source` is logged as `override`.
|
||
|
||
### Actually dispatch and log energy
|
||
|
||
`POST /dispatch` does the same routing work as `/route` and then calls the selected provider, streams the response if requested, and records the request as an `energy_observations` row.
|
||
|
||
```bash
|
||
curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
|
||
-d '{"task":"What is a Python context manager?"}'
|
||
```
|
||
|
||
- Provider response is proxied back, including streaming chunks.
|
||
- Energy, carbon, cost, and duration are scraped from SSE comments and logged.
|
||
|
||
### Report whether it actually worked
|
||
|
||
Every other quality signal is a proxy — structural checks know code *parses*,
|
||
the local LLM check guesses prose *looks* right. Only the client knows whether
|
||
the answer did the job, because it ran the tests.
|
||
|
||
```bash
|
||
# id comes from the completion body, or any stream chunk
|
||
curl -s localhost:8080/outcome -H 'content-type: application/json' \
|
||
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
|
||
```
|
||
|
||
- The only signal that survives streaming — a retry can't reach bytes already
|
||
sent, but a report arrives afterward and works either way.
|
||
- Successes count too: unlike the automatic checks, `feedback.py` folds a
|
||
reported `ok: true` in as a real success, not just failures.
|
||
- An unknown `request_id` returns `404` rather than being silently accepted.
|
||
|
||
### Point any OpenAI-compatible client at it
|
||
|
||
The `/v1` endpoints speak the OpenAI completions and models API.
|
||
|
||
```bash
|
||
curl -s localhost:8080/v1/models
|
||
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'
|
||
```
|
||
|
||
Ask for a **virtual router model** and the router picks a candidate, subject
|
||
to the profile you specify. The general form is `auto:<profile>`; asking for
|
||
just `auto` is shorthand for `auto:default`.
|
||
|
||
```bash
|
||
# Normal interactive routing (same as "auto")
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto","messages":...}'
|
||
|
||
# Restrict to local dispatch models (ollama-local provider only)
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto:locality","messages":...}'
|
||
|
||
# Restrict to models priced at ≤ $0.50/1M completion tokens
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto:onlycheaps","messages":...}'
|
||
|
||
# Restrict to frontier tier (tier 3)
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto:bigboybritches","messages":...}'
|
||
|
||
# Overnight/async work via flex rows
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto:batch","messages":...}'
|
||
```
|
||
|
||
| Profile | What it does |
|
||
|---|---|
|
||
| `auto` / `auto:default` | Normal quality-first routing; `-flex` rows excluded |
|
||
| `auto:batch` | Admits `-flex` rows that may be held during peak hours |
|
||
| `auto:locality` | Routes only to the `ollama-local` provider (local dispatch) |
|
||
| `auto:onlycheaps` | Limits to models at or below $0.50 per 1M completion tokens |
|
||
| `auto:bigboybritches` | Routes only to tier-3 (frontier) models |
|
||
|
||
Profiles are **candidate-set filters** — they narrow which models the router
|
||
may pick from, but do not change the ranking objective (quality-first, cost as
|
||
tiebreak). An unknown profile raises HTTP 422.
|
||
|
||
Operators can define custom profiles under `profiles:` in
|
||
`config/config.yaml`. Each entry accepts `provider`, `min_tier`/`max_tier`,
|
||
`max_cost_per_1m_completion`, `latency_tolerance`, and `allowed_model_ids`
|
||
as rewrite rules on top of the default profile.
|
||
|
||
- Ask for any real model id and the router dispatches directly, still logged.
|
||
- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines.
|
||
|
||
### Ask an image question
|
||
|
||
Cloud vision is not universal in the catalog. Requests carrying `image_url` parts
|
||
are routed only to vision-capable catalog rows. If none survives, the router can
|
||
fall back to a local vision model instead of returning 422.
|
||
|
||
```bash
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto","messages":[{"role":"user","content":[{"type":"text","text":"Describe this"},{"type":"image_url","image_url":{"url":"data:image/gif;base64,R0lGODlhAQABAAD/ACwAAAAAAQABAAACADs="}}]}]}'
|
||
```
|
||
|
||
- Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined.
|
||
- Configure the fallback in `config/config.yaml` under `local_vision:`.
|
||
|
||
### Force JSON output
|
||
|
||
Use `response_format` when you need structured output. The router treats this as
|
||
a hard capability requirement and only admits rows that declare `supports_json_mode = 1`.
|
||
|
||
```bash
|
||
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
|
||
-d '{"model":"auto","messages":[{"role":"user","content":"Return a JSON object with field answer"}],"response_format":{"type":"json_object"}}'
|
||
```
|
||
|
||
- `NULL` fails closed: an unknown flag means the capability cannot be confirmed.
|
||
- When streaming is not used, the structural verification verdict appears in the `X-Router-Verification` header.
|
||
|
||
### Watch it live
|
||
|
||
Three read-only ways to observe the router without spending quota:
|
||
|
||
```bash
|
||
# Aggregate JSON health/usage summary
|
||
curl -s localhost:8080/metrics | python -m json.tool
|
||
|
||
# Live SSE stream of routing decisions
|
||
curl -s localhost:8080/events/decisions
|
||
|
||
# Terminal dashboard with live routing feed
|
||
PYTHONPATH=src python -m tui
|
||
```
|
||
|
||
- **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency.
|
||
- **`GET /events/decisions`** is a Server-Sent Events stream; replays recent decisions, then streams new ones.
|
||
- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with live routing feed, category breakdown, detail popup, and quota panels.
|
||
|
||
### Compare against trivial baselines
|
||
|
||
`baseline_report.py` replays recent `route_decisions` against two counterfactuals — always cheapest and always best proficiency — using the current catalog and proficiency table. Read-only, no quota spend.
|
||
|
||
```bash
|
||
PYTHONPATH=src python baseline_report.py --since 2026-08-01
|
||
PYTHONPATH=src python baseline_report.py --since 2026-08-01 --csv
|
||
```
|
||
|
||
A high dominance share (selected equals cheapest) with a near-zero proficiency delta means routing's quality-first scorer is not earning its complexity for that slice.
|
||
|
||
### Admin web portal
|
||
|
||
For the full admin portal walkthrough, see [docs/admin-portal.md](docs/admin-portal.md).
|
||
|
||
### Probe routing without spending
|
||
|
||
`router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree.
|
||
|
||
```bash
|
||
PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
|
||
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
|
||
```
|
||
|
||
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
|
||
- Use `--category`, `--tier`, and `--context` to override the classifier deterministically.
|
||
|
||
## At a Glance
|
||
|
||
| Dimension | Detail |
|
||
|---|---|
|
||
| **Cost model** | Per-kWh, not per-token. Flat $8.00/kWh measured across the catalog. |
|
||
| **Latency** | Classification is the floor: ~5-15 s warm on a local reasoning model, up to ~120 s cold. A cloud classifier measured ~1 s. Verification is async and never blocks the response. |
|
||
| **Quality** | Quality is the objective; cost is a per-request kWh ceiling plus a tiebreak. Eco is logged but no longer optimized. Proficiency is an expected pass rate on real traffic, with benchmarks as a prior. |
|
||
| **Verification** | Every response structurally checked (free) + local LLM spot-check on long prose answers (~6 s, never blocks). Failures fold back into proficiency via `feedback.py`. |
|
||
| **Fault tolerance** | Classifier failure degrades to a mid-tier fallback rather than 502/503. SDK retry count is zero, to prevent `timeout_seconds` silently becoming a 3× wall-clock bound. A model that 5xxs trips a passive circuit breaker (cooldown + backoff, no health-check loop); a verification failure buys a tier-budgeted retry on a different candidate. |
|
||
| **API surface** | OpenAI-compatible `/v1` endpoints, streaming chunk proxy with SSE telemetry scraping. |
|
||
|
||
## Architecture
|
||
|
||
```
|
||
┌─────────────────────┐
|
||
incoming task ───▶│ Local Classifier │ Ollama (mistral-nemo:12b)
|
||
│ - task_category │ ~4s warm, ~120s cold cap
|
||
│ - task_tier │ temperature: 0, max 1024 tokens
|
||
│ - required_context │ max_retries: 0 (silent 3× cap guard)
|
||
│ - confidence │
|
||
└──────────┬───────────┘
|
||
▼
|
||
┌─────────────────────┐
|
||
│ Classifier Fallback │ tier 2 / general_chat on failure
|
||
│ (graceful degrade) │ — not a 502/503
|
||
└──────────┬───────────┘
|
||
▼
|
||
┌─────────────────┐
|
||
│ Escalation │ low-confidence tier bump
|
||
│ (optional) │ threshold: 0.6 confidence
|
||
└────────┬─────────┘
|
||
▼
|
||
┌─────────────────────┐
|
||
│ Hard Filters │ context window ≥ required
|
||
│ (routing.py) │ tier ≥ required
|
||
│ │ freshness (active, not stale)
|
||
│ │ access_level allowed
|
||
│ │ latency_class compatible
|
||
└──────────┬───────────┘
|
||
▼
|
||
┌─────────────────────┐
|
||
│ Quality-first select │ max proficiency, cheapest
|
||
│ (routing.py) │ among equals, under a
|
||
│ │ per-request kWh ceiling
|
||
└──────────┬───────────┘
|
||
▼
|
||
┌──────────────────────┐
|
||
│ Dispatcher / Provider │──▶ NeuralWatt / OpenRouter et al.
|
||
│ (FastAPI API) │──▶ OpenAI-compatible /v1
|
||
│ │ Streaming SSE + energy scrape
|
||
└───────┬──────────────┘
|
||
▼
|
||
┌─────────────────────────┐
|
||
│ Verification Pipeline │
|
||
│ Layer 1: Structural │ ast.parse, json.load, yaml.safe_load
|
||
│ Layer 2: Local LLM │ async, ~6s, >600 token gate
|
||
└──────┬──────────────────┘
|
||
│
|
||
▼
|
||
┌─────────────────────┐
|
||
│ feedback.py │ failures → proficiency updates
|
||
│ (on-demand agent) │ idempotent, attributed-only
|
||
└─────────────────────┘
|
||
```
|
||
|
||
Not pictured: a model that 5xxs is passively excluded from Hard Filters by
|
||
`circuit_breaker.py` until its cooldown clears, and a verification failure can
|
||
loop back into Dispatcher for a tier-budgeted retry before falling through to
|
||
`feedback.py` — see [docs/verification.md](docs/verification.md#retry-budget--iterationpy).
|
||
|
||
For the detailed module-by-module map, see [docs/architecture.md](docs/architecture.md).
|
||
|
||
## Tech Stack
|
||
|
||
| Layer | Technology |
|
||
|---|---|
|
||
| **Language** | Python 3.10+ (IO-bound provider APIs; iteration speed matters more than raw speed) |
|
||
| **Framework** | FastAPI + uvicorn |
|
||
| **Database** | SQLite (`router.db`) — decision table, energy observations, proficiency, verifications |
|
||
| **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN |
|
||
| **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works |
|
||
| **Cloud Providers** | OpenAI-compatible carriers via `dispatch_providers:` — NeuralWatt is the reference (per-request power telemetry); any OpenAI-compatible carrier works |
|
||
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
|
||
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
|
||
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
|
||
| **Testing** | `pytest` — 1923 tests across ~90 files, all offline |
|
||
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
|
||
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
|
||
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
|
||
|
||
Dependencies are pinned in `requirements.txt` — recreate the venv with those exact versions to avoid silent drift. **Bump deliberately**: the old `>=` ranges once jumped openai 2.53 → 3.0 and httpx → httpx2 without warning on a fresh venv; a service that restarts on boot shouldn't change its dependency tree underneath itself. Never use `>=`.
|
||
|
||
## Testing
|
||
|
||
```bash
|
||
python -m pytest # 1923 tests
|
||
python -m pytest --cov # with coverage
|
||
```
|
||
|
||
No test calls a provider or a local model — the pure modules take rows and
|
||
config as arguments, so the suite runs offline on a clean checkout.
|
||
|
||
`pyproject.toml` puts the repo root on `sys.path` for pytest — the modules live
|
||
at the root rather than in a package, so `pytest` (console script) and
|
||
`python -m pytest` would otherwise disagree about whether `import config`
|
||
resolves.
|
||
|
||
|
||
Representative test files (~90 files total; the list below is a useful subset, not exhaustive):
|
||
|
||
| Test file | What it covers |
|
||
|---|---|
|
||
| `test_scoring.py` | `normalize_inverted`, `composite_score`, None-handling |
|
||
| `test_routing.py` | Hard filters, candidate selection & ranking across all 9 categories × 3 tiers |
|
||
| `test_tiering.py` | Tier resolver precedence: override → reasoning → cost + context window → mid |
|
||
| `test_apply_tiering.py` | DB tiering pass, sanity guards |
|
||
| `test_poller_parsing.py` | Serving class, base model, access level parsing |
|
||
| `test_proficiency.py` | Blending rule, accumulation |
|
||
| `test_load_candidates.py` | `load_candidates`: cost/eco/proficiency join |
|
||
| `test_eval_scoring.py` | Code, exact, tool, and judge scoring |
|
||
| `test_task_set.py` | Reference solutions validating every `code` task's checks & `exact` answers |
|
||
| `test_verification.py` | Structural checks on every language, verifier edge cases, local-LLM gating |
|
||
| `test_feedback.py` | Failure identification, idempotency, attribution filtering |
|
||
| `test_iteration.py` | Retry budget per tier, and matching the retry to the failure kind |
|
||
| `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal |
|
||
| `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split |
|
||
| `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior |
|
||
| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
|
||
| `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
|
||
| `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
|
||
| `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy |
|
||
| `test_classifier_input.py` | Classifier framing: `_previous_context` scope, `_classifier_user_content` framing |
|
||
| `test_route_decisions.py` | `route_decisions` table, inline-create helper, config gate |
|
||
| `test_admin_*.py` | Admin portal: config/RBAC-style guards, overrides, providers, profiles, allowlist, triggers (a large family) |
|
||
| `test_local_dispatch*.py` | Local (Ollama) dispatch: fallback routing, local encoder/energy, seed sweep |
|
||
| `test_multi_provider*.py` | Multi-provider dispatch + poller; per-provider balance & attenuation |
|
||
| `test_circuit_breaker.py` | Passive circuit breaker: cooldown + backoff on model 5xx |
|
||
| `test_session_cache.py` | Per-session category/tier caching to skip repeat classifier round-trips |
|
||
| `test_outcome_attribution.py` | Outcome attribution & ambiguity refusal (streaming answer path) |
|
||
|
||
## Known Limitations & Open Items
|
||
|
||
- **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. See [CLAUDE.md](CLAUDE.md) "What's NOT built yet" for the live list.
|
||
- **Three models unsettled** — split-half stability varies too much for some model positions.
|
||
- **Retry does not reach streaming** — corrective attempts work only on the non-streaming path; `POST /outcome` is the streamed answer.
|
||
- **Local dispatch has its own limits** — no true streaming (answer is buffered), no `verifications` rows for local answers, and follow-ups are not coalesced. Routed requests in eligible categories now degrade to the local model when the cloud is unavailable (a fallback, not a preference — see [docs/routing.md](docs/routing.md)).
|
||
- **No auth** — the service holds a billable API key with no authentication of its own; loopback is the only guard.
|
||
- **Local energy is metered when enabled** — `local_energy:` in `config/config.yaml` gates classifier/verifier/vision draw on a separate `local_energy_observations` table (off by default, refuses `enabled: true` without a tariff rate). See [data-model.md](docs/data-model.md).
|
||
- **Context assembly (RAG)** is out of scope — the classifier sees the full conversation but does not perform document/code retrieval.
|