Files
6krrt/CLAUDE.md
adlee-was-taken 8b861f9267 docs: document local vision fallback and the monitoring stack
README gains the Local Vision fallback section (the Ollama/qwen3-vl path
for image requests with no cloud vision candidate — previously only the
cloud capability gates were documented) plus route_decisions table,
metrics.py/tui.py/router_cli.py in the module list, GET /metrics endpoint
and keys, and monitoring usage snippets.

CLAUDE gains metrics.py/tui.py/router_cli.py in what's-built and a
Monitoring: route_decisions persistence section.

Also corrects verifications.kind/verdict to match schema.sql.
2026-08-23 21:00:51 -04:00

59 KiB
Raw Permalink Blame History

Local LLM Model Router — project brief

README.md documents what is actually built, module by module. design/local-llm-model-router.md holds the architecture and rationale, including parts still unbuilt. This file is the working state + immediate next steps, and is the one to trust on what is currently true.

What this is

A router that uses a local model (served via Ollama) to classify incoming coding/documentation tasks — category, tier, required context size — and dispatch each task to the cheapest/best-fit open-weight model on Neuralwatt Cloud, weighted by cost, per-category proficiency, and a per-request energy ceiling.

Every measurement in this file was taken on one deployment against one provider account. They are recorded because the reasoning is worth more than the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it.

Neuralwatt is the only provider. OpenRouter was removed — the provider column and the (model_id, provider) primary key stay so a second provider can be added later without a migration.

Stack

  • Python (chosen over Rust — this is I/O-bound against provider APIs, not CPU-bound; iteration speed on the scoring/weighting logic matters more than raw execution speed at this scale)
  • SQLite for the decision table
  • Ollama for local classification, via any OpenAI-compatible endpoint — localhost:11434/v1, or an Ollama on another machine across a VPN (classifier.base_url)
  • FastAPI for the dispatcher service

Billing is per-kWh, not per-token — and neither is what scoring uses

Measured against the live API, 2026-08-11. Neuralwatt bills a flat $8.00 per kWh and the catalog's input_per_million / output_per_million prices are not what this account is charged. Confirmed across five models; cost_usd / energy_kwh came back 8.00 every time:

model list $/1M out completion tokens billed USD kWh $/kWh
deepseek-v4-flash 0.28 600 5.00e-06 5.88e-07 8.50*
gemma-4-31b 0.42 540 3.80e-04 4.75e-05 8.00
qwen3.6-35b-fast 1.15 600 2.53e-04 3.16e-05 8.01
kimi-k2.7-code-fast 4.00 600 2.25e-03 2.81e-04 8.00
kimi-k3-fast 15.00 539 2.17e-04 2.71e-05 8.00

* rounding — billed cost is quantized to ~1e-06.

The precise rule, validated against all 65 samples of the reference sweep (61/65 within 2%; the 4 outliers are microdollar rounding, not misses):

cost_usd = min( $8.00/kWh x energy_kwh ,  3 x list token price )

The ceiling bound in only 2 of 65 samples, both deepseek-v4-flash energy spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x $0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349.

List price ranks models backwards. Not approximately — invertedly:

list $/1M actually billed gCO2eq
deepseek-v4-flash 0.28 9.80e-05 8.90e-04
gemma-4-31b 0.42 5.02e-05 2.32e-04

deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon.

But cost and eco are NOT the same axis — the tempting simplification, and it is wrong. Cost tracks energy, but carbon is energy x the serving region's grid intensity, and models run in different regions:

grid gCO2/kWh models
FI ~49-50 most of the catalog; varies by time
FI (reported) 475 glm-5.2-fast, glm-5.2-flex
US-MIDA-PJM ~442 the kimi-k3 family

A 13.6x spread, so the two axes disagree: glm-5.2-fast is the 2nd cheapest model and only the 6th cleanest; kimi-k3-flex draws 3.7x less energy than kimi-k2.7-code while emitting 3.6x more carbon. Weighting them separately is load-bearing, and tests/test_routing.py pins it.

Superseded — cost no longer comes from the sweep at all. cost was the median measured USD over the reference sweep. That was measured to be WRONG for real traffic, because the reference workload is the wrong shape.

The sweep sends a 400-token prompt with a 400-token completion. Real agent traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts:

workload winner
reference sweep (400/400) glm-5.2-fast, 3.2x cheaper
realistic (70k prompt, short answer) deepseek-v4-flash, 5.0x cheaper

Same two models, opposite answer. glm-5.2-fast sits at attribution 0.006 on a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts and badly on real ones. deepseek-v4-flash barely moves (0.21 -> 0.25).

So routing.estimated_cost prices each request from catalog token prices, scaled to that request's actual shape (prompt size, assumed completion length, assumed_cache_rate). List price is not what gets billed, but billing is capped at 3x list, so it tracks the real ordering and bounds it — and on the one case that was checked live it agrees with the measurement in direction and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no sweep, and refreshes whenever the poller runs.

Three signals said deepseek-v4-flash — catalog token price (7.8x cheaper), NeuralWatt's own published per-request energy (~10x lower), and a live 70k measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the workload you actually run.

eco still comes from the sweep's median gCO2eq, and is still not an objective. flex_cost_multiplier is gone: a flex row's measured cost already is its flex cost.

Open, and worth knowing: NeuralWatt's model cards publish gross energy (~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is gross x attribution. GLM burns roughly 7x more actual electricity per request and charges ~5x less, because far more tenants share its GPUs. Anything built on eco inherits that inversion — the attributed carbon figure answers "what is my share", not "what was burned".

Energy attribution: signal that looks like noise

Billed energy decomposes exactly:

energy_kwh = avg_power_watts x duration_seconds x attribution_ratio

attribution_ratio is the request's share of a shared multi-tenant GPU pool. Up close it looks like pure noise — eight rapid identical calls to one model spanned 20x in billed energy, correlating +0.997 with the ratio while power and duration held steady. Two sweeps of the same 13 models with the same prompt disagreed by up to 36x.

Scoring on the pre-attribution product (power x duration) was tried, and it is wrong. Across the sweep:

spread
median attribution, between models 750x
typical spread within one model 1.8x

The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving concurrency, a stable per-model property, not weather. A model whose GPUs carry far more concurrent requests genuinely costs less per request, and that is most of the real cost difference in the catalog: deepseek-v4-flash bills ~1000x under its share of pool gross. Stripping attribution discards a 750x real signal to suppress a 1.8x one.

So scoring reads the attributed figures, and the median absorbs what noise remains. A split-half check on the 7-sample sweep (median of first three vs last four) shows that working:

  • 10 of 13 models agree within 1.4x — stable enough to route on
  • 3 do not: kimi-k2.7-code-fast (29x), kimi-k3 (14x), glm-5.2-flex (2.2x). Those need more samples before their position is trustworthy.

dispatcher.gross_energy_kwh remains as a diagnostic on the identity, not a scoring input.

Attribution drifts across hours, so sampling must too

Within about 30 minutes the billed figures reproduce (0.3-1.1x on a spot-check). Across hours they do not: between two sweeps, deepseek-v4-flash moved roughly 50x and qwen3.6-35b about 7x the other way — enough to invert their cost ranking. Attribution tracks pool load, and pool load tracks time of day.

More samples inside one sweep does not fix this; it measures one moment more precisely. Coverage across time does. load_candidates already takes the median over ALL seed_reference rows, so repeated sweeps accumulate into a median-across-time for free — hence llm-router-seed.timer, which runs a small sweep every 6 hours.

Until several sweeps have accumulated, treat the eco ordering as provisional. A single sweep's ranking is one sample of a moving quantity.

And none have accumulated since 6e729ad. That commit moved log_observation's trailing arguments to keyword-only without updating seed_energy.py, so every timer run since spent one billed completion and then died on TypeError — which is not a RequestException, so the per-sample except did not catch it. Fixed, and the sweep now has an offline end-to-end test, but the accumulation this section describes starts from the next run rather than from months of history.

What's built and working

  • schema.sqlmodels, proficiency, energy_observations. Applies cleanly (sqlite3 router.db < schema.sql). models carries the serving class columns (below); energy_observations carries real carbon/cost.
  • poller.py — fetches Neuralwatt's /models endpoint (public, unauthenticated), normalizes, upserts, marks stale rows. Verified against the live API: 19 models, and the metadata.pricing / metadata.capabilities / metadata.limits field mappings are confirmed correct.
  • config.yaml / config.py — weights, thresholds, provider settings, Pydantic-validated.
  • scoring.py — one normalize_inverted (cost and eco normalize identically; they differ only in what is fed to them) plus the weighted composite.
  • seed_energy.py — runs a fixed reference task N times per routable model and writes energy_observations rows tagged seed_reference. This is what makes cost and eco real numbers instead of the neutral 0.5. Re-run it after the catalog gains models: python seed_energy.py --samples 5 (13 models x 5 = 65 calls, and the whole sweep cost under a cent).
  • tiering.py / tier.py — pure tier resolver + the DB pass that applies it.
  • routing.py — pure hard filters and ranking.
  • dispatcher.py — FastAPI service. GET /health, POST /route (classify and pick, no provider call), POST /dispatch (route, call, log), plus an OpenAI-compatible GET /v1/models and POST /v1/chat/completions.
  • proficiency.py / proficiency_store.py — pure blending arithmetic, and the only path that writes the proficiency table (so blended_score and source can never drift from the inputs that produced them).
  • leaderboards.yaml / leaderboard.py — curated per-family priors and their importer. Ships empty by design; python leaderboard.py --check names every active family lacking a prior.
  • evals/tasks.yaml / eval_proficiency.py — the self-eval task set and runner. Four scoring kinds, objective wherever the category allows it: code executes the model's Python against checks, exact compares a normalized answer, tool inspects the tool call structurally, and only the four prose categories fall back to a judge.
  • capabilities.py — request-side capability detection. detect_capabilities reads an OpenAI-format body and returns RequestCapabilities (has_images, require_json_mode, tools_present, has_reasoning_request). It scans all messages for image_url parts and reads response_format.type against {json_object, json_schema}. Detection is read from the request body, not inferred by a classifier.
  • logs.py — the service's structured logging: a per-request trace id in a ContextVar, logfmt rendering, and journald priority prefixes emitted only when systemd actually owns stderr. logs.bind() exists because a StreamingResponse's generator is resumed in a fresh copy of the caller's context, so the ContextVar cannot reach it — the path all agent traffic takes. Level from logging.level, overridden by LLM_ROUTER_LOG_LEVEL.
  • metrics.py / GET /metrics — read-only observability aggregation. Keeps the dashboard helpers outside dispatcher.py to avoid an import cycle: dispatcher imports metrics for /health and /metrics, so metrics takes (conn, cfg) arguments and never imports dispatcher. Returns quota burn against plan_kwh_per_period, scoring coverage warnings, recent route_decisions, per-model aggregates over energy_observations, verification verdict mix, and top proficiency by category.
  • tui.py — Textual terminal dashboard over GET /metrics. A foreground entrypoint (python tui.py), not a service. textual is imported only here, so the router's dispatch path has no UI dependency.
  • router_cli.py — one-shot /route probe. Posts a task to the running router and prints the decision tree, or emits raw JSON with --json. Spends no quota because it only routes.
  • tests/ — 356 tests across 19 files, all passing, all offline. Verified on Python 3.10 and 3.14; nothing declares requires-python, so 3.10 is the tested floor rather than a promised one.

Monitoring: route_decisions persistence

route_decisions is the newest observability table (not a scoring input). It records one row per routing decision — route | dispatch | chat | passthrough | local_vision — with category, tier, selected model, runners up, estimated cost/proficiency, classifier latency, rejection reason, and request feature flags (tools, images, json_mode, streamed). It stores only a hashed session_key; session_dir, prompts, and answers are excluded, and a test enforces that the write path never stores conversation text.

The write is gated by logging.log_route_decisions and is best-effort: a failed write is logged and swallowed, because a decision record is worth having but never worth failing or slowing a request for. The table is created from code at module load (_ensure_route_decisions_table) and again on every write (ensure_route_decisions inside persist_route_decision), mirroring the proficiency_store.ensure_columns migration pattern: schema.sql is CREATE TABLE IF NOT EXISTS, but a live router.db predating this table needs the code-side migration. Both the module-load hook and the write-path guarantee are idempotent and leave existing rows intact.

Request-side capability gates

The router treats some capabilities as hard filters, read directly from the request body. Two new gates live in routing.rejection_reason and routing.select_candidates:

  • require_vision — active when the request carries image_url parts and routing.require_vision is true. A model passes only if its catalog row says supports_vision = 1. supports_vision = NULL fails closed: an unknown flag means the capability cannot be confirmed, and routing an image request to a model that might lack vision is a guaranteed provider 400.
  • require_json_mode — active when response_format.type is json_object or json_schema and routing.require_json_mode is true. A model passes only if supports_json_mode = 1; NULL also fails closed, for the same reason.

Both gates default to on in config.yaml, because a wrong guess produces a 400. This is a deliberate asymmetry against the tool-proficiency gate below: capability flags fail closed on unknown, while quality measurements (tool proficiency, energy) admit on absent evidence ("unproven, not bad").

tools_present remains a measurement gate, not a flag gate. Every routable catalog row has supports_tools = 1, so that flag would be inert. What matters is the measured tool_use_agentic proficiency. A model with no measured tool score is admitted; a measured score below routing.min_tool_proficiency is dropped — but only when the request carries a tools array. This gated filter ships disabled (null).

has_reasoning_request is detected in capabilities.py but is informational only. Reasoning stays on tiering: reasoning_default_enabled decides the tier, not supports_reasoning (which only means "the endpoint accepts a reasoning param" and is true for 17 of 19 rows). There is no reasoning capability gate.

Pass-through capability check

When a client pins a real model id in /v1/chat/completions, the router dispatches as asked. Before spending a provider call, _check_pinned_capabilities reads the model's supports_vision and supports_json_mode flags and returns a clear 422 if the pin cannot satisfy the request. The pin could never have worked, so failing early is better than an opaque provider 400.

Local Ollama vision fallback

local_vision: in config.yaml configures a fallback path for image requests that find no cloud vision candidate — the cloud catalog excludes the cost leader (deepseek) on vision, so without this fallback every image request that would otherwise have routed there 422s instead. It is a core feature, enabled by default (enabled: true, both in config.yaml and in LocalVisionConfig's own default, so a config that omits the section still gets it). Disable it explicitly (enabled: false) on a host with no local Ollama, or one that hasn't pulled the vision model.

When enabled and routing returns no selected model, _run_local_vision sends the original messages, with image_url parts intact, to a local Ollama model (qwen3-vl:4b by default). The local answer then replaces the completion: _local_vision_response returns a normal OpenAI-shaped response, including a stream-wrapped version for stream: true. It does not inject a caption into a cloud call, because the streaming proxy cannot rewrite bytes mid-stream.

Budget guards refuse images too numerous or too large before the local call is made. If the local call fails for any reason, it falls through to the ordinary 422 No model satisfies the hard filters so the failure is visible rather than a silent empty response.

Serving class: one base model, many rows

Neuralwatt ships ~6 base models as 19 catalog rows. The id suffixes are three orthogonal dimensions (hence glm-5.2-short-fast-flex), parsed by poller.parse_serving_class into columns:

suffix column meaning
-flex latency_class discounted async; held server-side during peak until a capacity gap opens
-fast reasoning_mode thinking disabled or capped to a short budget — not queue priority
-short context_variant 200K pool with a bounded reasoning budget

These rows carry identical catalog pricing, so without these columns all 7 GLM rows tie exactly and the router picks arbitrarily — which could send an interactive request to a preemptible flex row. latency_tolerance (interactive | batch) is therefore a hard filter in routing.py, not a weight. Suffixes are matched as whole --delimited segments so deepseek-v4-flash is not misread as a -fast row.

Access gating is prose-only

6 of 19 rows are restricted, and the catalog says so only in free text ("Private preview (grant-gated)", "(Canary)") — there is no structured field. poller.parse_access_level parses it into access_level, and routing.allowed_access_levels (default [public]) excludes them, so the dispatcher doesn't select a model that 403s. 13 of 19 rows are routable.

Tiering

Tier on reasoning_default_enabled (from metadata.reasoning.default_enabled, falling back to capabilities.reasoning), not supports_reasoning. capabilities.reasoning only means "the endpoint accepts a reasoning param" — it is true for 17 of 19 rows, and tiering on it put 17 models in tier 3 and left tier 1 empty. A -fast row does not inherit its sibling's tier 3. Cost is checked before the reasoning rule so $0.28/1M models can reach tier 1.

Cheapness is not a capability ceiling (tier1_context_max, default 512000). Tier is a floorrouting.py drops any row with tier < required_tier — so tier 1 means "simple work only", not "cheap". Deciding that on price alone put deepseek-v4-flash in tier 1 for no reason but its $0.28/1M completion price, which excluded it outright from every tier-2 request. It has a 1M advertised window and scores 1.00 on all three coding categories. That was the same substitution the cost axis already had to unlearn: price is a market signal, not a capability measurement.

Tier 1 now requires the model to be small and cheap. The gate reads the advertised context_window, whose catalog values are the clean market classes — 131056 / 199984 / 262128 / 1048560 — rather than effective_context_window, which varies within a class. 512000 sits in the empty band between the 256K and 1M classes with a 2x margin either side, so it is not fitted to any one model. The gate only ever demotes; a huge window never promotes an expensive model into tier 1, and a missing window does not block it, since absent evidence should not decide anything.

Distribution moved 4 / 6 / 9 -> 1 / 9 / 9 — only the three deepseek rows changed. With cost priced per-request, deepseek-v4-flash now wins coding_general at every context size (16.5x cheaper than kimi-k3 at 200k) and is still correctly absent from tool_use_agentic, where its measured 0.33 drops it out of the quality band. That is the eval data earning it the slot rather than a thumb on the scale — no model_tiers override was needed.

Why tools and reasoning stay on their existing signals

Tools stay on the measured tool_use_agentic proficiency gate, not a supports_tools flag gate. Every routable catalog row already has supports_tools = 1, so a flag gate would be inert. The real signal is the measured proficiency, because the observed failure is a model over-reaching for tools on a non-agentic prompt. routing.min_tool_proficiency captures that measurement and only applies when the request carries a tools array.

Reasoning stays on tiering (reasoning_default_enabled), not on a new flag gate. supports_reasoning only means the endpoint accepts a reasoning parameter, and that is true for 17 of 19 rows — nearly the whole catalog. has_reasoning_request is detected purely for observation. Making it a gate would add no useful filtering, because the decision of whether a request needs reasoning is already encoded in the requested tier.

The fail-closed asymmetry, stated plainly: capability flags fail closed on unknown; quality measurements admit on absent evidence. A missing supports_vision or supports_json_mode flag means "cannot confirm", so the model is dropped. A missing tool_use_agentic score or energy measurement means "unproven, not bad", so the model is admitted. The first wrong guess is a guaranteed 400; the second is just an empty data point that the neutral default handles.

Proficiency: category now changes routing

proficiency_score is the ONLY category-dependent term in the ranking, so until this table had data, task_category could not change a decision at all — the classifier computed it, the router paid ~10s for it, and then it made no difference. It does now: sweeping 9 categories x 3 tiers currently returns 5 distinct winners at both 50k and 120k of context.

context winners over 27 decisions
50k qwen3.6-35b (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2)
120k kimi-k2.7-code (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2)

This spread is recent, and how it got here is the useful part. For a long time all 27 decisions returned ONE model, and that was the correct answer at the time rather than a bug: with cost and eco both populated, qwen3.6-35b was Pareto-dominant — cheapest AND cleanest in the routable set, while scoring within quality_tolerance of the best. No defensible weighting picks anything else out of that.

Two corrections widened it, and neither was a tuning change:

  • Cost stopped being a benchmark average. It is now priced per request from catalog prices scaled to the request's shape, so the ranking depends on the workload instead of on a 400-token reference sweep that no real traffic resembles.
  • Tier stopped being inferred from price. deepseek-v4-flash was pinned to tier 1 for being cheap, which excluded it from every tier-2 request regardless of what any score said.

Note what changes between the two rows above: only the leader, and only because of the hard context filter. That is the filter working, not the scoring disagreeing with itself.

If you see one model win everything again, check for dominance before reaching for config. One winner is a legitimate outcome. The lever, if a genuinely different balance is wanted, is objective.quality_tolerance (how large a quality gap must be before it outranks a cost saving) or objective.max_energy_per_request (a hard ceiling). There is no weight to tune — quality is the objective and cost is the tiebreak.

What the task set actually found

The benchmark could not discriminate these models on coding. Every row scored exactly 1.00 on coding_general, coding_refactor and debugging — and that is after the tasks were deliberately hardened with touching intervals, full semver, present-but-falsy defaults, late-binding closures and a binary search that infinite-loops. Every model in this catalog is simply good at that class of problem, so cost decides coding routes, which is the right outcome.

Real traffic broke one of those ties, which the benchmark never could. coding_general now spans 0.86-1.00: glm-5.2-fast fell to 0.862 over 29 samples folded in by feedback.py from an actual agent session, and crossed self_eval_min_samples on the way, so it reads self_eval rather than self_eval_thin. That is the intended shape of this system — the 23-task benchmark establishes a floor, and your own traffic is what refines it. coding_refactor and debugging are still flat at 1.00, awaiting the same treatment.

A 1.00 can also be a sampling artifact, and docs_writing was one. At 2 samples per model the category read 0.70-1.00 with a model at the ceiling, and the router paid for that ceiling: kimi-k3-fast won every docs route. Six more benchmark passes moved every score and left NOTHING at 1.00:

model n=2 n=11-14
kimi-k3 0.85 0.973
kimi-k2.7-code 0.85 0.886
deepseek-v4-flash 0.80 0.864
kimi-k3-fast 1.00 0.864
qwen3.6-35b 0.85 0.800
gemma-4-31b 0.85 0.786

The winner moved to kimi-k2.7-code, 3.2x cheaper at 50k of context ($0.0441 -> $0.0136), with no config change — kimi-k3 scores higher but sits inside quality_tolerance, so cost breaks the tie. deepseek-v4-flash ($0.0024) misses the band by 0.009, which is the kind of margin the tolerance exists to describe rather than a verdict.

The whole spread rests on one rubric line, though. docs_function is effectively saturated — 1.00 on nine of every ten samples — and nearly every docs_gotcha deduction is the same omission: the model documents that order is preserved, that the first occurrence is kept, and what key does, then never says elements must be hashable. That is real discrimination, since it is a real property of the function, but one sentence is deciding a category. Treat this ordering as thinner than n=14 makes it look.

The self-judging guard costs sample density, and it shows up here. Most models reached n=14; kimi-k3 and kimi-k3-fast reached only 11, because those two are the ones diverted to the alternate judge qwen3.6-35b, which returns unparseable JSON more often than kimi-k3 does. The guard is still right — a model grading its own family is worse than a thinner sample — but the alternate judges should be picked for parseability, not just for being someone else.

The reflex when a category looks flat is to reach for quality_tolerance. Neither tie broken so far was broken that way: coding_general opened up when feedback.py folded in real traffic, and docs_writing opened up on six more benchmark passes. Both were samples, not settings. coding_refactor and debugging are still flat at 1.00 on 2-3 samples each — which is now a state this project has mistaken for a measurement once.

Current spread by category, widest first:

category spread
tool_use_agentic 0.33 - 1.00
summarization 0.60 - 1.00
reasoning_math 0.67 - 1.00
docs_writing 0.66 - 0.97
general_chat 0.80 - 1.00
translation 0.85 - 1.00
coding_general 0.86 - 1.00
coding_refactor, debugging flat at 1.00

What does discriminate is tool use, arithmetic traps, and prose. deepseek-v4-flash scores 1.00 on all three coding categories yet 0.33 on tool_use_agentic and 0.67 on reasoning_math. Verified live, not an artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes away?" — both times supplied — it calls two tools rather than subtracting. It over-reaches for tools, which is exactly the failure mode that matters in an agent loop. The router now avoids it for those categories while still picking it for coding.

Most rows still read source='self_eval_thin' (118 of 132): real measurement, but below self_eval_min_samples at 2-3 tasks per category per run. The 14 that have crossed it are all docs_writing, from the six extra passes above. Two paths thicken it, and they are complementary — re-run eval_proficiency.py to accumulate benchmark samples, or just use the router and let feedback.py fold in real outcomes. Both fold into a running mean rather than replacing, so samples add up across runs.

A score is only as fresh as the row it was copied to

Proficiency is a property of the weights, not the queue, so the eval harness scores one row per family and propagate_to_variants copies the result onto the serving variants — kimi-k3-flex gets kimi-k3's number, because no benchmark rates a -flex row separately.

That copy used to happen exactly once per variant, ever. The guard skipped any row with self_eval_samples > 0, meaning "measured directly, do not overwrite" — but inheritance copies the sample count too, so after the first propagation an inherited row was indistinguishable from a measured one and was never refreshed again. kimi-k3-flex sat at 0.85/n=2 while kimi-k3 moved to 0.973/n=11.

proficiency.inherited_from records the provenance that was missing, and the migration was the delicate half, not the fix: ADD COLUMN gives every existing row NULL, which reads as "measured here", so shipping the guard alone would have permanently frozen the exact rows it exists to unfreeze. The backfill infers provenance from the harness's own selection rule rather than guessing — eval_identities only ever evaluates standard rows plus flex rows with no standard equivalent, so a flex row that has one was never a candidate for direct evaluation, whatever its sample count claims. Everything else keeps NULL, which fails safe: NULL means "do not overwrite", so no real measurement can be lost to a wrong guess.

Confirmed on the live database, and on the catalog's one genuine exception — glm-5.2 is canary, so glm-5.2-flex is the routable row the harness scores directly, and its NULL is correct.

Harness bugs this shook out

Three separate defects, each of which scored the rig rather than the model, and each caught by reading per-task detail rather than the summary:

  • Token budget. max_tokens was shared between a reasoning model's trace and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and returned an EMPTY content field, scoring 0.00 on tasks it can plainly do. Now 24000, clamped per model (gemma-4-31b caps at 16384), and finish_reason: length skips the sample instead of scoring it.
  • One leading space. kimi-k2.7-code returns " def f(...)", which becomes IndentationError once the harness prepends its imports — 0.00 across all nine coding tasks for a model with "code" in its name.
  • Judge failures scored as model failures. 44% of judge calls returned unparseable output (the judge is itself a reasoning model and leaks its thinking despite response_format). Each was recorded as 0.0. Now the JSON is extracted from surrounding prose and an unusable reply yields no sample.

tests/test_task_set.py exists so this stops happening: it implements a reference solution for every code task and asserts it passes every check, recomputes every exact answer (one by brute force), and confirms each refactor target already passes its own checks while each debugging target fails. It immediately caught a check where the expected value was simply wrong — which would have docked every model on a task and been indistinguishable from genuine difficulty.

Tool competence is read from the request, not guessed at

Neither local classifier can identify agentic work. Asked to label six unambiguous tool-use prompts ("read the config then update the manifest", "run the tests and fix what fails"), qwen3.5 got 2/6 and mistral-nemo 1/6 — and mistral-nemo's misses collapse to general_chat, which is also the configured fallback_category, so qwen3.5's crashes land in the same place.

That mattered because tool_use_agentic has the widest proficiency spread in the table (0.33-1.00) and deepseek-v4-flash — the current winner on coding — sits at the bottom of it.

The fix was not a better classifier. Whether tools are on the table is stated in the request: every agent client sends a tools array, and chat_completions never looked at it. Reading it is exact and free.

It is applied as a hard filter, not a category override, and the distinction is load-bearing. The question is not "is this task agentic" but "can this model be trusted with tools that exist". The recorded failure is precisely the second one: deepseek-v4-flash was given a non-agentic prompt ("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times supplied) and called two tools rather than subtracting. A model that over-reaches is a hazard on every request where tools are available, whatever a classifier would have labelled the task.

So routing.min_tool_proficiency drops any candidate whose measured tool_use_agentic score is below it, but only when the request carries tools:

request winner on coding_general @ 50k
no tools deepseek-v4-flash ($0.0024)
tools present qwen3.6-35b ($0.0041)

Verified live through /v1/chat/completions with identical bodies differing only by the tools array. The cost of safety here is 1.7x on that route, paid only where tools exist.

It is currently set to null, i.e. OFF, deliberately and pending experiment. opencode sends tools on essentially every request, so with the filter on, deepseek-v4-flash is excluded from ordinary agent traffic and its ~7x cost advantage goes unused; with it off, that advantage applies and a model measured at 0.33 on tool use handles requests where tools are on the table. Which is right is an empirical question and the benchmark cannot answer it — the 0.33 comes from 3 tasks.

What settles it is POST /outcome: run with the filter off, let real pass/fail reports accumulate, and compare deepseek-v4-flash's tool_use_agentic proficiency before and after. That is the one signal here that knows whether the work actually worked, and feedback.py folds client outcomes in both directions, so success counts too.

0.5 sits in the empty band between the only two values the catalog holds (0.33 and 1.00), so it is not fitted to either. A model with no measured tool score is unproven rather than proven bad and is not dropped — the same rule as the tier-1 context gate. Config load refuses a routing.tool_use_category that is not a real category, because a name matching nothing yields NULL for every row and NULL means "do not disqualify": the filter would silently stop filtering.

Tier is an iteration budget, not just a floor

A tier used to mean only "do not route below this". It now also buys corrective attempts after a verification failure:

tier batch interactive
1 0 retries 0
2 1 1
3 2 1

Interactive is capped below its tier because every retry doubles time-to-answer, and in interactive use latency is a quality loss.

Retries are matched to the failure, since the causes differ:

  • truncated — raise the token budget on the same model; a different one would run out too. If there is no cap to raise, the model's own output ceiling is the wall, so escalate to a candidate that can emit more.
  • malformed — more tokens will not make unparseable output parse, so escalate to the next-ranked candidate.
  • ok / unverifiable — buy nothing. Retrying unverifiable would burn quota across the majority of prose traffic for no signal.

escalation.preemptive_on_low_confidence is now off by default. Bumping the tier because the classifier was unsure pays frontier prices before anything has gone wrong; spending after a check has actually failed is better on both mandates — the cheap attempt usually succeeds, and when it fails you have evidence rather than a hunch.

The only ground truth: POST /outcome

Everything else the router records is a proxy. Structural checks know whether code parses. The local checker guesses whether prose looks right. Neither knows whether the answer did the job — the client does, because it ran the tests.

# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'

Two things make this the highest-value signal available:

  • It is the only quality signal that survives streaming. A retry cannot reach a streamed response because the bytes are already gone; a report arrives afterwards and works either way. Every agent client streams.
  • Its successes count. feedback.py folds client outcomes in BOTH directions, unlike checks, where only failures count. A parser's 'ok' means the code parsed and would inflate scores; a client's 'succeeded' means the work worked.

An unknown request_id returns 404 rather than being quietly accepted — a client whose reports go nowhere should find out.

Verification: what local compute is actually good for

Local inference is a poor substitute for cloud completions here — 7-82x the energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s). But it is very good at stopping a cloud completion from being wasted, and completions are where all the money is: fitted on real traffic, a completion token costs 201x a prompt token.

check cost what it catches
structural (verification.py) free truncation, malformed code/JSON/YAML, empty answers
local LLM (Ollama) ~$6.9e-05, ~6s refusals, wrong-question answers, incoherence

Structural checks run on every response and never execute the code — they parse it. The local LLM check runs only on answers above verification.min_completion_tokens, in the background after the client has its response, because it costs ~15% of a median 193-token answer and only pays above ~600 tokens.

An agent turn is not a prose answer

Both checkers had mirror halves of the same blind spot, and real traffic is what found it. On the first genuine agent session (63 completions, a shipped feature, 349 passing tests, clean mypy and ruff):

checker called it actually
structural malformed: empty response x29 turns that ended in a tool call
local LLM cuts off mid-sentence x8 of 9 same turns, judged from the other side

A turn that calls a tool has empty or half-finished text by design. Both paths now take has_tool_calls and return unverifiable; in verify_response that check outranks even finish_reason == 'length', because stopping mid-sentence at a call boundary is not a budget overrun. worth_local_check declines outright, which also stops paying ~6s of local inference to mis-grade a tool call.

Had feedback.py run before this, it would have applied ~12 false failures to the models that had just shipped the feature. That is the fourth harness bug in this project that would have scored the rig rather than the model, and the first caught by real traffic instead of a synthetic test. The pre-fix rows are kept but set model_attributable = 0, so the record survives without steering routing.

client_capped was also over-applied. It marked every verdict non-attributable whenever the client set max_tokens — and opencode always does — so real failures were invisible to feedback. A client's token cap explains a truncated verdict and nothing else; it is now scoped to exactly that.

feedback.py folds observed failures into proficiency, so routing learns from your traffic rather than only the 23-task benchmark. Only failures are folded in: a structural 'ok' means the code parsed, not that it was correct, and recording those as 1.0 would flatten every score toward the ceiling. Failures the model did not cause — a client's own tight max_tokens truncating the answer — are recorded but excluded.

The classifier is the latency floor

Every routed request pays a full local classification round-trip before a single upstream token is requested, so the classifier model is the single biggest lever on interactive latency.

The default is mistral-nemo:12b, and it replaced qwen3.5:latest on measurement. Same prompts, same system prompt, temperature: 0, cold load excluded — 14 unambiguous category cases and 5 tier probes:

qwen3.5 mistral-nemo:12b
category correct 10/14 9/14
tier correct 1/5 (2 of them crashed) 3/5
hard failures 4 of 19 calls (21%) 0
latency mean / median / max 6.6 / 5.5 / 15.6s 1.7 / 1.8 / 1.9s

Category accuracy is a wash. What decides it is the failure column and the tail. Every one of those 4 failures is the runaway-thinking-trace mode below: ~15s spent to produce no JSON, which degrades to source: "fallback" — tier 2, general_chat. A reasoning model is the wrong tool for a job whose entire output is ~45 tokens of JSON.

mistral-nemo does not reason by default, which is why its tail is flat: its slowest call (1.94s) is faster than qwen3.5's median. End-to-end /route went ~10s -> ~1.7s. Both scored 5/5 on the local verification task, so the verifier moved with it and only one model stays resident.

Four settings keep this usable, each fixing a failure seen in practice:

  • max_retries=0 on the classifier client. The OpenAI SDK retries twice by default, so timeout_seconds silently became a 3x wall-clock bound — a request hung past 250s on a 120s setting and logged nothing.
  • max_output_tokens: 1024. Bounds a REASONING model's chain of thought. Unbounded, it cascades: Ollama keeps generating after the client gives up and serializes per model, so one runaway request queues every later request behind it and the timeouts spread. 256 was too tight — the trace consumed the budget and the model was truncated before emitting any JSON, which surfaced as an empty response. Does not bind for mistral-nemo; kept because it costs nothing unused and is the only guard if a reasoning model is swapped back in.
  • max_input_chars: 8000. The classifier decides a category and a tier; it does not need the document. Feeding it one is harmful, not merely wasteful — on a ~20k-token prompt qwen3.5 spent 28.7s and returned empty while mistral-nemo spent 41.8s echoing the input back inside its JSON. Both land on source: "fallback", i.e. 30-40s of local inference buying the answer an instant failure would have given. Clamped to head + tail (the instruction sits at one end or the other; never the middle), the same prompts classify correctly in ~2.2s. Nothing is lost: chat_completions measures the real conversation with estimate_prompt_tokens and takes the larger value.
  • fallback_tier / fallback_category. A classifier that times out, errors, or returns garbage now degrades to a configured mid tier flagged source: "fallback" instead of returning 502/503. The caller is a coding agent that would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier.
  • temperature: 0. At the default, the same prompt classified tier 2 then tier 1 on consecutive calls and routed to two different models.

But ~10s of local overhead on every message is a real tax for an interactive agent, where the upstream answer itself may take 2s. Still unaddressed: classify once per session rather than per message, cache by prompt hash, or skip classification for short prompts.

"Local" means your hardware, not this machine

The premise is that a local LLM classifies the task before a cloud model answers it. That does not require the GPU to be in the machine you are typing on, and usually it isn't — most developers already have WireGuard or a VPN back to a home lab. So the normal shape is router and editor on the laptop, Ollama on the workstation, classifier.base_url pointing across the tunnel.

classifier.base_url / api_key_env / model take any OpenAI-compatible endpoint. Verified end to end against a non-loopback address, classifier and verifier both.

Ollama binds 127.0.0.1 by default, so this fails with connection-refused until the serving host applies deploy/ollama-over-vpn.conf. Bind it to the VPN address, not 0.0.0.0: Ollama has no auth of any kind, so anything that can reach the port can run inference and enumerate your models, and 0.0.0.0 publishes it on whatever wifi the laptop is sitting on.

A cloud endpoint also satisfies that interface, and measured better than the local one — worth knowing before assuming local is the cheap side. Five prompts, same system prompt, temperature: 0:

local qwen3.5 (RTX 6000) NeuralWatt deepseek-v4-flash
mean latency 11.58s (4.95-15.76) 1.02s
categories agreed with the label 2 of 4 5 of 5
hard failures 1 of 5 (empty after 15.76s) 0
energy per call ~7e-05 kWh, on your meter 1.17e-05 kWh attributed
cost per 1,000 calls electricity + 6.6GB resident $0.093 (0.19% of quota)

11x faster, more accurate, and less attributed energy, on a machine with a 24GB card sitting idle. The one hard failure is the documented runaway-thinking-trace mode — qwen3.5 spent 15.76s and emitted no JSON, degrading to source: "fallback", i.e. a silent tier-2 guess. deepseek-v4-flash does not reason by default, so that mode does not exist for it. End-to-end /route goes ~10s -> ~1.0s.

Local classification was assumed to be the cheap option because local compute felt free. It is not free, it is just unbilled — the same substitution the cost axis had to unlearn. The shipped default is still local Ollama, because switching spends quota and that is a deployment choice, not a code one.

The verifier follows, but only to another Ollama

The local LLM check speaks Ollama's native /api/chat (the only way to set think: False), so it follows the classifier across a VPN but not to a cloud provider. It used to derive its URL by stripping /v1 off classifier.base_url, which meant moving the classifier at all would have pointed it at <that host>/api/chat. It now has its own verification.base_url and verification.model.

verification.model may be null only while both run on one host. Config load refuses the null once the hostnames differ, because the failure is silent: observed directly, with the classifier on NeuralWatt the verifier POSTed deepseek-v4-flash to localhost:11434, 404'd, caught it, logged "local verification unavailable" and recorded no sample. Verification would have looked enabled while producing nothing.

Structural verification needs no model at all — it is pure Python — so verification.local_llm_enabled: false leaves a host with no local inference fully functional, minus the refusal/incoherence class of failure.

Note also that opencode sends ~32K prompt tokens of system prompt and tool definitions on a trivial request, so the measured-size floor in estimate_prompt_tokens does real work — the classifier's own estimate for that request was two orders of magnitude low.

What's NOT built yet — pick up here

  1. Leaderboard priors are unfilled. leaderboards.yaml ships empty on purpose — inventing plausible-looking benchmark numbers would put fabricated data straight into routing, the same failure as the provider's static_fallback carbon constant this project already excludes. Until real sourced figures go in, a newly listed NeuralWatt family has no prior and relies entirely on self-eval accumulating. python leaderboard.py --check lists what is missing.

  2. Sampling depth for three models — now eco-only. 7 samples/model gives split-half agreement within 1.4x for 10 of 13, but kimi-k2.7-code-fast (29x), kimi-k3 (14x) and glm-5.2-flex (2.2x) are still unsettled. This no longer touches cost, which is priced per-request from the catalog, so it only affects eco — which is not an objective. Low priority unless eco comes back.

  3. Retry does not reach streaming. The iteration budget (iteration.py) retries after a failed check, but only on the non-streaming path — once bytes have gone to the client there is nothing to take back. Buffering to fix that would cost streaming itself, a worse trade for interactive work. POST /outcome is the answer for streamed traffic: it arrives afterwards, so it works identically either way.

  4. Session-directory attribution picks the wrong directory. The opencode plugin now genuinely reports — 31 accepted, 3 refused as ambiguous, 22 succeeded / 4 failed from real test runs, all folded in. But the "most frequent path" heuristic resolved one session to a dependency's source directory inside the project's own virtualenv (24 mentions) rather than the project being edited (22), because reading a library's source outweighed writing the code under test. It degrades safely — an ambiguous session is refused rather than misattributed, which is what the 3 × 409 were — but the heuristic needs to weight writes over reads, or anchor on the client's cwd instead of a path histogram.

  5. Local energy is not on the ledger. The router meters what NeuralWatt bills and reports, but the electricity its own classifier and verifier burn on local hardware is invisible to it. The whole premise is spending cheap local compute to avoid wasting expensive cloud compute, and right now only one side of that trade is measured — which is how "local is free" survived as an assumption long enough to be wrong (a hosted classifier measured faster, more accurate and lower attributed energy than the local one). Closing it means metering local draw (nvidia-smi / RAPL / a smart plug) and pricing it against a real tariff — utility rate data, ideally the user's own plan, including time-of-use bands. That would also make the local-vs-cloud comparison an actual number rather than a shrug.

Known open questions

  • Answered: cost and eco stay separate axes — grid intensity spans 13.6x across the catalog, so they rank models differently.
  • Answered: the GLM rows reporting grid_id: FI at 475 gCO2/kWh were carbon_source: static_fallback — a substituted constant, not a measurement. They are now excluded from eco rather than trusted. Still worth asking NeuralWatt why the fallback keeps the original grid_id, since that is what made it look like a real regional difference.
  • Three models still fail a split-half stability check at 7 samples. Is the instability real (variable serving conditions) or an artifact of when the sweep ran? Re-sweeping at a different hour would tell.
  • Answered, and the question no longer parses: tier-1 composites used to sit within 0.009 of each other because min-max normalization compressed them. There is no composite any more — ranking is quality first, cost as the tiebreak inside quality_tolerance — so nothing normalizes and nothing compresses.
  • Answered: the eval set exists (evals/tasks.yaml, 23 tasks, four scoring kinds) and tests/test_task_set.py keeps it honest. The open part is narrower now — coding_refactor and debugging are still flat at 1.00 across every model, so those tasks discriminate nothing and either need hardening again or should be conceded as non-discriminating. Try samples before hardening. docs_writing looked flat at the top too, and six more passes spread it 0.66-0.97 without touching a task; two samples per model is not enough to tell a saturated task from an unsampled one.
  • How much context-assembly (RAG-style retrieval) belongs in the classifier step vs. a separate pre-step? Leaning decoupled, undecided.
  • Should eco_score use real-time grid carbon intensity per request or a stable per-model average? Currently the latter, from the reference sweep. grid_carbon_intensity and grid_id are logged per observation, so this stays answerable from data without a re-run.

Config is strict: an unknown key is an error

Pydantic ignores extra keys by default, which means a typo or a misplaced setting loads cleanly, does nothing, and still looks configured. Every config model now inherits StrictModel (extra="forbid"), so both of these fail at load rather than silently:

verification.max_input_chars     # right key, wrong section
routing.min_tool_proficency      # sic

This is not hypothetical. max_input_chars shipped into the verification: block instead of classifier: and was accepted and discarded — it happened to match the code default, so behaviour was correct and the file was a lie. Editing it would have done nothing.

The corollary worth keeping: every knob belongs in config.yaml, not only in a Pydantic default. A default the file never mentions is invisible to anyone tuning it. classifier.outcome_attribution_window_seconds was removed in the same pass — it was declared, never read, and shadowed the verification one that actually is.

Setup

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < schema.sql
cp .env.example .env   # fill in NEURALWATT_API_KEY
python poller.py       # populate the catalog
python tier.py         # resolve tiers
python config.py       # sanity-check config loads
python -m uvicorn dispatcher:app --reload

Then set what is deployment-specific in config.yaml: classifier.model and classifier.base_url for your Ollama, and objective.plan_kwh_per_period to your own plan's quota (it is reported in /health as burn against the allowance; it does not gate anything).

Ollama must be reachable with the classifier model pulled — the name must match classifier.model in config.yaml:

ollama pull mistral-nemo:12b

local_vision ships enabled as a core feature (see below), so also pull its model unless you're turning it off:

ollama pull qwen3-vl:4b

It does not have to be on this machine. To use one across a VPN, point classifier.base_url and verification.base_url at it and apply deploy/ollama-over-vpn.conf on the serving host — Ollama binds loopback-only by default and will otherwise refuse.

requirements.txt is pinned. Recreating the venv with the old >= ranges jumped openai 2.53 -> 3.0 and httpx -> httpx2 without warning; a service that restarts on boot shouldn't change its dependency tree underneath itself. Bump deliberately.

pyproject.toml puts the repo root on sys.path for pytest — the modules live at the root rather than in a package, so pytest (console script) and python -m pytest would otherwise disagree about whether import config resolves.

Run as a service

deploy/ holds five systemd user units (dispatcher, plus a timer and a oneshot service each for the poller and the seed sweep) and one drop-in for a system Ollama — see deploy/README.md for install and operation. In short:

echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer

The dispatcher binds 127.0.0.1:8080. The poller timer is load-bearing, not housekeeping: stale_after_days is 3 with exclude_stale: true, so an unpolled catalog eventually marks every row stale and the router returns zero candidates for everything.

The service holds a billable API key and has no auth of its own. Loopback bind is the only thing standing between the open internet and your allowance; add auth before widening --host.

The same applies to an Ollama shared over a VPN — it has no auth either, so deploy/ollama-over-vpn.conf binds it to the VPN address rather than 0.0.0.0, which would publish it on whatever network the client happens to be on.

Pointing a coding agent at it

The /v1 endpoints are OpenAI-compatible, so any normal client works — opencode, an SDK, plain curl. Repo-local opencode.json is already wired up, so running opencode from a clone of this repo routes by default. For global use, merge provider.llm-router into ~/.config/opencode/opencode.json.

model name behavior
auto router picks; flex rows excluded so nothing is held during peak
auto:batch router picks; flex rows admitted, for overnight/async work
any real model id dispatched as asked, still logged

Streaming is proxied chunk by chunk rather than buffered, so tokens still render as they arrive. NeuralWatt emits its energy and cost blocks as SSE comment lines (: energy {...}) before data: [DONE] — ordinary clients ignore comments, so the stream passes through untouched while the router reads the telemetry on the way past. Without that, streamed calls would log no energy at all, which is most of the point of this project.

Try it

# Pick a model without spending anything
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
  -d '{"task":"Refactor this 800-line Django view into service objects."}'

# Same, but admit flex rows (held during peak, fine for overnight work)
curl -s -X POST localhost:8080/route -H 'content-type: application/json' \
  -d '{"task":"nightly code review","latency_tolerance":"batch"}'

# Actually call the winner and log energy/cost
curl -s -X POST localhost:8080/dispatch -H 'content-type: application/json' \
  -d '{"task":"What is a Python context manager?"}'

# The OpenAI-compatible surface any client can use
curl -s localhost:8080/v1/models
curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}'

# What did routing actually cost and burn?
sqlite3 -header -column router.db \
  "SELECT model_id, prompt_tokens, completion_tokens, energy_kwh, carbon_g_co2eq, cost_usd
     FROM energy_observations ORDER BY id DESC LIMIT 10;"

/route takes task_category, task_tier, and required_context_tokens overrides, which skip the classifier — useful for testing routing changes deterministically.