Adds seed_energy.py, which runs a fixed reference task N times per routable model and writes energy_observations rows tagged 'seed_reference'. Scoring now reads mean measured USD billed and mean measured gCO2eq from those rows instead of the catalog's list price. List price was not merely imprecise, it was inverted: deepseek-v4-flash lists 33% cheaper than gemma-4-31b ($0.28 vs $0.42 per 1M) but costs 95% more to run ($9.80e-05 vs $5.02e-05) and emits 284% more carbon. Billing is min($8.00/kWh x energy, 3 x list token price), validated against all 65 samples; the ceiling bound twice, both deepseek energy spikes, matching to the cent. That rule stays in the docs as rationale — no billing formula is in the code, since measured cost already accounts for it, including the flex tier. flex_cost_multiplier is therefore removed rather than left as a guess. Cost and eco deliberately stay separate axes. Collapsing them looked right — cost = 8 x energy, so they seemed to be one signal — but carbon is energy times the serving region's grid intensity, and the catalog spans 37 gCO2/kWh (FI) to 505 (US-MIDA-PJM). They rank models differently: glm-5.2-fast is 2nd cheapest and 6th cleanest; kimi-k3-flex draws 3.7x less energy than kimi-k2.7-code while emitting 3.6x more carbon. A test pins it. Only reference-workload observations steer routing. Organic traffic stays logged for accounting but is excluded, because per-request energy varies more with request shape than with the model — 19x across organic traffic on a single model, purely from differing prompt and completion sizes. cost_score and eco_score were byte-identical implementations; both now alias one normalize_inverted. Effect: 27 routing decisions (9 categories x 3 tiers) went from 2 distinct models to 3, and tier 1 flipped from deepseek-v4-flash to gemma-4-31b. Category still cannot influence a decision — proficiency is the only category-dependent term and remains empty, which is now the single highest-value gap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018xTPER7K8fNyKiuqNvTCTa
17 KiB
Local LLM Model Router
Status: design draft Owner: Aaron Lee Hardware: RTX 6000 24GB (local classifier/context-assembly), routes to Neuralwatt Cloud
Revised 2026-08-11 after first live run. Neuralwatt is now the only provider (OpenRouter dropped). Two findings from real API traffic change this document materially: billing is per-kWh, not per-token (§4.1), and one base model appears as several catalog rows that differ only by serving class (§3.4). Sections below are annotated where the original design assumed otherwise.
1. Goal
Use a local model on the RTX 6000 to analyze incoming coding/documentation tasks — pulling in relevant architecture docs and code context — and route each task to the cheapest/best-fit open-weight model across multiple providers, rather than defaulting every task to a single expensive cloud model. Router overhead runs on local hardware (electricity, not API spend), so the marginal cost of a routing decision is near zero.
2. Architecture
┌─────────────────────┐
incoming task ───▶│ Context Assembler │ (RAG-lite: pulls relevant docs/code)
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Local Classifier │ (RTX 6000, e.g. Qwen2.5-32B)
│ - task_tier │
│ - required_context │
│ - confidence │
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Decision Table │◀── scheduled pollers (pricing/benchmarks)
│ (SQLite, weighted) │
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Dispatcher │──▶ Neuralwatt Cloud
│ (OpenAI-compatible) │──▶ (future providers)
└──────────┬───────────┘
▼
┌─────────────────────┐
│ Verify / Escalate │ (local model spot-checks output)
└─────────────────────┘
Fits the same pattern as vigilar: scheduled pollers publish updates, a service holds current state, consumers read from it rather than hitting live APIs per-request.
3. Decision table schema
Two tables: a models table (one row per model+provider) and a proficiency table (one row
per model+category), joined at query time.
3.1 models
| column | type | notes |
|---|---|---|
model_id |
text | canonical model name |
provider |
text | neuralwatt today; column kept for future providers |
cost_per_1m_prompt |
real | listed token price |
cost_per_1m_completion |
real | listed token price |
cost_per_1m_prompt_cached |
real | e.g. Neuralwatt bills cached prefix at 25% |
energy_mwh_per_req |
real | Neuralwatt-only; null for token-only providers |
carbon_g_co2eq_per_req |
real | derived from energy × grid intensity, where available |
context_window |
int | max tokens, provider-advertised |
effective_context_window |
int | advertised minus a safety margin (see §3.2) |
tier |
int | 1 (cheap/simple) – 3 (frontier reasoning) |
availability |
text | active, deprecated, stale |
last_updated |
timestamp |
3.2 Context window criteria
Advertised context length isn't the number to route on directly — quality tends to degrade before the hard limit, and you need headroom for the response itself. Rules of thumb to encode:
effective_context_window= advertised context × 0.75, minus a reserved output budget (e.g. 4K tokens reserved for completion on coding tasks). Adjust the 0.75 factor per model once you have real data — some models hold up much better near their limit than others, and that's worth capturing per-model rather than as one global constant.- Classifier estimates
required_contextfor the task (assembled docs/code + prompt overhead), not a vague size bucket — an actual token estimate, so it can be compared directly againsteffective_context_window. - Hard filter, not a weighted factor: any model where
effective_context_window < required_contextis excluded from candidates entirely before scoring runs — a cheap model that can't physically hold the context isn't a cost tradeoff, it's disqualified. - Context size tiers (for quick filtering / dashboarding, not routing itself):
- small: ≤ 8K
- medium: 8K–32K
- large: 32K–128K
- xlarge: 128K+
- Overflow fallback: if required context exceeds every available model's effective window, the router should hand back to the context assembler to trim/summarize rather than silently picking the largest model and truncating mid-task — truncation failures are worse than a slower round-trip.
3.3 proficiency (benchmark scores by category)
One row per (model, category):
| column | type | notes |
|---|---|---|
model_id |
text | |
category |
text | see category list below |
score |
real | 0–1, normalized |
source |
text | leaderboard, self_eval, blended |
sample_size |
int | for self-run evals, so thin data can be flagged/discounted |
last_updated |
timestamp |
Starting category set (extend as needed — these map to the task types you're actually routing):
| category | example task | typical source |
|---|---|---|
coding_general |
boilerplate, small functions | Aider polyglot, self-eval |
coding_refactor |
multi-file refactors, architecture-aware edits | self-eval (most predictive here) |
debugging |
root-cause / fix-from-error-trace | self-eval |
docs_writing |
docstrings, README, comments | self-eval, cheap to score |
summarization |
log/alert summarization (vigilar-style) | self-eval |
translation |
natural-language translation | leaderboard |
reasoning_math |
multi-step logic, math-adjacent | LiveBench, leaderboard |
tool_use_agentic |
function calling, multi-step agent tasks | leaderboard + self-eval |
general_chat |
catch-all / low-stakes | leaderboard |
Blending rule: score = 0.3 × leaderboard_score + 0.7 × self_eval_score once
sample_size for self-eval crosses a minimum threshold (e.g. 10 runs); before that, fall back
to leaderboard_score alone so a thin/noisy self-eval doesn't dominate early. Store both
components, not just the blend, so you can re-weight later without re-running evals.
Category selection at routing time: the local classifier tags each incoming task with a
primary category (and optionally a secondary one for hybrid tasks, e.g. "refactor + docs").
proficiency_score used in the composite scoring formula (§4) is the score for that specific
category — not an average across categories — so a model that's great at docs_writing but
mediocre at coding_refactor doesn't get miscredited on refactor tasks.
3.4 Serving class (added after first live run)
Neuralwatt ships ~6 base models as 19 catalog rows. The id suffixes are three orthogonal
dimensions — hence ids like glm-5.2-short-fast-flex — and they are not visible anywhere
except the id string and the prose description:
| suffix | column | meaning |
|---|---|---|
-flex |
latency_class |
discounted async tier; requests held server-side during peak until a capacity gap opens |
-fast |
reasoning_mode |
thinking disabled or capped to a short budget. Not queue priority |
-short |
context_variant |
200K pool with a bounded reasoning budget |
These rows carry identical catalog pricing, so they tie exactly on every scored dimension.
Left to the weighted score, a coin flip decides whether an interactive request lands on a
preemptible flex row. So latency tolerance is a third hard filter, not a weight:
interactive (default) excludes -flex; batch admits it.
Access gating is likewise prose-only — 6 of 19 rows say "Private preview (grant-gated)" or
"(Canary)" in free text with no structured field. Those are parsed into access_level and
excluded by default, since routing to one earns a 403 at dispatch rather than a bad answer.
4. Weighted scoring
Router computes a composite score per candidate model and picks the max, after three hard filters are applied (not weighted — these disqualify a candidate outright):
effective_context_window >= required_context(§3.2)tier >= required_tier(from classifier)- serving class compatible with the request's latency tolerance, and
access_levelactually reachable by this account (§3.4)
score = (w_cost × cost_score)
+ (w_eco × eco_score)
+ (w_prof × proficiency_score[task_category])
- cost_score: normalized inverse cost (cheapest candidate = 1.0, scaled down from there). ⚠️ See §4.1 — the input this currently reads is not the price being billed.
- eco_score: normalized inverse
carbon_g_co2eq_per_req, which Neuralwatt reports per-completion; models without a carbon figure get a neutral/default score rather than being penalized, so a model new to the table isn't downweighted just for lacking history. - proficiency_score: from the benchmark refresh — general leaderboard score blended with your own eval set (see §5), weighted toward the latter over time as it accumulates.
w_cost, w_eco, w_prof live in a small config file (not hardcoded) so the balance can shift
per-project or per-mood — e.g. crank w_eco up for background/batch jobs where latency and
cost don't matter much, crank w_prof up for anything touching production code.
Hard floor, not just a weight: task tier from the classifier acts as a minimum bar (don't let a great cost/eco score smuggle a task into a model that's actually too weak for it) — the weights decide within the eligible set, not whether a bad-fit model can win purely on price.
4.1 Billing is per-kWh — the cost axis needs rethinking
The line above ("take whichever the account is actually billed under") turned out to be the
whole ballgame. Measured across five models on 2026-08-11, cost_usd / energy_kwh came back
8.00 every time: Neuralwatt bills a flat $8.00/kWh and the catalog's
input_per_million / output_per_million figures are not what this account is charged.
The ranking consequence is not subtle. On the same prompt:
| model | list $/1M out | billed USD | kWh |
|---|---|---|---|
kimi-k2.7-code-fast |
4.00 | 2.25e-03 | 2.81e-04 |
kimi-k3-fast |
15.00 | 2.17e-04 | 2.71e-05 |
The model listing at a quarter the price cost 10x more to run. Any cost_score computed
from list price ranks these backwards.
This collapses two of the three scoring axes into one. cost_usd = 8.00 x energy_kwh exactly,
so w_cost and w_eco are weighting the same underlying quantity — 0.6 of the composite on
one signal, with proficiency left holding 0.4. Cost and eco separate only through grid carbon
intensity (grid_carbon_intensity_gco2perkwhr, 37.0 on grid FI at time of writing), which is
a real distinction only when routing across regions.
Two further wrinkles before this can be fixed:
- Energy is noisy. Identical token counts (18 -> 900) billed anywhere from 1.8e-05 to 8.8e-05 — a 4x spread run to run. A single observation is not a per-model estimate.
- Cold start.
energy_observationsis empty until real traffic accrues, so energy-based scoring needs a seeding sweep (fixed prompt, N runs per routable model) before it can drive anything.
4.2 Resolved: measured cost, and the two axes stay separate
Both sub-scores now read measured values from a fixed reference workload
(seed_energy.py, 13 models x 5 samples, tagged seed_reference): mean USD actually billed
for cost, mean gCO2eq for eco. No billing formula lives in the code — the rule in §4.1 is
documentation for why list price is not used, not logic.
The tempting simplification — cost = 8 x energy, therefore cost and eco are one axis — is
wrong, and the sweep is what caught it. Cost tracks energy, but carbon is energy x the
serving region's grid intensity, and the catalog spans 37 gCO2/kWh (FI) to 505
(US-MIDA-PJM). So the axes disagree: glm-5.2-fast is the 2nd cheapest model and the 6th
cleanest; kimi-k3-flex draws 3.7x less energy than kimi-k2.7-code while emitting 3.6x more
carbon. Collapsing them would have silently picked a side.
Only reference-workload observations feed scoring. Organic traffic stays in the table for cost accounting but is excluded from routing, because per-request energy varies far more with the shape of a request than with the model — across organic traffic on a single model it spanned 19x purely from differing prompt and completion sizes. Averaging that in would rank models by what they happened to be asked.
Effect on the hard-filter-plus-score pipeline: with list-price scoring, 27 routing decisions
(9 categories x 3 tiers) produced 2 distinct models. With measured scoring, 3 — and tier 1
flipped from deepseek-v4-flash to gemma-4-31b, correctly. Category still has zero influence,
because proficiency remains empty and is the only category-dependent term (§3.3).
5. Freshness / scheduled jobs
Two independent pollers, different cadences:
- Pricing poller (hourly–daily): hits
Neuralwatt /v1/models(unauthenticated for public listing), writes cost/context/serving-class/access columns, flags rows unchanged > N days asstale. Note the catalog carries no energy fields at all — energy arrives only per-completion, so the dispatcher populatesenergy_observations, not this poller. - Benchmark poller (weekly, or on-demand): blends general leaderboard data (LMSYS Arena, LiveBench, Aider polyglot) with a small self-run eval set built from your actual task patterns (docstring quality, refactor correctness, doc-formatting) — the self-run set is more work but far more predictive of real routing quality than a general leaderboard number.
Alert (MQTT topic, matching vigilar's pattern) on: price change beyond some threshold, model deprecated/removed, or a row going stale — these are the silent-failure modes that break a router without it noticing.
6. Escalation path
More important than the initial routing decision. Needs to be cheap and low-friction:
- Downstream agent/human can flag "wrong tier" → bumps the task one tier and retries
- Local verify step (optional): before accepting cheap-tier output, local model does a fast sanity pass; on failure, auto-escalate rather than surfacing bad output
- Track escalation rate per task-type — a category that escalates often is a signal the classifier's tier mapping needs adjusting, not that escalation is broken
7. Metrics to track (not just savings)
- $ saved vs. always-frontier baseline
- kWh / carbon g CO2eq per task-batch (Neuralwatt gives this natively)
- Escalation rate (overall and per task-type)
- Silent failure rate — cheap-tier output accepted but later found wrong (harder to measure, worth a periodic manual audit sample)
A router that's cheaper but quietly worse isn't a win — savings and error rate need to be reported together.
8. Open questions
- How much of "required_context" should the local classifier assemble itself vs. defer to a separate RAG step? (Leaning: keep these decoupled — context assembly is its own mini-project.)
- Self-run eval set: what's the minimum viable set of representative tasks to start scoring proficiency meaningfully, without it becoming a maintenance burden on its own?
- Whether
eco_scoreshould factor in grid carbon intensity at request time (Neuralwatt exposes real-time grid data) vs. a stable per-model average — real-time is more accurate but adds volatility to routing decisions that may not be worth it for marginal gains.
9. Next build steps
Poller: Neuralwatt → SQLite decision tabledone, verified against the live APIMinimal FastAPI dispatcher: reads table, applies weighted score, calls provider, returnsdone —/health,/route(dry run),/dispatchLocal classifier prompt/schemadone — the allowed category list is injected from config at call time, since the model otherwise invents labels that join against nothingRe-base the cost axisdone (§4.2) — cost and eco now read measured billed USD and gCO2eq rather than list priceSeeddone —energy_observationswith a fixed-prompt sweepseed_energy.py- Populate
proficiency— the last inert axis. 0.4 of the composite weight does nothing, andtask_categorycannot change a routing decision until it does - Wire into one real workflow (coding agent subagent spin-off) end to end before generalizing