# Local LLM router config # All weights, thresholds, and provider settings live here so they can be # tuned without touching code. Loaded/validated by config.py. objective: # Quality is the objective. Cost is a constraint and a tiebreak. Eco is # logged per request but is NOT optimized here — that judgement is made # outside this router. # # This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4). # Measurement killed it: turning the cost weight from 0.4 to ZERO changed # the winner in only 2 of 6 categories, so the blend was never steering on # quality — while 60% of every decision adjudicated fractions of a cent # (all real traffic to date totals $0.07). # Proficiency differences smaller than this are treated as equal and the # cheaper model wins. This is measurement noise, not preference: scores # currently rest on 2-3 samples per category, so a 0.05 gap is # indistinguishable from sampling variation and paying for it buys noise. # Narrow it as samples accumulate. # Above this pace ratio, the dashboard raises an alarm. # 1.25 = burning 25% faster than the plan allows. plan_pace_warn_ratio: 1.25 quality_tolerance: 0.1 # Cost is priced per-request from catalog prices, NOT from a benchmark # sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper # than deepseek-v4-flash; on a realistic 70k-token prompt deepseek is 5.0x # cheaper. Attribution inverts with prompt size, so a fixed-shape benchmark # cannot rank models for a workload of another shape. Catalog prices scaled # to the actual request agree with the live measurement, cost nothing, and # need no sweep. # Share of prompt tokens served from the provider's prefix cache. Agent # clients resend the whole conversation each turn, so most of it is a hit. # # Measured token-weighted across 50 sessions and 40.7M tokens (2026-08-23): # 91.7% overall, 92.6% on the sessions above 400k tokens, which are the ones # that carry the cost. The previous 0.84 came from 2.2M tokens of much # earlier traffic. # # Raising it changed NO winner in any of the nine categories at 60k context # — checked before editing. It is here because the number should be true, # not because the routing needed it. assumed_cache_rate: 0.917 # Completion length assumed when pricing a request. Real sessions here median # around 200-400 completion tokens against enormous prompts. assumed_completion_tokens: 500 # Per-request ceiling on measured ENERGY, in kWh. null disables it. # # Denominated in kWh rather than dollars. For scale: the reference task runs # ~5e-06 kWh on the cheapest model and ~2.2e-04 on the most expensive. # Overage is billed against the account's credit balance; # plan_kwh_per_period gates nothing. max_energy_per_request: # The subscription's kWh allowance per billing period, for reporting burn in # /health. This is a planning figure only: per-request traffic is never # refused for exceeding plan_kwh_per_period — it gates nothing. Set to match # your plan; null disables the report. NeuralWatt also returns # allowance_remaining_usd per request, which is logged for /metrics, but that # is a dollar figure while the plan is denominated in energy. plan_kwh_per_period: 6.25 # Hours of recent balance history used to estimate the burn rate. The most # recent monotonically-decreasing segment of allowance_remaining_usd values # (segments split at each balance increase) is examined over this window. quota_burn_window_hours: 24 # Projected hours of runway below which /metrics and dashboards emit a # low-warning boolean; only fires when a burn rate estimate exists and the # projected remainder is positive but short. quota_runway_warning_hours: 6 # A burn estimate needs at least this many balance samples in the most recent # monotonically-decreasing segment (post-top-up resets the segment); below # which the estimate is None with an explanatory runway note. quota_burn_min_segment_samples: 3 # ...and the segment must span at least this many hours, otherwise the # estimate is None with an explanatory note (never a wild extrapolation). quota_burn_min_segment_hours: 0.5 # The day-of-month your NeuralWatt subscription billing cycle resets. Set # this to YOUR real billing day so the admin quota modal shows a genuine # next-reset date instead of a misleading rolling-window start. null disables # the feature and the modal shows "not configured". Valid range: 1-28 (skip # 29-31 to avoid shorter-month edge cases). billing_reset_day: 6 # Rejection-rate detection over route_decisions rows that selected no model # (422 "no model satisfies the hard filters"). The signal is novelty or rate — # NEVER mere presence: this deployment routinely has ~3 rejections/hour of # ordinary over-large tier-3 requests that are behaving as designed (6 in the # last 24h, 17 all-time at 2026-09-05), so a count-only tripwire would be # permanently on. # # alert window (hours): rejections this recent are counted per # (task_tier, digit-normalized reason) group. rejection_warning_window_hours: 1 # baseline window (hours) BEFORE the alert window: groups absent here are # "new". A new group warns from 2 occurrences — the 2026-09-04 vision # incident produced exactly 2 and no rate threshold can sit below routine # noise yet above that. rejection_warning_baseline_hours: 24 # a group already present in the baseline warns only at this many # occurrences inside the alert window: 6 = 2x the observed routine hourly # peak, far below the dozens/hour a deprecation flare produces. rejection_warning_min_count: 6 # How far back /metrics looks to decide which models the router has actually # picked. Long on purpose: this answers "has this model EVER been chosen", # and a model that only wins one category can go days between selections. # Feeds two very different reports -- see metrics.selection_coverage. selection_coverage_window_hours: 168 # Prefix-cache rate measured from energy_observations, and the expiry check # on assumed_cache_rate above. That constant was measured ONCE, on # 2026-08-23, and it is the highest-leverage term in routing.estimated_cost # on a 100k-token prompt — a premise that large should not go unchecked just # because it was true when it was written. # # The comparison is a DIVERGENCE from assumed_cache_rate, not an absolute # floor. "Is the cache working" is the wrong question; a deployment whose # real rate is 0.60 is not broken, it is mispriced, and a floor would say # nothing about the number the cost model is actually built on. # # Trailing window, in hours. Long for the same reason # selection_coverage_window_hours is: only a fraction of rows carry a # reported cached count, so a 1h window is usually empty. cache_rate_window_hours: 168 # Absolute divergence from assumed_cache_rate that warns. 0.10 sits well # outside ordinary session-mix drift (measured 2026-09-13: 0.879 neuralwatt # / 0.901 openrouter against the assumed 0.917) and well inside the 0.952 vs # 0.478 same-model/switched gap the waves plan is chasing. cache_rate_warn_margin: 0.10 # Minimum reported-cache observations before either warning fires, per # aggregate and per (provider, model) group. Observations, not tokens: one # 200k-token prompt outweighs a hundred ordinary turns, so a token floor # would let a single request's luck read as a measurement. cache_rate_warn_min_observations: 25 # --- Report-only measurement series. Neither is read by routing. --------- # # Cost-estimator calibration: routing.estimated_cost's prediction against # the provider's own billed figure, per (provider, model), joined on # request_id. Measured 2026-09-12 at ~100% telemetry coverage, the estimator # is high by 1.6x-13x depending on the model. # # The scale error is NOT the finding — a uniform overestimate reorders # nothing, because the ranking is a comparison and every candidate moves # together. The SPREAD in that error is the finding: it priced # deepseek/deepseek-v4-flash below qwen3.6-35b while the bill said the # reverse. /metrics reports `spread` for exactly that reason. # # Nothing applies these factors. They are here so their stability can be # judged first — this project has already mistaken one moment of a moving # per-model quantity for a constant (see CLAUDE.md, "Attribution drifts # across hours, so sampling must too"). # # Trailing window, in hours. Long for the same reason # cache_rate_window_hours is: only rows with both an estimate and a billed # figure qualify, and a short window is usually too thin to read. cost_calibration_window_hours: 168 # Joined observations before a (provider, model) factor is marked # `sufficient` and allowed to set the reported spread. Groups below it are # still listed with their counts — the count is itself information — but a # ratio over three requests must not become the headline. cost_calibration_min_observations: 25 # Router-observed latency: router_wall_seconds and router_ttft_seconds from # energy_observations, p50 and p95 per (provider, model). These are the # ROUTER's clock, not the provider's duration_seconds — which is why this # can see OpenRouter at all, since OpenRouter reports no duration. # # It found z-ai/glm-5.3-flash, a model with "flash" in its name, at a p50 of # 9.6s to first token against 1.6s for deepseek-v4-flash on NeuralWatt. # Reported, not scored: latency is not an objective here, and one reading of # a quantity that tracks pool load is not grounds to make it one. latency_window_hours: 168 # Observations before a group's percentiles are marked `sufficient`. Applied # SEPARATELY to the wall and TTFT counts, because TTFT is streaming-only by # nature and a buffered deployment legitimately has fewer of them. latency_min_observations: 25 # Gate: with this off, routing is byte-identical to today. # Flip on only after the Wave 1 post-restart baseline day; # see plans/token-waste-waves.md Wave 2 gate. incumbent_cache_pricing: false # Challenger cache-rate dial. # - null (default): neutral, follows assumed_cache_rate → today's ranking exactly # including the incumbent; this is the off position for tuning. # - 0.0: challengers priced as fully cold prompts (maximum incumbent advantage). # - Any value between is a partial cache-penalty — the whole wave dials # from off to full here, no revert needed. incumbent_challenger_cache_rate: # Seconds between refreshes of the measured per-(provider, model) cache # rates. The TTL is measured from the first call after the process started # (time.monotonic), so a brief post-restart cold period is expected. incumbent_rate_refresh_seconds: 300 # Minimum reported-cache observations before a per-(provider, model) # cache rate is trusted for pricing decisions — the independent pricing # floor. Tuning this does NOT move the cache-rate warning floor # (cache_rate_warn_min_observations): the two knobs answer different # questions with different failure costs. # # Shipped with a documented default (25) so the code always reads a # concrete float when the penalty is engaged. incumbent_rate_min_observations: 25 # Credit-aware routing attenuation (OFF BY DEFAULT). # When enabled, a provider configured with a balance_url (e.g. OpenRouter's # prepaid account balance) gets its *comparison cost* inflated inside the # quality-first ranker as its account balance nears zero. Quality bands still # win; this only shifts ties. Providers whose balance comes from per-request # energy telemetry allowance_remaining_usd (e.g. NeuralWatt's overage-billed # subscription) are ALWAYS multiplier 1.0 regardless of their reading, so a # low soft_floor_usd never biases routing toward an attenuated provider just # because a telemetry provider's allowance reads near zero under normal use. # # This knob is deliberately NOT exposed in the admin UI persisted-config # allowlist (_CONFIG_ALLOWLIST in admin.py) or _ProviderUpdateBody; enabling # or tuning it requires editing this file and restarting llm-router.service # (the dispatcher's module-level cfg binds at import). credit_attenuation: enabled: false soft_floor_usd: 5.0 # balance >= this -> multiplier 1.0 zero_floor_usd: 0.0 # balance <= this -> max_multiplier max_multiplier: 5.0 # maximum cost-inflation at/below zero floor refresh_seconds: 300 # cache duration for resolved multipliers context: safety_factor: 0.75 # fraction of advertised context treated as usable default_output_reserve_tokens: 4096 # Ceiling on the output reserve, as a fraction of the usable window. A # provider's advertised max_output_tokens is normally a small per-request # cap, but OpenRouter reports max_completion_tokens -- "the most you may # ASK for", 0.8-0.9 of context on a dozen rows. Subtracting that whole # left 12 models at an effective context of 0, silently unroutable. # Only the catalog-derived reserve is capped; a per_model_overrides # reserve is a measurement and is used as written. max_output_reserve_fraction: 0.5 # Per-model exceptions to the two settings above, for a row whose real # limits you have measured. The global factor has to hold for the whole # catalog, so it is deliberately pessimistic; a model you have actually # pushed to its limit deserves its own number. # # Both keys are optional and each falls back to the global on its own. # An override of 0 reserve tokens means zero, not "unset". # # per_model_overrides: # qwen3.6-35b: # safety_factor: 0.85 # output_reserve_tokens: 8192 # # Takes effect on the next `python poller.py` -- effective_context_window is # computed at poll time, not per request. per_model_overrides: {} tiers: # Maps a tier number to a human label, purely for logging/dashboards. 1: "cheap / simple" 2: "mid / general" 3: "frontier / high-stakes" tiering: # Auto-tiering pass knobs. cheap_completion_max is the completion-cost # (per 1M tokens) ceiling below which a model is eligible for tier 1. # model_tiers overrides the heuristic per model_id and applies to ALL # providers (limitation vs a (model_id, provider) key). cheap_completion_max: 1.00 # Advertised context_window at or above which a model is NOT eligible for # tier 1, whatever it costs. Tier is a capability FLOOR (routing drops any # row with tier < required_tier), so tier 1 means "simple work only" — and # deciding that on price alone excluded deepseek-v4-flash from every tier-2 # request purely for being $0.28/1M, despite a 1M window and 1.00 on all # three coding categories. # # 512000 sits in the empty band between the catalog's 256K class (262128) # and its 1M class (1048560) — a 2x margin either side, so it is not fitted # to any one model. gemma-4-31b (256K) stays tier 1; the 1M rows do not. tier1_context_max: 512000 model_tiers: # This model lives on Ollama, not the NeuralWatt cloud catalog, so the # tiering heuristic never sees it -- src/tier.py OVERWRITES models.tier on # every run, so this pin is the ONLY thing that survives. If you swap the # local dispatch model, change the KEY here too or it silently loses its # tier. Pinned to tier 1 (cheap / simple) because it is targeted at # lightweight summarization and diff-checking — precisely the use-cases # tier 1 was designed for. qwen2.5-coder-router:14b: 1 proficiency: # Blending rule: leaderboard vs self-eval, once self-eval sample size # crosses the threshold below. Below threshold, leaderboard score alone # is used so thin self-eval data doesn't dominate. self_eval_min_samples: 10 leaderboard_weight: 0.3 self_eval_weight: 0.7 # Prior strength (k) for the empirical-Bayes shrunken estimate: the peer # prior contributes k pseudo-observations, pulling each noisy per-model # score toward the global average. 20 pseudo-observations damps thin # self-eval data (n < 100 per model) without erasing the per-model signal. outcome_prior_strength: 20 categories: - coding_general - coding_refactor - debugging - docs_writing - summarization # file_summarization + diff_checking are served by the local dispatch model # (see local_dispatch_models: section below). # They are kept adjacent to summarization since this model is targeted # at lightweight summarization and diff-checking tasks. - file_summarization - diff_checking - translation - reasoning_math - tool_use_agentic - general_chat exploration: # Epsilon-greedy exploration: on the epsilon share of requests, picks the # hard-filter-eligible candidate with the fewest outcome samples (tie-break: # lowest cost) instead of the highest-score model. Deliberately ON by default # because the system is a ranking loop — without periodic exploration it # converges on whatever happens to be sampled most, creating exposure bias. # Adjust once real exploration traffic (was_exploration=True in # route_decisions) shows how often the exploration path picks differently # from the greedy path. enabled: true # Probability of exploration per request. 0.03 = ~3% of requests take an # exploratory path, giving ~97% exploit on known winners while still # occasionally sampling under-explored candidates. epsilon: 0.03 # Cap: exploration candidate cost must be <= this multiple of the ranking # winner's cost. 4.0 gives headroom — an under-sampled model can be more # expensive than the greedy winner without eating the explore budget on # wildly off-target picks. max_cost_ratio: 4.0 # Only tiers 1 and 2 models are eligible for exploration. Tier 3 (frontier) # is too expensive to spend on random sampling — save the explore budget for # models where being wrong costs less. max_tier: 2 escalation: enabled: true max_tier: 3 # Bump the tier when the classifier is unsure of its own call. DEFAULT OFF: # this pays frontier prices on a hunch, before anything has gone wrong. The # iteration budget below spends after a check has actually failed, which is # strictly better on both mandates — the cheap attempt usually succeeds and # costs nothing extra, and when it fails you have evidence. preemptive_on_low_confidence: false min_confidence_before_bump: 0.6 iteration: # A tier is not only a capability floor, it is a budget for getting the # answer right. These are corrective attempts AFTER a verification failure, # not speculative retries. # # Retries are matched to the failure: a truncated answer gets a bigger token # budget on the SAME model (a different one would also run out), while a # malformed answer escalates to the next-best candidate (more tokens will # not make unparseable output parse). enabled: true attempts_by_tier: 1: 0 # cheap/simple — one shot; iterating costs more than it is worth 2: 1 3: 2 # Interactive requests are capped below their tier's budget regardless of # tier: every retry doubles time-to-answer, and in interactive use latency # IS a quality loss. Batch work does not care. max_attempts_interactive: 1 pinch: # Relevance-based context pruning (ported from the MIT-licensed llmrouter's # "pinch"). This is an OPTIONAL, pre-dispatch stage: when a conversation # exceeds budget_tokens, the provider-bound messages are trimmed BEFORE any # paid token is sent upstream. User/assistant/system messages are always # kept verbatim; only old TOOL RESULTS are shortened or dropped, because # they carry the bulk of a long agent session's tokens and are least needed # in full by the time the next turn is answered. # # On by default. Tool results can only be dropped safely because tool # outputs are idempotent enough for a placeholder; a wrong guess here loses # context, so start conservative (large budget, small reduction) and watch # route_decisions / pinch stats on real traffic before widening it. enabled: true budget_tokens: 50000 keep_last_turns: 4 max_summarize_chars: 4000 # keep_last_turns has no size limit inside it: an entire autonomous # tool-call loop with no new user message can be one protected turn, and # one outsized tool result inside it (a full verbose test run, a huge file # read) ships verbatim regardless of size. Measured live 2026-09-06: a # 324k-token conversation shrank only ~8% because nearly all of it sat # inside the protected window. This closes that gap: any tool result # inside the protected window over this many characters still gets the # same head/tail elision candidates get. Deliberately a much higher bar # than max_summarize_chars -- recent results are more likely to still # matter -- so it only catches true outliers. Set to null to disable. protected_max_chars: 20000 # Prefix-stability probe (Wave 1 item 1.3 of plans/token-waste-waves.md). # The provider bills the longest byte-identical PREFIX of a prompt at the # cached rate, so rewriting an early message re-bills everything after it. # The uniform pruning path compresses a contiguous positional region, so an # append cannot disturb it; the relevance path below compresses a prefix of # a relevance-ORDERED list, so a growing deficit pulls in one more candidate # each turn at an arbitrary message POSITION. Offline that rewrites 75% of a # payload's tokens; the live cache rate on pruned turns is 0.924, which is # not what that should look like. This settles it on real traffic. # # When on, each decision row records prefix_divergence_index, # prefix_tokens_after_divergence and prefix_prev_message_count -- three # integers. HASHES ONLY, and not even those: the per-message digests live in # process memory for exactly one turn and never reach the database, so # nothing reversible is stored and a restart costs one comparison per # session. Measured on the live median payload (98k tokens in, 74k out): # ~1.0 ms per turn, against a request path whose floor is a provider # round-trip of 1.4-2.0 s, plus ~10 bytes per decision row. prefix_probe: true relevance: # Off by default, matching every other new-and-unproven knob in this # project — and specifically requires pinch.enabled too, since this has no # effect otherwise. Ship it, watch route_decisions / pinch stats on real # traffic, then decide the default. enabled: true # An EMBEDDING model, not a chat model — this must not point at # classifier.model or verification.model. Pull one on the same Ollama: # ollama pull nomic-embed-text model: "nomic-embed-text" base_url: "http://localhost:11434/v1" timeout_seconds: 10 # Below this many trim-eligible candidates, skip the embedding call # entirely and fall back to uniform compression — a network round trip # to rank one candidate decides nothing. min_candidates: 2 session_cache: # In-memory per-session classification cache. Remembers the last # task_category / task_tier decision for each session for a few minutes, so # a long agent session skips the ~1-2s classifier round-trip on every turn. # Capability flags (tools / images / json) are NEVER cached — they are read # fresh from each request. Fallback classifications are NEVER cached. No # persistence: a restart just reclassifies each session once. # # Off by default, matching every other new-and-unproven knob in this # project: ship it, watch route_decisions.source="cached" on real traffic, # then decide the right default. enabled: true staleness_seconds: 1200 circuit_breaker: # Passive availability circuit breaker, on by default. When enabled, a model # that returns 5xx is temporarily excluded from routing with exponential # backoff; recovery is passive (the next real request becomes the probe once # the cooldown passes). It has a low-risk failure mode even when wrong. enabled: true initial_cooldown_seconds: 30 max_cooldown_seconds: 600 backoff_multiplier: 2.0 routing: # Access gating is prose-only in the NeuralWatt catalog ("Private preview # (grant-gated)", "(Canary)"), so the poller parses it into access_level and # routing excludes anything not listed here. Add 'preview'/'canary' only if # the account actually holds the grant — otherwise dispatch earns a 403. allowed_access_levels: - public # '-flex' rows are held server-side during peak until a capacity gap opens. # That's correct for overnight/batch agent work and wrong for anything # interactive, so a request has to opt in via latency_tolerance. default_latency_tolerance: interactive # 'interactive' | 'batch' # Operator's default stance on routing to '-flex' serving-class rows for # requests that do not state one explicitly. A 4-position scale: # no-flex never route to a flex row # auto decide per request (the default; keeps existing behavior) # prefer-flex flex first, standard as fallback # force-flex flex only # 'force-flex' is the dangerous global default: it bypasses the # latency_tolerance: interactive hard filter for EVERY request, so even a # request that opted into interactive routing would admit rows that are held # server-side during peak. Use it only when you are certain the caller can # tolerate flex latency globally. default_flex_preference: auto # Bare `auto` resolves to this profile. This is the profile name the router # uses when the client does not specify one explicitly. It must name a # built-in profile or a config-defined profile; the value is validated at # load and will be editable from the admin Controls page in a later wave. default_profile: "default" # Minimum tool_use_agentic proficiency required of a model when the REQUEST # carries tool definitions. A filter, not a weight, because it is a # capability requirement rather than a preference. # # It does NOT ask whether the task is agentic — it asks whether the model # can be trusted with tools that are on the table. The measured failure is # the second one: deepseek-v4-flash scores 0.33 here, and the recorded case # is a NON-agentic prompt ("it is 1:20pm, my meeting is at 3pm, how many # minutes?") where it called two tools instead of subtracting. A model that # over-reaches is a hazard wherever tools exist, not only where a classifier # would say "agentic" — which it cannot do anyway: asked to identify six # unambiguous tool-use prompts, the local models managed 2/6 and 1/6. # Whether tools are present is stated in the request body. Read it. # # 0.5 sits in the empty band between the only two values the catalog # currently holds (0.33 and 1.00), so it is not fitted to either. A model # with NO measured tool score is unproven rather than proven bad, and is not # dropped. # # CURRENTLY DISABLED (null) pending experimentation. The trade being # measured: opencode sends `tools` on essentially every request, so with the # filter on, deepseek-v4-flash is excluded from ordinary agent traffic and # the ~7x cost advantage on coding routes goes unused. With it off, that # advantage applies — and a model measured at 0.33 on tool use handles # requests where tools are on the table. # # What would settle it is outcome data, not another benchmark: run with it # off, let POST /outcome report real pass/fail, and compare # tool_use_agentic proficiency for deepseek before and after. That is the # one signal here that knows whether the work actually worked. min_tool_proficiency: tool_use_category: tool_use_agentic # Request-side capability gates. These read the request body (image parts, # response_format) and hard-restrict to models whose catalog row declares the # capability. Unlike min_tool_proficiency these are NOT proficiency gates — # a wrong guess is a guaranteed provider 400, so they gate by default and # fail closed when the catalog flag is unknown. require_vision: true require_json_mode: true # There is deliberately no flex discount knob. Whether a flex row is usable # at all is a hard filter above (latency_tolerance), not a price adjustment # -- being held server-side during peak is a latency property, and the # catalog advertises flex and standard at the same token price anyway. local_vision: # Local Ollama vision fallback. Used ONLY when a request carries image # parts and routing finds NO cloud model that supports_vision — the cloud # catalog currently excludes the cost leader (deepseek) on vision, so a # fallback is what keeps image requests working instead of 422ing. # It speaks the OpenAI-compatible /v1 surface, so this is an Ollama endpoint # and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host. enabled: true base_url: "http://localhost:11434/v1" api_key_env: # A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag. # Measured live: the base tag comes up at Ollama's own default num_ctx # (32768) and costs 9.4GB loaded — resident alongside the classifier's # pre-fix 14GB, that left 1.9GB free on a 24GB card. num_ctx cannot be set # per-request here: verified live that Ollama's OpenAI-compatible endpoint # (0.22.0) silently ignores it under every field shape tried, so it has to # be baked into the model tag itself: # printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision # ollama create qwen3-vl-router:4b -f Modelfile.vision # # Reduced 16384 -> 8192 on 2026-09-03 to buy back KV cache. Most of this # model's footprint is KV, not weights: 3.3GB on disk but 6.9GB resident at # 16384, and 5.6GB at 8192. That 1.3GB is what lets the dispatch model run # at num_ctx 32768 (17GB) and still leave vision RESIDENT -- 23.2GB of 24GB # together. Without it, loading vision EVICTS the dispatch model entirely # and the next classification pays a ~6s cold reload on the latency floor. # # This is the one path that still has no measured ceiling: it gets the RAW # message list (dispatcher.py calls it BEFORE pinch pruning), so 8192 is a # judgement call, not a derived minimum. If image requests start failing on # context, raise this FIRST and drop the dispatch model to num_ctx 16384 to # pay for it. Budget guards (max_images, max_image_bytes) bound the image # side but not the conversation text around it. # # Do NOT swap this for a 1B-class model to save memory. Tested 2026-09-03: # moondream (1B, 2048 ctx) answered a real dashboard screenshot with "a # spreadsheet ... possibly related to business decisions or financial # analysis" -- it read no title, no column, no value, and confabulated a # plausible description instead. qwen3-vl:4b read the page title, quoted the # subtitle verbatim and recovered the row count. Screenshots are the actual # workload here, and dense small text is exactly where tiny VLMs fail. model: "qwen3-vl-router:4b" timeout_seconds: 60 max_images: 4 max_image_bytes: 9437184 local_dispatch_models: # Local Ollama models that can serve dispatch (not just classification/vision). # Each entry is a model that the router can route traffic to — the same path # that picks cloud models from the NeuralWatt catalog, but calling a local # OpenAI-compatible endpoint instead. # # These models MUST ALSO appear under proficiency.categories in eligible_categories # and MUST have a tier pin under tiering.model_tiers, because src/tier.py # OVERWRITES models.tier on every run (the config override is the only pin). # # To add a model: copy this block and change the values. Keep the comment # describing how to build the model tag (Modelfile recipe), because the num_ctx # baked into the tag is what Ollama's /v1 endpoint actually uses — setting it # per-request through the OpenAI API is not supported by Ollama. - # DEFAULT. Measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5), and a # sound starting point for any 24-32GB card -- the numbers below are real, # not aspirational. It is a default rather than a recommendation only in # the sense that YOUR hardware and YOUR taste in models should win: # docs/local-models.md documents the measurement method so you can swap in # whatever you actually want to run and defend the choice with your own # numbers instead of inheriting these. # # Dormancy under the DEFAULT profile (measured 2026-09-02/03): quality-gated # ranking — the local row must score within objective.quality_tolerance # (0.10) of the cloud leader before its price advantage is even consulted. # Today it does not: 0.767 vs 0.95 on file_summarization (gap 0.183), so # this row is dormant under the default profile. That is by design, not a # bug. Do NOT widen quality_tolerance. It fires as a fallback when the cloud # is unavailable; see docs/routing.md § Local dispatch branch. # # qwen2.5-coder-router:14b serves BOTH local dispatch and classification, # so only one model stays resident. Selected 2026-09-02/03 by measurement: # # file_summarization (n=6, judge-scored -- treat +-0.15 as a tie): # deepseek-v4-flash (cloud) 0.95 <- local costs ~0.18 of quality # qwen2.5-coder:32b 0.833 (22GB: evicts vision, 3.4s classify) # qwen2.5-coder:14b Q8_0 0.817 (19GB, 1.5x slower, gain within noise) # qwen2.5-coder:14b Q4_K_M 0.767 <- chosen # nemotron-mini:4b 0.25 (the plan's original pick; FABRICATED # on 3 of 6, and detected 0/4 bugs) # # classification was 11/14 with 0 hard failures for EVERY variant above -- # quantization and context size changed only speed, never accuracy. # Q4_K_M @ 32k: 1.14s mean. Q8_0 @ 32k: 3.79s, because it needs 24GB and # Ollama spills 17% to CPU (watch the "17%/83%" column in `ollama ps`). # # num_ctx 32768 costs ~4GB of KV cache over 16384 (13GB -> 17GB resident). # That is affordable here ONLY because the vision model was retagged to # num_ctx 8192; together they are 23.2GB of 24GB, with ~800MB spare. On a # smaller card, drop this to 16384 first -- KV cache is the cheapest GB to # reclaim, and the classifier never needs it (classifier.max_input_chars # clamps to 8000 chars, ~2.7k tokens). # # Modelfile recipe (run on an Ollama host): # ollama pull qwen2.5-coder:14b # printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch # ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch model_id: "qwen2.5-coder-router:14b" base_url: "http://localhost:11434/v1" api_key_env: timeout_seconds: 180 context_window: 32768 max_output_tokens: 2048 tier: 1 # Restricted to file_summarization and diff_checking. The router's hard # filter (routing.py) will exclude this model for any other category. Do # not widen this without measuring the new category first -- this model # scored 0/4 at detecting bugs before the task set was fixed, and a # confidently wrong answer is worse than an honest error. eligible_categories: - file_summarization - diff_checking verification: # Structural checks (parse the code, never run it) are free, pure Python and # always on — they need no model and run anywhere, down to an RPi. # This section governs the LOCAL LLM check, which is not free. # # Set local_llm_enabled: false on a host with no usable local inference. # Structural checking and POST /outcome both keep working; only the # refusal/incoherence class of failure stops being caught. local_llm_enabled: true # This check speaks Ollama's NATIVE API (/api/chat with think=False), which # no cloud provider offers, so it is configured SEPARATELY from the # classifier rather than derived from it. Deriving it meant that pointing # classification anywhere else sent these requests to /api/chat. # # Point it at the SAME Ollama as the classifier when that is across a VPN; # leave the model null in that case, since the hosts then match. base_url: "http://localhost:11434" # null means "whatever classifier.model is", correct only while both run on # the same local Ollama — config load REFUSES the null once the hosts differ, # because the fallback would name a model this Ollama has never heard of and # the verifier would fail silently. Stated explicitly here anyway, matching # classifier.model exactly (see that field's comment) so both calls hit the # SAME resident Ollama instance, loaded once at its Modelfile-tagged # context size rather than two separately-sized copies. model: "qwen2.5-coder-router:14b" # Only check answers this large. Measured on real traffic: a local check # costs ~15% of a median 193-token answer, so it would only pay if such # answers failed more than ~15% of the time. At 1,500 completion tokens the # break-even failure rate drops to ~1.9%, which is plausible. Below the # threshold a check costs more than the risk it removes. min_completion_tokens: 600 # The check runs AFTER the response has gone back to the client, so it never # adds its ~6s to anyone's latency. It exists to learn which models fail on # real work, not to gate answers. timeout_seconds: 60 max_output_tokens: 1024 # How far back /outcome looks when a report carries no request_id and no # matching directory. A test run follows the completion that caused it within # seconds, so this is deliberately short: a wide window sweeps in sessions # that finished long ago and makes every report look ambiguous. If more than # one conversation was active inside it, the report is refused rather than # guessed at. outcome_attribution_window_seconds: 120 freshness: stale_after_days: 3 # Router refuses to route to a model whose row is stale/deprecated, # regardless of how good its score would otherwise be. exclude_stale: true exclude_deprecated: true # Allowlisting a model does nothing until a poll ingests it, and the timer # runs every 2 hours -- so a model added through /admin sat invisible until # the next tick, looking like the allowlist had not worked. # # Debounce, not delay: each further edit restarts the clock, so adding five # models one at a time costs one poll, not five. 0 disables the trigger and # leaves the timer as the only path. repoll_after_allowlist_change_seconds: 5.0 database: path: "router.db" classifier: # The local LLM that classifies each task before a cloud model answers it. # "Local" means your hardware, not necessarily this machine — the box with # the GPU is usually not the laptop you are typing on. It speaks plain # OpenAI-compatible chat completions, so point it wherever Ollama lives: # # same machine base_url: http://localhost:11434/v1 # over a VPN base_url: http://:11434/v1 # (serving host needs deploy/ollama-over-vpn.conf; Ollama # binds loopback-only by default and will refuse) # # Any OpenAI-compatible endpoint works, so a cloud model can classify too — # set api_key_env for one that checks a key. Worth knowing before assuming # local is the cheap option: on five prompts NeuralWatt's deepseek-v4-flash # classified in 1.02s mean against qwen3.5's 11.58s on an RTX 6000, agreed # with the label 5/5 against 2/4, and cost $0.093 per thousand calls. Local # inference is not free, it is unbilled. provider: "ollama" base_url: "http://localhost:11434/v1" # Unset means unauthenticated, which is the Ollama case. Name the env var # holding the key when the endpoint actually checks one. api_key_env: # A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag # — must match a model `ollama list` reports. Ollama loads a model at its # library Modelfile's default context unless told otherwise, and the base # tag never was: measured live, mistral-nemo:12b came up at num_ctx=32768 # (14GB of a 24GB card) though max_input_chars + the system prompt + # max_output_tokens need well under a quarter of that. num_ctx cannot be # set per-request here: verified live that Ollama's OpenAI-compatible # endpoint (0.22.0) silently ignores it under every field shape tried (a # 200 comes back, the loaded context never changes), so it has to be baked # into the tag itself: # printf 'FROM mistral-nemo:12b\nPARAMETER num_ctx 8192\n' > Modelfile.router # ollama create mistral-nemo-router:12b -f Modelfile.router # 8192 leaves 2x headroom over the real requirement. verification.model # points at this same tag (see its own comment) so classify and verify # share one resident instance at one context size, rather than risking a # reload thrash from two differently-sized copies of the same base model. model: "qwen2.5-coder-router:14b" # A cold Ollama took 43s to answer the first classification, which blew the # old 30s ceiling and returned 503 to the client. Warm it is ~5s. The # ceiling is for a cold model load, not the steady state. timeout_seconds: 120 # Classification must be reproducible: at the default temperature the same # task was classified tier 2 then tier 1 on consecutive calls, which routed # it to two different models. Routing that changes under an identical # prompt is untraceable. temperature: 0 # Hard cap on the classifier's generation, to bound a failure mode that # cascades on REASONING models: they emit a long chain of thought, blow past # timeout_seconds, and Ollama keeps generating after the client gives up AND # serializes per model, so one runaway request queues every later one behind # it. Measured on qwen3.5, which failed this way on 4 of 19 calls — each one # a 15s wait ending in a silent fallback. # # The current default (mistral-nemo) does not reason, so this does not bind # for it. Kept anyway: it costs nothing when unused and is the only thing # standing between a swapped-in reasoning model and that cascade. If you do # swap one in, note 256 was too tight — the trace consumed the whole budget # and the model was truncated before emitting any JSON. max_output_tokens: 1024 # Where routing lands when the classifier times out, errors, or returns # something unparseable. A local model being slow should degrade routing, # not refuse the request — the caller is a coding agent that would rather # have a mid-tier answer than a 502. # Ceiling on the text handed to the classifier (head + tail, middle elided). # It decides a category and a tier; it does not need the document, and # feeding it one is harmful rather than merely wasteful. Measured on a ~20k # token prompt: qwen3.5 spent 28.7s and returned empty (budget consumed by # its reasoning trace), mistral-nemo spent 41.8s echoing the input back # inside its JSON. Both land on source: "fallback" — the same answer an # instant failure gives, after 30-40s of local inference. # # Nothing is lost: chat_completions measures the real conversation with # estimate_prompt_tokens and takes the larger value, so required_context # never depends on what the classifier saw. 0 disables clamping. max_input_chars: 8000 # When the chat path supplies the previous turn as context (see the pinch / # context notes), frame the classifier input as "Context: / Message: # " so a short follow-up inherits the prior turn's complexity # instead of being classified in isolation as trivial. context_framing: true fallback_tier: 2 fallback_category: general_chat # --- which implementation is PRIMARY ----------------------------------- # This is a peer concept to the fallback cascade below, not a replacement # for it: whichever mode is primary, a failure still walks the SAME # cascade (stale session -> session history -> cloud_fallback -> the # fallback_tier/fallback_category above). # # local_llm (default) the model/base_url above, unchanged. # cloud_llm a cloud model answers PRIMARY, not just as a fallback # after local fails. Requires exactly one of # cloud_primary or cloud_primary_auto below. # local_encoder a small, non-generative classifier (see the encoder: # block below) that cannot exhibit the runaway-reasoning # failure mode documented above, at the cost of only # producing task_category -- task_tier falls back to # fallback_tier for this mode. Zero-shot, not trained on # your traffic: this router never stores raw task text. mode: local_llm # Read only when mode: cloud_llm. Exactly one of these two: # cloud_primary: # base_url: https://api.neuralwatt.com/v1 # model: deepseek-v4-flash # api_key_env: NEURALWATT_API_KEY # timeout_seconds: 2 # max_output_tokens: 1024 # cloud_primary_auto: true # resolve the cheapest routable model, live # Read only when mode: local_encoder. Every field has a default, so # `encoder: {}` is enough to opt in. # encoder: # model: facebook/bart-large-mnli # device: cpu # confidence_threshold: 0.5 # --- fallback cascade --------------------------------------------------- # When the local classifier fails, the router walks: stale session cache -> # this session's history in route_decisions -> the optional cloud # classifier below -> fallback_tier/fallback_category above. The first two # steps are free and local. # # Global backoff after a classifier failure. This is what bounds cloud # spend during a sustained outage: at most one cloud attempt per window # across ALL requests, not one per request. It is also reused as the window # in which an account-level provider refusal suppresses the cloud step, # since an out-of-credit account makes that call a guaranteed waste. cooldown_seconds: 30 # Degradation warning on /metrics. Silent below degraded_warn_min # decisions in 24h, because a share computed over a handful of requests is # noise and a flapping warning is one nobody reads. degraded_warn_min: 20 degraded_warn_threshold: 0.5 # Optional. ABSENT BY DEFAULT, which is what makes the cascade cost # nothing: with no cloud_fallback block the router degrades straight to the # static guess. Uncomment and point it at any OpenAI-compatible endpoint to # trade a little money for a real classification during a local outage. # A configured api_key_env whose variable is missing degrades to the static # guess rather than failing the request -- unlike the primary classifier, # which raises, because this one is only reached when things already broke. # cloud_fallback: # base_url: https://api.neuralwatt.com/v1 # model: deepseek-v4-flash # api_key_env: NEURALWATT_API_KEY # timeout_seconds: 2 # max_output_tokens: 1024 response_format: "json" # ask Ollama to constrain output to valid JSON # The labels a classifier may return. A SUBSET of proficiency.categories, # validated as one at config load, and deliberately not the same list: # proficiency.categories is the scoring axis (what a model is good at), # this is the set a classifier is asked to choose between. # # tool_use_agentic is excluded. It describes what a turn mechanically DOES # rather than what it is for, and an agent turn is always both -- measured # 2026-09-14 on live traffic with the session cache off, 31 of 31 # consecutive turns classified tool_use_agentic and every one routed to the # slowest model in the catalog (14.8s TTFT, p95 40s). Per-turn # classification became accurate and routing got worse. The signal it was # standing in for is already read exactly and for free from the request's # own `tools` array by routing.min_tool_proficiency. # # WHAT THIS COSTS, so the next person does not discover it: with the # classifier unable to emit tool_use_agentic, no new POST /outcome report # attributes to that category, so its proficiency scores FREEZE at their # current values (e.g. qwen3.6-35b at 0.902 over 154 samples). The tool # filter keeps reading them; they simply stop accumulating. Accepted for # now. # # Considered and REJECTED as the fix: attributing outcomes to # tool_use_agentic whenever the request carried a `tools` array. opencode # sends `tools` on essentially every request, so the score would converge # on each model's overall pass rate and stop discriminating the exact thing # the category exists to detect -- a model that OVER-reaches for tools on # work that did not need them. A frozen honest score beats a live # meaningless one. # # Omit this key entirely to mean "every proficiency category". candidate_categories: - coding_general - coding_refactor - debugging - docs_writing - summarization - file_summarization - diff_checking - translation - reasoning_math - general_chat # The dispatcher appends the authoritative category list from # classifier.candidate_categories (above) to this prompt at call time. Do # not enumerate the categories here as well — a hand-copied list drifts, and # a category the model invents joins against nothing in the proficiency # table. system_prompt: | You are a task router. Given a task description and any attached context, respond with ONLY a JSON object with these fields: { "task_category": one of the allowed categories listed below, "task_tier": integer 1-3, where 1 is cheap/simple, 2 is mid/general, and 3 is frontier/high-stakes, "required_context_tokens": integer estimate of prompt+context token count, "confidence": float 0-1 } dispatch_providers: neuralwatt: base_url: "https://api.neuralwatt.com/v1" api_key_env: "NEURALWATT_API_KEY" has_energy_telemetry: true # false: NeuralWatt reports its bill in a top-level `cost` block (and as a # `: cost {...}` SSE comment when streaming), so no request-side opt-in is # needed and none is sent. See the openrouter entry below for what the # flag does when it is on. reports_cost_in_usage: false enabled: true openrouter: base_url: "https://openrouter.ai/api/v1" api_key_env: "OPENROUTER_API_KEY" # Account-level balance poll; OpenRouter reports per-completion energy via # a separate allowance_remaining field, so this URL is for the prepaid pool. balance_url: "https://openrouter.ai/api/v1/credits" has_energy_telemetry: false # true: send OpenRouter's own `usage: {"include": true}` opt-in on every # upstream request. `stream_options.include_usage` is the OpenAI spelling # and gets token counts; this is the separate OpenRouter one, and without # it the accounting block is simply absent -- no billed `usage.cost` and # no `usage.prompt_tokens_details.cached_tokens`. Both ride on this one # opt-in, so the two coverage figures move together. # # It does NOT bring generation timing: OpenRouter serves that only from # its separate /api/v1/generation endpoint, so `duration_seconds` stays # NULL on OpenRouter rows. Turn this off only for a provider that returns # accounting unasked, or one that rejects the key. reports_cost_in_usage: true enabled: true require_allowlist: true local_compute: # "Gaming mode", inverted: true means the router may use local hardware. # Set it to false when you stop Ollama to give the GPU back to something # else, and the router SKIPS every local call instead of discovering the # outage one 120s timeout at a time: classifier, /health probe, local # verification, the local-vision fallback, and local dispatch rows (which # drop out of the candidate set rather than being picked and then failing). # # ONE flag that the call sites read, not a macro that rewrites # verification.local_llm_enabled / local_vision.enabled / local_energy.enabled. # Those keep their own meanings; this is an outer AND over all of them, so # turning it back on restores exactly the state you left. # # Config load REFUSES false unless classifier.cloud_fallback is configured. # Skipping the local classifier does not make classification remote, it makes # it a static guess -- and that guess is recorded as general_chat, a fully # scored category, so it looks like real classification afterwards. # Uncommenting the cloud_fallback example above is the intended setup path. enabled: true local_energy: # Local energy metering for the router's own Ollama calls (classifier, # verifier, local-vision fallback). Cloud providers expose per-request energy, # but these calls run on your own hardware and their electricity is real even # though it never shows up on the NeuralWatt bill. # # OFF by default. Enable only after setting a real per-kWh rate below. # Config load REFUSES enabled: true with a null tariff because a null rate # would record cost_usd = 0 instead of "unknown". enabled: false # Power sampler. Only "nvidia_smi" is supported today. It shells out rather # than importing pynvml so the router has no new dependency and degrades # cleanly on non-NVIDIA hosts. meter: "nvidia_smi" # How often to sample GPU power while a local call is running. 0.25s was the # original llmrouter pinch sample rate and is fine as a starting point. sample_interval_seconds: 0.25 # YOUR real per-kWh electricity rate. Set this before enabling metering. # # Example (do not use this number unless it matches your bill): # tariff_usd_per_kwh: 0.12 # # Set 2026-09-02 from the user's Great Lakes Energy bill. This is the # MARGINAL rate -- what one more kWh actually costs -- not the all-in # effective rate of 0.183 ($410.15 / 2,244 kWh). The difference is the # $52.30/month of fixed charges, which are paid whether or not the GPU # runs, so loading them onto compute would overstate local cost by ~15% # and bias routing toward the cloud for a reason that is not real. # 14.742c energy + 0.188c PSCR + 0.316c EO = 15.246c, x 1.04 MI sales # tax = 15.856c/kWh. tariff_usd_per_kwh: # YOUR grid carbon intensity in grams of CO2 equivalent per kWh. Optional: # null disables the carbon estimate but still records energy and cost. # # Example (do not use this number unless it matches your grid): # grid_intensity_g_per_kwh: 475 grid_intensity_g_per_kwh: logging: # Whether to write a row to energy_observations for every completion. Off # means no cost accounting, no reference sweep and no /outcome attribution, # so leave it on unless you are debugging. # # There is no log_path. Nothing ever wrote a file -- the dispatcher prints to # stderr and the systemd unit hands that to the journal (`journalctl --user # -u llm-router -f`), so the setting named a destination that did not exist. log_energy_observations: true # Whether to write a row to route_decisions for every routing decision # (route | dispatch | chat | passthrough | local-vision). Off leaves the # monitoring TUI's decision history empty; it does not change what routes. log_route_decisions: true # debug | info | warning | error. # # info gives one line per request: what it was classified as, which model won, # what it cost, how long each stage took. debug adds why -- every candidate # that was dropped and by which filter, the ranking with scores, the # classifier's raw reply. It logs no conversation text at any level; prompts # here run 60k-150k tokens and the journal is on disk. # # LLM_ROUTER_LOG_LEVEL overrides this, so a running service can be turned up # without editing a tracked file: # # systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug" # systemctl --user restart llm-router level: info