At 32k the 14b is 16.5-17.8 GB resident and the local_decision classifier (qwen3.5:4b, 5.9 GB at num_ctx 8192) evicts it; at 16k it is 12.26 GB and the pair sits at ~19.9 of 24 GB (plans/local-decision-classifier-results.md, sections 7 and 9). context_window follows the tag, which the operator rebuilds at 16384 after this merges; the config must shrink first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
1142 lines
58 KiB
YAML
1142 lines
58 KiB
YAML
# Local LLM router config
|
|
# All weights, thresholds, and provider settings live here so they can be
|
|
# tuned without touching code. Loaded/validated by config.py.
|
|
|
|
objective:
|
|
# Quality is the objective. Cost is a constraint and a tiebreak. Eco is
|
|
# logged per request but is NOT optimized here — that judgement is made
|
|
# outside this router.
|
|
#
|
|
# This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4).
|
|
# Measurement killed it: turning the cost weight from 0.4 to ZERO changed
|
|
# the winner in only 2 of 6 categories, so the blend was never steering on
|
|
# quality — cost decided almost nothing while consuming 40% of every decision.
|
|
#
|
|
# The old comment cited $0.07 total spend as evidence it was "fractions of a
|
|
# cent" — that figure expired long ago. /metrics now surfaces current all-time
|
|
# spend via cumulative_spend_series() and warns via cumulative_spend_warnings()
|
|
# once it has grown past the tiebreak's "harmless fractions of a cent" frame,
|
|
# so the justification is checked live rather than embedded as a stale number.
|
|
|
|
# Proficiency differences smaller than this are treated as equal and the
|
|
# cheaper model wins. This is measurement noise, not preference: scores once
|
|
# rested on 2-3 samples per category, and a 0.05 gap is indistinguishable
|
|
# from sampling variation when the evidence is that thin.
|
|
#
|
|
# /metrics now enforces the "narrow it" promise: once average sample depth
|
|
# (outcome_samples + self_eval_samples) crosses the threshold in
|
|
# proficiency_sample_depth_warnings(), the premise has expired and the warning
|
|
# fires — so the loose tolerance is called out rather than silently prolonged.
|
|
# That threshold defaults to 20; see objective.proficiency_depth_warn_min_samples.
|
|
# Above this pace ratio, the dashboard raises an alarm.
|
|
# 1.25 = burning 25% faster than the plan allows.
|
|
plan_pace_warn_ratio: 1.25
|
|
quality_tolerance: 0.1
|
|
|
|
# Cost is priced per-request from catalog prices, NOT from a benchmark
|
|
# sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper
|
|
# than deepseek-v4-flash; on a realistic 70k-token prompt deepseek is 5.0x
|
|
# cheaper. Attribution inverts with prompt size, so a fixed-shape benchmark
|
|
# cannot rank models for a workload of another shape. Catalog prices scaled
|
|
# to the actual request agree with the live measurement, cost nothing, and
|
|
# need no sweep.
|
|
|
|
# Share of prompt tokens served from the provider's prefix cache. Agent
|
|
# clients resend the whole conversation each turn, so most of it is a hit.
|
|
#
|
|
# Measured token-weighted across 50 sessions and 40.7M tokens (2026-08-23):
|
|
# 91.7% overall, 92.6% on the sessions above 400k tokens, which are the ones
|
|
# that carry the cost. The previous 0.84 came from 2.2M tokens of much
|
|
# earlier traffic.
|
|
#
|
|
# Raising it changed NO winner in any of the nine categories at 60k context
|
|
# — checked before editing. It is here because the number should be true,
|
|
# not because the routing needed it.
|
|
assumed_cache_rate: 0.917
|
|
|
|
# Completion length assumed when pricing a request. Real sessions here median
|
|
# around 200-400 completion tokens against enormous prompts.
|
|
assumed_completion_tokens: 500
|
|
|
|
# Per-request ceiling on measured ENERGY, in kWh. null disables it.
|
|
#
|
|
# Denominated in kWh rather than dollars. For scale: the reference task runs
|
|
# ~5e-06 kWh on the cheapest model and ~2.2e-04 on the most expensive.
|
|
# Overage is billed against the account's credit balance;
|
|
# plan_kwh_per_period gates nothing.
|
|
max_energy_per_request:
|
|
|
|
# The subscription's kWh allowance per billing period, for reporting burn in
|
|
# /health. This is a planning figure only: per-request traffic is never
|
|
# refused for exceeding plan_kwh_per_period — it gates nothing. Set to match
|
|
# your plan; null disables the report. NeuralWatt also returns
|
|
# allowance_remaining_usd per request, which is logged for /metrics, but that
|
|
# is a dollar figure while the plan is denominated in energy.
|
|
plan_kwh_per_period: 6.25
|
|
|
|
# Hours of recent balance history used to estimate the burn rate. The most
|
|
# recent monotonically-decreasing segment of allowance_remaining_usd values
|
|
# (segments split at each balance increase) is examined over this window.
|
|
quota_burn_window_hours: 24
|
|
# Projected hours of runway below which /metrics and dashboards emit a
|
|
# low-warning boolean; only fires when a burn rate estimate exists and the
|
|
# projected remainder is positive but short.
|
|
quota_runway_warning_hours: 6
|
|
# A burn estimate needs at least this many balance samples in the most recent
|
|
# monotonically-decreasing segment (post-top-up resets the segment); below
|
|
# which the estimate is None with an explanatory runway note.
|
|
quota_burn_min_segment_samples: 3
|
|
# ...and the segment must span at least this many hours, otherwise the
|
|
# estimate is None with an explanatory note (never a wild extrapolation).
|
|
quota_burn_min_segment_hours: 0.5
|
|
|
|
# The day-of-month your NeuralWatt subscription billing cycle resets. Set
|
|
# this to YOUR real billing day so the admin quota modal shows a genuine
|
|
# next-reset date instead of a misleading rolling-window start. null disables
|
|
# the feature and the modal shows "not configured". Valid range: 1-28 (skip
|
|
# 29-31 to avoid shorter-month edge cases).
|
|
billing_reset_day: 6
|
|
|
|
# Rejection-rate detection over route_decisions rows that selected no model
|
|
# (422 "no model satisfies the hard filters"). The signal is novelty or rate —
|
|
# NEVER mere presence: this deployment routinely has ~3 rejections/hour of
|
|
# ordinary over-large tier-3 requests that are behaving as designed (6 in the
|
|
# last 24h, 17 all-time at 2026-09-05), so a count-only tripwire would be
|
|
# permanently on.
|
|
#
|
|
# alert window (hours): rejections this recent are counted per
|
|
# (task_tier, digit-normalized reason) group.
|
|
rejection_warning_window_hours: 1
|
|
# baseline window (hours) BEFORE the alert window: groups absent here are
|
|
# "new". A new group warns from 2 occurrences — the 2026-09-04 vision
|
|
# incident produced exactly 2 and no rate threshold can sit below routine
|
|
# noise yet above that.
|
|
rejection_warning_baseline_hours: 24
|
|
# a group already present in the baseline warns only at this many
|
|
# occurrences inside the alert window: 6 = 2x the observed routine hourly
|
|
# peak, far below the dozens/hour a deprecation flare produces.
|
|
rejection_warning_min_count: 6
|
|
|
|
# How far back /metrics looks to decide which models the router has actually
|
|
# picked. Long on purpose: this answers "has this model EVER been chosen",
|
|
# and a model that only wins one category can go days between selections.
|
|
# Feeds two very different reports -- see metrics.selection_coverage.
|
|
selection_coverage_window_hours: 168
|
|
|
|
# Prefix-cache rate measured from energy_observations, and the expiry check
|
|
# on assumed_cache_rate above. That constant was measured ONCE, on
|
|
# 2026-08-23, and it is the highest-leverage term in routing.estimated_cost
|
|
# on a 100k-token prompt — a premise that large should not go unchecked just
|
|
# because it was true when it was written.
|
|
#
|
|
# The comparison is a DIVERGENCE from assumed_cache_rate, not an absolute
|
|
# floor. "Is the cache working" is the wrong question; a deployment whose
|
|
# real rate is 0.60 is not broken, it is mispriced, and a floor would say
|
|
# nothing about the number the cost model is actually built on.
|
|
#
|
|
# Trailing window, in hours. Long for the same reason
|
|
# selection_coverage_window_hours is: only a fraction of rows carry a
|
|
# reported cached count, so a 1h window is usually empty.
|
|
cache_rate_window_hours: 168
|
|
# Absolute divergence from assumed_cache_rate that warns. 0.10 sits well
|
|
# outside ordinary session-mix drift (measured 2026-09-13: 0.879 neuralwatt
|
|
# / 0.901 openrouter against the assumed 0.917) and well inside the 0.952 vs
|
|
# 0.478 same-model/switched gap the waves plan is chasing.
|
|
cache_rate_warn_margin: 0.10
|
|
# Minimum reported-cache observations before either warning fires, per
|
|
# aggregate and per (provider, model) group. Observations, not tokens: one
|
|
# 200k-token prompt outweighs a hundred ordinary turns, so a token floor
|
|
# would let a single request's luck read as a measurement.
|
|
cache_rate_warn_min_observations: 25
|
|
|
|
# Premise-expiry check for quality_tolerance's "2-3 per category" claim.
|
|
# /metrics warns when average proficiency sample depth exceeds this — the
|
|
# point where the old thin-data justification is clearly obsolete and the
|
|
# loose tolerance should be narrowed.
|
|
proficiency_depth_warn_min_samples: 20
|
|
# Minimum proficiency rows before the depth check fires. Prevents a warning
|
|
# from a near-empty table (novelty-or-rate convention).
|
|
proficiency_depth_warn_min_rows: 10
|
|
|
|
# Premise-expiry check for the cost-as-tiebreak justification.
|
|
# /metrics warns when all-time SUM(cost_usd) exceeds this — the point where
|
|
# the "fractions of a cent / $0.07" claim is no longer the honest frame.
|
|
cumulative_spend_warn_usd: 50.0
|
|
# Minimum priced rows before the spend check fires. Prevents a warning from
|
|
# a DB with no priced rows.
|
|
cumulative_spend_warn_min_rows: 10
|
|
|
|
# --- Report-only measurement series. Neither is read by routing. ---------
|
|
#
|
|
# Cost-estimator calibration: routing.estimated_cost's prediction against
|
|
# the provider's own billed figure, per (provider, model), joined on
|
|
# request_id. Measured 2026-09-12 at ~100% telemetry coverage, the estimator
|
|
# is high by 1.6x-13x depending on the model.
|
|
#
|
|
# The scale error is NOT the finding — a uniform overestimate reorders
|
|
# nothing, because the ranking is a comparison and every candidate moves
|
|
# together. The SPREAD in that error is the finding: it priced
|
|
# deepseek/deepseek-v4-flash below qwen3.6-35b while the bill said the
|
|
# reverse. /metrics reports `spread` for exactly that reason.
|
|
#
|
|
# Nothing applies these factors. They are here so their stability can be
|
|
# judged first — this project has already mistaken one moment of a moving
|
|
# per-model quantity for a constant (see CLAUDE.md, "Attribution drifts
|
|
# across hours, so sampling must too").
|
|
#
|
|
# Trailing window, in hours. Long for the same reason
|
|
# cache_rate_window_hours is: only rows with both an estimate and a billed
|
|
# figure qualify, and a short window is usually too thin to read.
|
|
cost_calibration_window_hours: 168
|
|
# Joined observations before a (provider, model) factor is marked
|
|
# `sufficient` and allowed to set the reported spread. Groups below it are
|
|
# still listed with their counts — the count is itself information — but a
|
|
# ratio over three requests must not become the headline.
|
|
cost_calibration_min_observations: 25
|
|
|
|
# Router-observed latency: router_wall_seconds and router_ttft_seconds from
|
|
# energy_observations, p50 and p95 per (provider, model). These are the
|
|
# ROUTER's clock, not the provider's duration_seconds — which is why this
|
|
# can see OpenRouter at all, since OpenRouter reports no duration.
|
|
#
|
|
# It found z-ai/glm-5.3-flash, a model with "flash" in its name, at a p50 of
|
|
# 9.6s to first token against 1.6s for deepseek-v4-flash on NeuralWatt.
|
|
# Reported, not scored: latency is not an objective here, and one reading of
|
|
# a quantity that tracks pool load is not grounds to make it one.
|
|
latency_window_hours: 168
|
|
# Observations before a group's percentiles are marked `sufficient`. Applied
|
|
# SEPARATELY to the wall and TTFT counts, because TTFT is streaming-only by
|
|
# nature and a buffered deployment legitimately has fewer of them.
|
|
latency_min_observations: 25
|
|
|
|
# Gate: with this off, routing is byte-identical to today.
|
|
# Flip on only after the Wave 1 post-restart baseline day;
|
|
# see plans/token-waste-waves.md Wave 2 gate.
|
|
incumbent_cache_pricing: false
|
|
|
|
# Challenger cache-rate dial.
|
|
# - null (default): neutral, follows assumed_cache_rate → today's ranking exactly
|
|
# including the incumbent; this is the off position for tuning.
|
|
# - 0.0: challengers priced as fully cold prompts (maximum incumbent advantage).
|
|
# - Any value between is a partial cache-penalty — the whole wave dials
|
|
# from off to full here, no revert needed.
|
|
incumbent_challenger_cache_rate:
|
|
|
|
# Seconds between refreshes of the measured per-(provider, model) cache
|
|
# rates. The TTL is measured from the first call after the process started
|
|
# (time.monotonic), so a brief post-restart cold period is expected.
|
|
incumbent_rate_refresh_seconds: 300
|
|
|
|
# Minimum reported-cache observations before a per-(provider, model)
|
|
# cache rate is trusted for pricing decisions — the independent pricing
|
|
# floor. Tuning this does NOT move the cache-rate warning floor
|
|
# (cache_rate_warn_min_observations): the two knobs answer different
|
|
# questions with different failure costs.
|
|
#
|
|
# Shipped with a documented default (25) so the code always reads a
|
|
# concrete float when the penalty is engaged.
|
|
incumbent_rate_min_observations: 25
|
|
|
|
# Window (seconds) for the conversation adoption counter in /metrics.
|
|
# Without this, the counter may report stale adoption figures from a period
|
|
# when the feature was still rolling out. Null or absent = no windowing
|
|
# (all time); 0 is rejected by the config validator. For a rolling 7-day
|
|
# window: 604800.
|
|
adoption_window_seconds: 604800
|
|
|
|
# Credit-aware routing attenuation (OFF BY DEFAULT).
|
|
# When enabled, a provider configured with a balance_url (e.g. OpenRouter's
|
|
# prepaid account balance) gets its *comparison cost* inflated inside the
|
|
# quality-first ranker as its account balance nears zero. Quality bands still
|
|
# win; this only shifts ties. Providers whose balance comes from per-request
|
|
# energy telemetry allowance_remaining_usd (e.g. NeuralWatt's overage-billed
|
|
# subscription) are ALWAYS multiplier 1.0 regardless of their reading, so a
|
|
# low soft_floor_usd never biases routing toward an attenuated provider just
|
|
# because a telemetry provider's allowance reads near zero under normal use.
|
|
#
|
|
# This knob is deliberately NOT exposed in the admin UI persisted-config
|
|
# allowlist (_CONFIG_ALLOWLIST in admin.py) or _ProviderUpdateBody; enabling
|
|
# or tuning it requires editing this file and restarting llm-router.service
|
|
# (the dispatcher's module-level cfg binds at import).
|
|
credit_attenuation:
|
|
enabled: false
|
|
soft_floor_usd: 5.0 # balance >= this -> multiplier 1.0
|
|
zero_floor_usd: 0.0 # balance <= this -> max_multiplier
|
|
max_multiplier: 5.0 # maximum cost-inflation at/below zero floor
|
|
refresh_seconds: 300 # cache duration for resolved multipliers
|
|
|
|
context:
|
|
safety_factor: 0.75 # fraction of advertised context treated as usable
|
|
default_output_reserve_tokens: 4096
|
|
# Ceiling on the output reserve, as a fraction of the usable window. A
|
|
# provider's advertised max_output_tokens is normally a small per-request
|
|
# cap, but OpenRouter reports max_completion_tokens -- "the most you may
|
|
# ASK for", 0.8-0.9 of context on a dozen rows. Subtracting that whole
|
|
# left 12 models at an effective context of 0, silently unroutable.
|
|
# Only the catalog-derived reserve is capped; a per_model_overrides
|
|
# reserve is a measurement and is used as written.
|
|
max_output_reserve_fraction: 0.5
|
|
# Per-model exceptions to the two settings above, for a row whose real
|
|
# limits you have measured. The global factor has to hold for the whole
|
|
# catalog, so it is deliberately pessimistic; a model you have actually
|
|
# pushed to its limit deserves its own number.
|
|
#
|
|
# Both keys are optional and each falls back to the global on its own.
|
|
# An override of 0 reserve tokens means zero, not "unset".
|
|
#
|
|
# per_model_overrides:
|
|
# qwen3.6-35b:
|
|
# safety_factor: 0.85
|
|
# output_reserve_tokens: 8192
|
|
#
|
|
# Takes effect on the next `python poller.py` -- effective_context_window is
|
|
# computed at poll time, not per request.
|
|
per_model_overrides: {}
|
|
|
|
tiers:
|
|
# Maps a tier number to a human label, purely for logging/dashboards.
|
|
1: "cheap / simple"
|
|
2: "mid / general"
|
|
3: "frontier / high-stakes"
|
|
|
|
tiering:
|
|
# Auto-tiering pass knobs. cheap_completion_max is the completion-cost
|
|
# (per 1M tokens) ceiling below which a model is eligible for tier 1.
|
|
# model_tiers overrides the heuristic per model_id and applies to ALL
|
|
# providers (limitation vs a (model_id, provider) key).
|
|
cheap_completion_max: 1.00
|
|
|
|
# Advertised context_window at or above which a model is NOT eligible for
|
|
# tier 1, whatever it costs. Tier is a capability FLOOR (routing drops any
|
|
# row with tier < required_tier), so tier 1 means "simple work only" — and
|
|
# deciding that on price alone excluded deepseek-v4-flash from every tier-2
|
|
# request purely for being $0.28/1M, despite a 1M window and 1.00 on all
|
|
# three coding categories.
|
|
#
|
|
# 512000 sits in the empty band between the catalog's 256K class (262128)
|
|
# and its 1M class (1048560) — a 2x margin either side, so it is not fitted
|
|
# to any one model. gemma-4-31b (256K) stays tier 1; the 1M rows do not.
|
|
tier1_context_max: 512000
|
|
|
|
model_tiers:
|
|
# This model lives on Ollama, not the NeuralWatt cloud catalog, so the
|
|
# tiering heuristic never sees it -- src/tier.py OVERWRITES models.tier on
|
|
# every run, so this pin is the ONLY thing that survives. If you swap the
|
|
# local dispatch model, change the KEY here too or it silently loses its
|
|
# tier. Pinned to tier 1 (cheap / simple) because it is targeted at
|
|
# lightweight summarization and diff-checking — precisely the use-cases
|
|
# tier 1 was designed for.
|
|
qwen2.5-coder-router:14b: 1
|
|
|
|
proficiency:
|
|
# Blending rule: leaderboard vs self-eval, once self-eval sample size
|
|
# crosses the threshold below. Below threshold, leaderboard score alone
|
|
# is used so thin self-eval data doesn't dominate.
|
|
self_eval_min_samples: 10
|
|
leaderboard_weight: 0.3
|
|
self_eval_weight: 0.7
|
|
# Prior strength (k) for the empirical-Bayes shrunken estimate: the peer
|
|
# prior contributes k pseudo-observations, pulling each noisy per-model
|
|
# score toward the global average. 20 pseudo-observations damps thin
|
|
# self-eval data (n < 100 per model) without erasing the per-model signal.
|
|
outcome_prior_strength: 20
|
|
categories:
|
|
- coding_general
|
|
- coding_refactor
|
|
- debugging
|
|
- docs_writing
|
|
- summarization
|
|
# file_summarization + diff_checking are served by the local dispatch model
|
|
# (see local_dispatch_models: section below).
|
|
# They are kept adjacent to summarization since this model is targeted
|
|
# at lightweight summarization and diff-checking tasks.
|
|
- file_summarization
|
|
- diff_checking
|
|
- translation
|
|
- reasoning_math
|
|
- tool_use_agentic
|
|
- general_chat
|
|
|
|
exploration:
|
|
# Epsilon-greedy exploration: on the epsilon share of requests, picks the
|
|
# hard-filter-eligible candidate with the fewest outcome samples (tie-break:
|
|
# lowest cost) instead of the highest-score model. Deliberately ON by default
|
|
# because the system is a ranking loop — without periodic exploration it
|
|
# converges on whatever happens to be sampled most, creating exposure bias.
|
|
# Adjust once real exploration traffic (was_exploration=True in
|
|
# route_decisions) shows how often the exploration path picks differently
|
|
# from the greedy path.
|
|
enabled: true
|
|
# Probability of exploration per request. 0.03 = ~3% of requests take an
|
|
# exploratory path, giving ~97% exploit on known winners while still
|
|
# occasionally sampling under-explored candidates.
|
|
epsilon: 0.03
|
|
# Cap: exploration candidate cost must be <= this multiple of the ranking
|
|
# winner's cost. 4.0 gives headroom — an under-sampled model can be more
|
|
# expensive than the greedy winner without eating the explore budget on
|
|
# wildly off-target picks.
|
|
max_cost_ratio: 4.0
|
|
# Only tiers 1 and 2 models are eligible for exploration. Tier 3 (frontier)
|
|
# is too expensive to spend on random sampling — save the explore budget for
|
|
# models where being wrong costs less.
|
|
max_tier: 2
|
|
|
|
escalation:
|
|
enabled: true
|
|
max_tier: 3
|
|
|
|
# Bump the tier when the classifier is unsure of its own call. DEFAULT OFF:
|
|
# this pays frontier prices on a hunch, before anything has gone wrong. The
|
|
# iteration budget below spends after a check has actually failed, which is
|
|
# strictly better on both mandates — the cheap attempt usually succeeds and
|
|
# costs nothing extra, and when it fails you have evidence.
|
|
preemptive_on_low_confidence: false
|
|
min_confidence_before_bump: 0.6
|
|
|
|
iteration:
|
|
# A tier is not only a capability floor, it is a budget for getting the
|
|
# answer right. These are corrective attempts AFTER a verification failure,
|
|
# not speculative retries.
|
|
#
|
|
# Retries are matched to the failure: a truncated answer gets a bigger token
|
|
# budget on the SAME model (a different one would also run out), while a
|
|
# malformed answer first retries the same model (cache-preserving — ~0.92
|
|
# hit rate vs ~0.35 on switch) before escalating to the next-best candidate.
|
|
enabled: true
|
|
attempts_by_tier:
|
|
1: 0 # cheap/simple — one shot; iterating costs more than it is worth
|
|
2: 1
|
|
3: 2
|
|
|
|
# Interactive requests are capped below their tier's budget regardless of
|
|
# tier: every retry doubles time-to-answer, and in interactive use latency
|
|
# IS a quality loss. Batch work does not care.
|
|
max_attempts_interactive: 1
|
|
|
|
# Maximum conversation prompt-tokens for which escalation (cache-destroying
|
|
# re-bill on a different model) is allowed. Same-model retries preserve the
|
|
# provider's prompt cache and are not gated — the marginal cost is only the
|
|
# answer generation, not a full prompt re-bill. 0 = no limit (backward-
|
|
# compatible default). Set to e.g. 65536 to suppress cold-model escalation
|
|
# for conversations whose prompt is larger than 64K tokens.
|
|
max_rebill_prompt_tokens: 0
|
|
|
|
pinch:
|
|
# Relevance-based context pruning (ported from the MIT-licensed llmrouter's
|
|
# "pinch"). This is an OPTIONAL, pre-dispatch stage: when a conversation
|
|
# exceeds budget_tokens, the provider-bound messages are trimmed BEFORE any
|
|
# paid token is sent upstream. User/assistant/system messages are always
|
|
# kept verbatim; only old TOOL RESULTS are shortened or dropped, because
|
|
# they carry the bulk of a long agent session's tokens and are least needed
|
|
# in full by the time the next turn is answered.
|
|
#
|
|
# On by default. Tool results can only be dropped safely because tool
|
|
# outputs are idempotent enough for a placeholder; a wrong guess here loses
|
|
# context, so start conservative (large budget, small reduction) and watch
|
|
# route_decisions / pinch stats on real traffic before widening it.
|
|
enabled: true
|
|
budget_tokens: 50000
|
|
keep_last_turns: 4
|
|
max_summarize_chars: 4000
|
|
# keep_last_turns has no size limit inside it: an entire autonomous
|
|
# tool-call loop with no new user message can be one protected turn, and
|
|
# one outsized tool result inside it (a full verbose test run, a huge file
|
|
# read) ships verbatim regardless of size. Measured live 2026-09-06: a
|
|
# 324k-token conversation shrank only ~8% because nearly all of it sat
|
|
# inside the protected window. This closes that gap: any tool result
|
|
# inside the protected window over this many characters still gets the
|
|
# same head/tail elision candidates get. Deliberately a much higher bar
|
|
# than max_summarize_chars -- recent results are more likely to still
|
|
# matter -- so it only catches true outliers. Set to null to disable.
|
|
protected_max_chars: 20000
|
|
# Prefix-stability probe (Wave 1 item 1.3 of plans/token-waste-waves.md).
|
|
# The provider bills the longest byte-identical PREFIX of a prompt at the
|
|
# cached rate, so rewriting an early message re-bills everything after it.
|
|
# The uniform pruning path compresses a contiguous positional region, so an
|
|
# append cannot disturb it; the relevance path below compresses a prefix of
|
|
# a relevance-ORDERED list, so a growing deficit pulls in one more candidate
|
|
# each turn at an arbitrary message POSITION. Offline that rewrites 75% of a
|
|
# payload's tokens; the live cache rate on pruned turns is 0.924, which is
|
|
# not what that should look like. This settles it on real traffic.
|
|
#
|
|
# When on, each decision row records prefix_divergence_index,
|
|
# prefix_tokens_after_divergence and prefix_prev_message_count -- three
|
|
# integers. HASHES ONLY, and not even those: the per-message digests live in
|
|
# process memory for exactly one turn and never reach the database, so
|
|
# nothing reversible is stored and a restart costs one comparison per
|
|
# session. Measured on the live median payload (98k tokens in, 74k out):
|
|
# ~1.0 ms per turn, against a request path whose floor is a provider
|
|
# round-trip of 1.4-2.0 s, plus ~10 bytes per decision row.
|
|
prefix_probe: true
|
|
relevance:
|
|
# ON by default. The embed is now budget-gated (it only fires when the
|
|
# conversation exceeds pinch.budget_tokens and pruning would actually
|
|
# happen), so the relevance path is safe to leave on. Still requires
|
|
# pinch.enabled (no effect otherwise) and still needs nomic-embed-text on
|
|
# the configured Ollama (pull: ollama pull nomic-embed-text).
|
|
enabled: true
|
|
# An EMBEDDING model, not a chat model — this must not point at
|
|
# classifier.model or verification.model. Pull one on the same Ollama:
|
|
# ollama pull nomic-embed-text
|
|
model: "nomic-embed-text"
|
|
base_url: "http://localhost:11434/v1"
|
|
timeout_seconds: 10
|
|
# Below this many trim-eligible candidates, skip the embedding call
|
|
# entirely and fall back to uniform compression — a network round trip
|
|
# to rank one candidate decides nothing.
|
|
min_candidates: 2
|
|
|
|
session_cache:
|
|
# In-memory per-session classification cache. Remembers the last
|
|
# task_category / task_tier decision for each session for a few minutes, so
|
|
# a long agent session skips the ~1-2s classifier round-trip on every turn.
|
|
# Capability flags (tools / images / json) are NEVER cached — they are read
|
|
# fresh from each request. Fallback classifications are NEVER cached. No
|
|
# persistence: a restart just reclassifies each session once.
|
|
#
|
|
# Off by default, matching every other new-and-unproven knob in this
|
|
# project: ship it, watch route_decisions.source="cached" on real traffic,
|
|
# then decide the right default.
|
|
enabled: true
|
|
staleness_seconds: 1200
|
|
|
|
circuit_breaker:
|
|
# Passive availability circuit breaker, on by default. When enabled, a model
|
|
# that returns 5xx is temporarily excluded from routing with exponential
|
|
# backoff; recovery is passive (the next real request becomes the probe once
|
|
# the cooldown passes). It has a low-risk failure mode even when wrong.
|
|
enabled: true
|
|
initial_cooldown_seconds: 30
|
|
max_cooldown_seconds: 600
|
|
backoff_multiplier: 2.0
|
|
|
|
routing:
|
|
# Access gating is prose-only in the NeuralWatt catalog ("Private preview
|
|
# (grant-gated)", "(Canary)"), so the poller parses it into access_level and
|
|
# routing excludes anything not listed here. Add 'preview'/'canary' only if
|
|
# the account actually holds the grant — otherwise dispatch earns a 403.
|
|
allowed_access_levels:
|
|
- public
|
|
|
|
# '-flex' rows are held server-side during peak until a capacity gap opens.
|
|
# That's correct for overnight/batch agent work and wrong for anything
|
|
# interactive, so a request has to opt in via latency_tolerance.
|
|
default_latency_tolerance: interactive # 'interactive' | 'batch'
|
|
|
|
# Operator's default stance on routing to '-flex' serving-class rows for
|
|
# requests that do not state one explicitly. A 4-position scale:
|
|
# no-flex never route to a flex row
|
|
# auto decide per request (the default; keeps existing behavior)
|
|
# prefer-flex flex first, standard as fallback
|
|
# force-flex flex only
|
|
# 'force-flex' is the dangerous global default: it bypasses the
|
|
# latency_tolerance: interactive hard filter for EVERY request, so even a
|
|
# request that opted into interactive routing would admit rows that are held
|
|
# server-side during peak. Use it only when you are certain the caller can
|
|
# tolerate flex latency globally.
|
|
default_flex_preference: auto
|
|
|
|
# Bare `auto` resolves to this profile. This is the profile name the router
|
|
# uses when the client does not specify one explicitly. It must name a
|
|
# built-in profile or a config-defined profile; the value is validated at
|
|
# load and will be editable from the admin Controls page in a later wave.
|
|
default_profile: "default"
|
|
|
|
# Minimum tool_use_agentic proficiency required of a model when the REQUEST
|
|
# carries tool definitions. A filter, not a weight, because it is a
|
|
# capability requirement rather than a preference.
|
|
#
|
|
# It does NOT ask whether the task is agentic — it asks whether the model
|
|
# can be trusted with tools that are on the table. The measured failure is
|
|
# the second one: deepseek-v4-flash scores 0.33 here, and the recorded case
|
|
# is a NON-agentic prompt ("it is 1:20pm, my meeting is at 3pm, how many
|
|
# minutes?") where it called two tools instead of subtracting. A model that
|
|
# over-reaches is a hazard wherever tools exist, not only where a classifier
|
|
# would say "agentic" — which it cannot do anyway: asked to identify six
|
|
# unambiguous tool-use prompts, the local models managed 2/6 and 1/6.
|
|
# Whether tools are present is stated in the request body. Read it.
|
|
#
|
|
# 0.5 sits in the empty band between the only two values the catalog
|
|
# currently holds (0.33 and 1.00), so it is not fitted to either. A model
|
|
# with NO measured tool score is unproven rather than proven bad, and is not
|
|
# dropped.
|
|
#
|
|
# CURRENTLY DISABLED (null) pending experimentation. The trade being
|
|
# measured: opencode sends `tools` on essentially every request, so with the
|
|
# filter on, deepseek-v4-flash is excluded from ordinary agent traffic and
|
|
# the ~7x cost advantage on coding routes goes unused. With it off, that
|
|
# advantage applies — and a model measured at 0.33 on tool use handles
|
|
# requests where tools are on the table.
|
|
#
|
|
# What would settle it is outcome data, not another benchmark: run with it
|
|
# off, let POST /outcome report real pass/fail, and compare
|
|
# tool_use_agentic proficiency for deepseek before and after. That is the
|
|
# one signal here that knows whether the work actually worked.
|
|
min_tool_proficiency:
|
|
tool_use_category: tool_use_agentic
|
|
|
|
# Request-side capability gates. These read the request body (image parts,
|
|
# response_format) and hard-restrict to models whose catalog row declares the
|
|
# capability. Unlike min_tool_proficiency these are NOT proficiency gates —
|
|
# a wrong guess is a guaranteed provider 400, so they gate by default and
|
|
# fail closed when the catalog flag is unknown.
|
|
require_vision: true
|
|
require_json_mode: true
|
|
|
|
# There is deliberately no flex discount knob. Whether a flex row is usable
|
|
# at all is a hard filter above (latency_tolerance), not a price adjustment
|
|
# -- being held server-side during peak is a latency property, and the
|
|
# catalog advertises flex and standard at the same token price anyway.
|
|
|
|
local_vision:
|
|
# Local Ollama vision fallback. Used ONLY when a request carries image
|
|
# parts and routing finds NO cloud model that supports_vision — the cloud
|
|
# catalog currently excludes the cost leader (deepseek) on vision, so a
|
|
# fallback is what keeps image requests working instead of 422ing.
|
|
# It speaks the OpenAI-compatible /v1 surface, so this is an Ollama endpoint
|
|
# and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host.
|
|
enabled: true
|
|
base_url: "http://localhost:11434/v1"
|
|
api_key_env:
|
|
# A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag.
|
|
# Measured live: the base tag comes up at Ollama's own default num_ctx
|
|
# (32768) and costs 9.4GB loaded — resident alongside the classifier's
|
|
# pre-fix 14GB, that left 1.9GB free on a 24GB card. num_ctx cannot be set
|
|
# per-request here: verified live that Ollama's OpenAI-compatible endpoint
|
|
# (0.22.0) silently ignores it under every field shape tried, so it has to
|
|
# be baked into the model tag itself:
|
|
# printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision
|
|
# ollama create qwen3-vl-router:4b -f Modelfile.vision
|
|
#
|
|
# Reduced 16384 -> 8192 on 2026-09-03 to buy back KV cache. Most of this
|
|
# model's footprint is KV, not weights: 3.3GB on disk but 6.9GB resident at
|
|
# 16384, and 5.6GB at 8192. That 1.3GB is what lets the dispatch model run
|
|
# at num_ctx 32768 (17GB) and still leave vision RESIDENT -- 23.2GB of 24GB
|
|
# together. Without it, loading vision EVICTS the dispatch model entirely
|
|
# and the next classification pays a ~6s cold reload on the latency floor.
|
|
#
|
|
# This is the one path that still has no measured ceiling: it gets the RAW
|
|
# message list (dispatcher.py calls it BEFORE pinch pruning), so 8192 is a
|
|
# judgement call, not a derived minimum. If image requests start failing on
|
|
# context, raise this FIRST and drop the dispatch model to num_ctx 16384 to
|
|
# pay for it. Budget guards (max_images, max_image_bytes) bound the image
|
|
# side but not the conversation text around it.
|
|
#
|
|
# Do NOT swap this for a 1B-class model to save memory. Tested 2026-09-03:
|
|
# moondream (1B, 2048 ctx) answered a real dashboard screenshot with "a
|
|
# spreadsheet ... possibly related to business decisions or financial
|
|
# analysis" -- it read no title, no column, no value, and confabulated a
|
|
# plausible description instead. qwen3-vl:4b read the page title, quoted the
|
|
# subtitle verbatim and recovered the row count. Screenshots are the actual
|
|
# workload here, and dense small text is exactly where tiny VLMs fail.
|
|
model: "qwen3-vl-router:4b"
|
|
timeout_seconds: 60
|
|
max_images: 4
|
|
max_image_bytes: 9437184
|
|
|
|
local_dispatch_models:
|
|
# Local Ollama models that can serve dispatch (not just classification/vision).
|
|
# Each entry is a model that the router can route traffic to — the same path
|
|
# that picks cloud models from the NeuralWatt catalog, but calling a local
|
|
# OpenAI-compatible endpoint instead.
|
|
#
|
|
# These models MUST ALSO appear under proficiency.categories in eligible_categories
|
|
# and MUST have a tier pin under tiering.model_tiers, because src/tier.py
|
|
# OVERWRITES models.tier on every run (the config override is the only pin).
|
|
#
|
|
# To add a model: copy this block and change the values. Keep the comment
|
|
# describing how to build the model tag (Modelfile recipe), because the num_ctx
|
|
# baked into the tag is what Ollama's /v1 endpoint actually uses — setting it
|
|
# per-request through the OpenAI API is not supported by Ollama.
|
|
- # DEFAULT. Measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5), and a
|
|
# sound starting point for any 24-32GB card -- the numbers below are real,
|
|
# not aspirational. It is a default rather than a recommendation only in
|
|
# the sense that YOUR hardware and YOUR taste in models should win:
|
|
# docs/local-models.md documents the measurement method so you can swap in
|
|
# whatever you actually want to run and defend the choice with your own
|
|
# numbers instead of inheriting these.
|
|
#
|
|
# Dormancy under the DEFAULT profile (measured 2026-09-02/03): quality-gated
|
|
# ranking — the local row must score within objective.quality_tolerance
|
|
# (0.10) of the cloud leader before its price advantage is even consulted.
|
|
# Today it does not: 0.767 vs 0.95 on file_summarization (gap 0.183), so
|
|
# this row is dormant under the default profile. That is by design, not a
|
|
# bug. Do NOT widen quality_tolerance. It fires as a fallback when the cloud
|
|
# is unavailable; see docs/routing.md § Local dispatch branch.
|
|
#
|
|
# qwen2.5-coder-router:14b serves BOTH local dispatch and classification,
|
|
# so only one model stays resident. Selected 2026-09-02/03 by measurement:
|
|
#
|
|
# file_summarization (n=6, judge-scored -- treat +-0.15 as a tie):
|
|
# deepseek-v4-flash (cloud) 0.95 <- local costs ~0.18 of quality
|
|
# qwen2.5-coder:32b 0.833 (22GB: evicts vision, 3.4s classify)
|
|
# qwen2.5-coder:14b Q8_0 0.817 (19GB, 1.5x slower, gain within noise)
|
|
# qwen2.5-coder:14b Q4_K_M 0.767 <- chosen
|
|
# nemotron-mini:4b 0.25 (the plan's original pick; FABRICATED
|
|
# on 3 of 6, and detected 0/4 bugs)
|
|
#
|
|
# classification was 11/14 with 0 hard failures for EVERY variant above --
|
|
# quantization and context size changed only speed, never accuracy.
|
|
# Q4_K_M @ 32k: 1.14s mean. Q8_0 @ 32k: 3.79s, because it needs 24GB and
|
|
# Ollama spills 17% to CPU (watch the "17%/83%" column in `ollama ps`).
|
|
#
|
|
# num_ctx 16384 measured (see plans/local-decision-classifier-results.md
|
|
# sections 7 and 9): the 14b at 32k is 16.5-17.8 GB resident, at 16k it is
|
|
# 12.26 GB. That headroom is what lets the local_decision classifier
|
|
# qwen3.5:4b (num_ctx 8192, 5.9 GB) coexist: 12.3 + 5.9 = ~19.9 of 24 GB
|
|
# (measured 2026-09-28). Local vision is off because its model no longer
|
|
# fits.
|
|
#
|
|
# Modelfile recipe (run on an Ollama host):
|
|
# ollama pull qwen2.5-coder:14b
|
|
# printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch
|
|
# ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch
|
|
model_id: "qwen2.5-coder-router:14b"
|
|
base_url: "http://localhost:11434/v1"
|
|
api_key_env:
|
|
timeout_seconds: 180
|
|
context_window: 16384
|
|
max_output_tokens: 2048
|
|
tier: 1
|
|
# Restricted to file_summarization and diff_checking. The router's hard
|
|
# filter (routing.py) will exclude this model for any other category. Do
|
|
# not widen this without measuring the new category first -- this model
|
|
# scored 0/4 at detecting bugs before the task set was fixed, and a
|
|
# confidently wrong answer is worse than an honest error.
|
|
eligible_categories:
|
|
- file_summarization
|
|
- diff_checking
|
|
|
|
verification:
|
|
# Structural checks (parse the code, never run it) are free, pure Python and
|
|
# always on — they need no model and run anywhere, down to an RPi.
|
|
# This section governs the LOCAL LLM check, which is not free.
|
|
#
|
|
# Set local_llm_enabled: false on a host with no usable local inference.
|
|
# Structural checking and POST /outcome both keep working; only the
|
|
# refusal/incoherence class of failure stops being caught.
|
|
local_llm_enabled: true
|
|
|
|
# This check speaks Ollama's NATIVE API (/api/chat with think=False), which
|
|
# no cloud provider offers, so it is configured SEPARATELY from the
|
|
# classifier rather than derived from it. Deriving it meant that pointing
|
|
# classification anywhere else sent these requests to <that host>/api/chat.
|
|
#
|
|
# Point it at the SAME Ollama as the classifier when that is across a VPN;
|
|
# leave the model null in that case, since the hosts then match.
|
|
base_url: "http://localhost:11434"
|
|
# null means "whatever classifier.model is", correct only while both run on
|
|
# the same local Ollama — config load REFUSES the null once the hosts differ,
|
|
# because the fallback would name a model this Ollama has never heard of and
|
|
# the verifier would fail silently. Stated explicitly here anyway, matching
|
|
# classifier.model exactly (see that field's comment) so both calls hit the
|
|
# SAME resident Ollama instance, loaded once at its Modelfile-tagged
|
|
# context size rather than two separately-sized copies.
|
|
model: "qwen2.5-coder-router:14b"
|
|
|
|
# Only check answers this large. Measured on real traffic: a local check
|
|
# costs ~15% of a median 193-token answer, so it would only pay if such
|
|
# answers failed more than ~15% of the time. At 1,500 completion tokens the
|
|
# break-even failure rate drops to ~1.9%, which is plausible. Below the
|
|
# threshold a check costs more than the risk it removes.
|
|
min_completion_tokens: 600
|
|
|
|
# The check runs AFTER the response has gone back to the client, so it never
|
|
# adds its ~6s to anyone's latency. It exists to learn which models fail on
|
|
# real work, not to gate answers.
|
|
timeout_seconds: 60
|
|
max_output_tokens: 1024
|
|
|
|
# How far back /outcome looks when a report carries no request_id and no
|
|
# matching directory. A test run follows the completion that caused it within
|
|
# seconds, so this is deliberately short: a wide window sweeps in sessions
|
|
# that finished long ago and makes every report look ambiguous. If more than
|
|
# one conversation was active inside it, the report is refused rather than
|
|
# guessed at.
|
|
outcome_attribution_window_seconds: 120
|
|
|
|
freshness:
|
|
stale_after_days: 3
|
|
# Router refuses to route to a model whose row is stale/deprecated,
|
|
# regardless of how good its score would otherwise be.
|
|
exclude_stale: true
|
|
exclude_deprecated: true
|
|
# Allowlisting a model does nothing until a poll ingests it, and the timer
|
|
# runs every 2 hours -- so a model added through /admin sat invisible until
|
|
# the next tick, looking like the allowlist had not worked.
|
|
#
|
|
# Debounce, not delay: each further edit restarts the clock, so adding five
|
|
# models one at a time costs one poll, not five. 0 disables the trigger and
|
|
# leaves the timer as the only path.
|
|
repoll_after_allowlist_change_seconds: 5.0
|
|
|
|
database:
|
|
path: "router.db"
|
|
|
|
classifier:
|
|
# The local LLM that classifies each task before a cloud model answers it.
|
|
# "Local" means your hardware, not necessarily this machine — the box with
|
|
# the GPU is usually not the laptop you are typing on. It speaks plain
|
|
# OpenAI-compatible chat completions, so point it wherever Ollama lives:
|
|
#
|
|
# same machine base_url: http://localhost:11434/v1
|
|
# over a VPN base_url: http://<vpn-ip>:11434/v1
|
|
# (serving host needs deploy/ollama-over-vpn.conf; Ollama
|
|
# binds loopback-only by default and will refuse)
|
|
#
|
|
# Any OpenAI-compatible endpoint works, so a cloud model can classify too —
|
|
# set api_key_env for one that checks a key. Worth knowing before assuming
|
|
# local is the cheap option: on five prompts NeuralWatt's deepseek-v4-flash
|
|
# classified in 1.02s mean against qwen3.5's 11.58s on an RTX 6000, agreed
|
|
# with the label 5/5 against 2/4, and cost $0.093 per thousand calls. Local
|
|
# inference is not free, it is unbilled.
|
|
provider: "ollama"
|
|
base_url: "http://localhost:11434/v1"
|
|
# Unset means unauthenticated, which is the Ollama case. Name the env var
|
|
# holding the key when the endpoint actually checks one.
|
|
api_key_env:
|
|
# A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag
|
|
# — must match a model `ollama list` reports. Ollama loads a model at its
|
|
# library Modelfile's default context unless told otherwise, and the base
|
|
# tag never was: measured live, mistral-nemo:12b came up at num_ctx=32768
|
|
# (14GB of a 24GB card) though max_input_chars + the system prompt +
|
|
# max_output_tokens need well under a quarter of that. num_ctx cannot be
|
|
# set per-request here: verified live that Ollama's OpenAI-compatible
|
|
# endpoint (0.22.0) silently ignores it under every field shape tried (a
|
|
# 200 comes back, the loaded context never changes), so it has to be baked
|
|
# into the tag itself:
|
|
# printf 'FROM mistral-nemo:12b\nPARAMETER num_ctx 8192\n' > Modelfile.router
|
|
# ollama create mistral-nemo-router:12b -f Modelfile.router
|
|
# 8192 leaves 2x headroom over the real requirement. verification.model
|
|
# points at this same tag (see its own comment) so classify and verify
|
|
# share one resident instance at one context size, rather than risking a
|
|
# reload thrash from two differently-sized copies of the same base model.
|
|
model: "qwen2.5-coder-router:14b"
|
|
# A cold Ollama took 43s to answer the first classification, which blew the
|
|
# old 30s ceiling and returned 503 to the client. Warm it is ~5s. The
|
|
# ceiling is for a cold model load, not the steady state.
|
|
timeout_seconds: 120
|
|
# Classification must be reproducible: at the default temperature the same
|
|
# task was classified tier 2 then tier 1 on consecutive calls, which routed
|
|
# it to two different models. Routing that changes under an identical
|
|
# prompt is untraceable.
|
|
temperature: 0
|
|
# Hard cap on the classifier's generation, to bound a failure mode that
|
|
# cascades on REASONING models: they emit a long chain of thought, blow past
|
|
# timeout_seconds, and Ollama keeps generating after the client gives up AND
|
|
# serializes per model, so one runaway request queues every later one behind
|
|
# it. Measured on qwen3.5, which failed this way on 4 of 19 calls — each one
|
|
# a 15s wait ending in a silent fallback.
|
|
#
|
|
# The current default (mistral-nemo) does not reason, so this does not bind
|
|
# for it. Kept anyway: it costs nothing when unused and is the only thing
|
|
# standing between a swapped-in reasoning model and that cascade. If you do
|
|
# swap one in, note 256 was too tight — the trace consumed the whole budget
|
|
# and the model was truncated before emitting any JSON.
|
|
max_output_tokens: 1024
|
|
|
|
# Where routing lands when the classifier times out, errors, or returns
|
|
# something unparseable. A local model being slow should degrade routing,
|
|
# not refuse the request — the caller is a coding agent that would rather
|
|
# have a mid-tier answer than a 502.
|
|
# Ceiling on the text handed to the classifier (head + tail, middle elided).
|
|
# It decides a category and a tier; it does not need the document, and
|
|
# feeding it one is harmful rather than merely wasteful. Measured on a ~20k
|
|
# token prompt: qwen3.5 spent 28.7s and returned empty (budget consumed by
|
|
# its reasoning trace), mistral-nemo spent 41.8s echoing the input back
|
|
# inside its JSON. Both land on source: "fallback" — the same answer an
|
|
# instant failure gives, after 30-40s of local inference.
|
|
#
|
|
# Nothing is lost: chat_completions measures the real conversation with
|
|
# estimate_prompt_tokens and takes the larger value, so required_context
|
|
# never depends on what the classifier saw. 0 disables clamping.
|
|
max_input_chars: 8000
|
|
|
|
# When the chat path supplies the previous turn as context (see the pinch /
|
|
# context notes), frame the classifier input as "Context: <prev> / Message:
|
|
# <current>" so a short follow-up inherits the prior turn's complexity
|
|
# instead of being classified in isolation as trivial.
|
|
context_framing: true
|
|
|
|
fallback_tier: 2
|
|
fallback_category: general_chat
|
|
|
|
# --- which implementation is PRIMARY -----------------------------------
|
|
# This is a peer concept to the fallback cascade below, not a replacement
|
|
# for it: whichever mode is primary, a failure still walks the SAME
|
|
# cascade (stale session -> session history -> cloud_fallback -> the
|
|
# fallback_tier/fallback_category above).
|
|
#
|
|
# local_llm (default) the model/base_url above, unchanged.
|
|
# cloud_llm a cloud model answers PRIMARY, not just as a fallback
|
|
# after local fails. Requires exactly one of
|
|
# cloud_primary or cloud_primary_auto below.
|
|
# local_encoder a small, non-generative classifier (see the encoder:
|
|
# block below) that cannot exhibit the runaway-reasoning
|
|
# failure mode documented above, at the cost of only
|
|
# producing task_category -- task_tier falls back to
|
|
# fallback_tier for this mode. Zero-shot, not trained on
|
|
# your traffic: this router never stores raw task text.
|
|
mode: local_llm
|
|
|
|
# Read only when mode: cloud_llm. Exactly one of these two:
|
|
# cloud_primary:
|
|
# base_url: https://api.neuralwatt.com/v1
|
|
# model: deepseek-v4-flash
|
|
# api_key_env: NEURALWATT_API_KEY
|
|
# timeout_seconds: 2
|
|
# max_output_tokens: 1024
|
|
# cloud_primary_auto: true # resolve the cheapest routable model, live
|
|
|
|
# Read only when mode: local_encoder. Every field has a default, so
|
|
# `encoder: {}` is enough to opt in.
|
|
# encoder:
|
|
# model: facebook/bart-large-mnli
|
|
# device: cpu
|
|
# confidence_min: 0.5
|
|
# # Design sketch only (see _TierFeatureClassifier in local_encoder.py):
|
|
# # predict task_tier from non-textual request features. Not wired.
|
|
# tier_from_features: false
|
|
# tier_feature_fields: []
|
|
|
|
# --- fallback cascade ---------------------------------------------------
|
|
# When the local classifier fails, the router walks: stale session cache ->
|
|
# this session's history in route_decisions -> the optional cloud
|
|
# classifier below -> fallback_tier/fallback_category above. The first two
|
|
# steps are free and local.
|
|
#
|
|
# Global backoff after a classifier failure. This is what bounds cloud
|
|
# spend during a sustained outage: at most one cloud attempt per window
|
|
# across ALL requests, not one per request. It is also reused as the window
|
|
# in which an account-level provider refusal suppresses the cloud step,
|
|
# since an out-of-credit account makes that call a guaranteed waste.
|
|
cooldown_seconds: 30
|
|
# Degradation warning on /metrics. Silent below degraded_warn_min
|
|
# decisions in 24h, because a share computed over a handful of requests is
|
|
# noise and a flapping warning is one nobody reads.
|
|
degraded_warn_min: 20
|
|
degraded_warn_threshold: 0.5
|
|
# Optional. ABSENT BY DEFAULT, which is what makes the cascade cost
|
|
# nothing: with no cloud_fallback block the router degrades straight to the
|
|
# static guess. Uncomment and point it at any OpenAI-compatible endpoint to
|
|
# trade a little money for a real classification during a local outage.
|
|
# A configured api_key_env whose variable is missing degrades to the static
|
|
# guess rather than failing the request -- unlike the primary classifier,
|
|
# which raises, because this one is only reached when things already broke.
|
|
# cloud_fallback:
|
|
# base_url: https://api.neuralwatt.com/v1
|
|
# model: deepseek-v4-flash
|
|
# api_key_env: NEURALWATT_API_KEY
|
|
# timeout_seconds: 2
|
|
# max_output_tokens: 1024
|
|
response_format: "json" # ask Ollama to constrain output to valid JSON
|
|
# The labels a classifier may return. A SUBSET of proficiency.categories,
|
|
# validated as one at config load, and deliberately not the same list:
|
|
# proficiency.categories is the scoring axis (what a model is good at),
|
|
# this is the set a classifier is asked to choose between.
|
|
#
|
|
# tool_use_agentic is excluded. It describes what a turn mechanically DOES
|
|
# rather than what it is for, and an agent turn is always both -- measured
|
|
# 2026-09-14 on live traffic with the session cache off, 31 of 31
|
|
# consecutive turns classified tool_use_agentic and every one routed to the
|
|
# slowest model in the catalog (14.8s TTFT, p95 40s). Per-turn
|
|
# classification became accurate and routing got worse. The signal it was
|
|
# standing in for is already read exactly and for free from the request's
|
|
# own `tools` array by routing.min_tool_proficiency.
|
|
#
|
|
# WHAT THIS COSTS, so the next person does not discover it: with the
|
|
# classifier unable to emit tool_use_agentic, no new POST /outcome report
|
|
# attributes to that category, so its proficiency scores FREEZE at their
|
|
# current values (e.g. qwen3.6-35b at 0.902 over 154 samples). The tool
|
|
# filter keeps reading them; they simply stop accumulating. Accepted for
|
|
# now.
|
|
#
|
|
# Considered and REJECTED as the fix: attributing outcomes to
|
|
# tool_use_agentic whenever the request carried a `tools` array. opencode
|
|
# sends `tools` on essentially every request, so the score would converge
|
|
# on each model's overall pass rate and stop discriminating the exact thing
|
|
# the category exists to detect -- a model that OVER-reaches for tools on
|
|
# work that did not need them. A frozen honest score beats a live
|
|
# meaningless one.
|
|
#
|
|
# Omit this key entirely to mean "every proficiency category".
|
|
candidate_categories:
|
|
- coding_general
|
|
- coding_refactor
|
|
- debugging
|
|
- docs_writing
|
|
- summarization
|
|
- file_summarization
|
|
- diff_checking
|
|
- translation
|
|
- reasoning_math
|
|
- general_chat
|
|
# The dispatcher appends the authoritative category list from
|
|
# classifier.candidate_categories (above) to this prompt at call time. Do
|
|
# not enumerate the categories here as well — a hand-copied list drifts, and
|
|
# a category the model invents joins against nothing in the proficiency
|
|
# table.
|
|
system_prompt: |
|
|
You are a task router. Given a task description and any attached context,
|
|
respond with ONLY a JSON object with these fields:
|
|
{
|
|
"task_category": one of the allowed categories listed below,
|
|
"task_tier": integer 1-3, where 1 is cheap/simple, 2 is mid/general,
|
|
and 3 is frontier/high-stakes,
|
|
"required_context_tokens": integer estimate of prompt+context token count,
|
|
"confidence": float 0-1
|
|
}
|
|
|
|
dispatch_providers:
|
|
neuralwatt:
|
|
base_url: "https://api.neuralwatt.com/v1"
|
|
api_key_env: "NEURALWATT_API_KEY"
|
|
has_energy_telemetry: true
|
|
# false: NeuralWatt reports its bill in a top-level `cost` block (and as a
|
|
# `: cost {...}` SSE comment when streaming), so no request-side opt-in is
|
|
# needed and none is sent. See the openrouter entry below for what the
|
|
# flag does when it is on.
|
|
reports_cost_in_usage: false
|
|
enabled: true
|
|
|
|
openrouter:
|
|
base_url: "https://openrouter.ai/api/v1"
|
|
api_key_env: "OPENROUTER_API_KEY"
|
|
# Account-level balance poll; OpenRouter reports per-completion energy via
|
|
# a separate allowance_remaining field, so this URL is for the prepaid pool.
|
|
balance_url: "https://openrouter.ai/api/v1/credits"
|
|
has_energy_telemetry: false
|
|
# true: send OpenRouter's own `usage: {"include": true}` opt-in on every
|
|
# upstream request. `stream_options.include_usage` is the OpenAI spelling
|
|
# and gets token counts; this is the separate OpenRouter one, and without
|
|
# it the accounting block is simply absent -- no billed `usage.cost` and
|
|
# no `usage.prompt_tokens_details.cached_tokens`. Both ride on this one
|
|
# opt-in, so the two coverage figures move together.
|
|
#
|
|
# It does NOT bring generation timing: OpenRouter serves that only from
|
|
# its separate /api/v1/generation endpoint, so `duration_seconds` stays
|
|
# NULL on OpenRouter rows. Turn this off only for a provider that returns
|
|
# accounting unasked, or one that rejects the key.
|
|
reports_cost_in_usage: true
|
|
enabled: true
|
|
require_allowlist: true
|
|
|
|
local_compute:
|
|
# "Gaming mode", inverted: true means the router may use local hardware.
|
|
# Set it to false when you stop Ollama to give the GPU back to something
|
|
# else, and the router SKIPS every local call instead of discovering the
|
|
# outage one 120s timeout at a time: classifier, /health probe, local
|
|
# verification, the local-vision fallback, and local dispatch rows (which
|
|
# drop out of the candidate set rather than being picked and then failing).
|
|
#
|
|
# ONE flag that the call sites read, not a macro that rewrites
|
|
# verification.local_llm_enabled / local_vision.enabled / local_energy.enabled.
|
|
# Those keep their own meanings; this is an outer AND over all of them, so
|
|
# turning it back on restores exactly the state you left.
|
|
#
|
|
# Config load REFUSES false unless classifier.cloud_fallback is configured.
|
|
# Skipping the local classifier does not make classification remote, it makes
|
|
# it a static guess -- and that guess is recorded as general_chat, a fully
|
|
# scored category, so it looks like real classification afterwards.
|
|
# Uncommenting the cloud_fallback example above is the intended setup path.
|
|
enabled: true
|
|
|
|
local_energy:
|
|
# Local energy metering for the router's own Ollama calls (classifier,
|
|
# verifier, local-vision fallback). Cloud providers expose per-request energy,
|
|
# but these calls run on your own hardware and their electricity is real even
|
|
# though it never shows up on the NeuralWatt bill.
|
|
#
|
|
# OFF by default. Enable only after setting a real per-kWh rate below.
|
|
# Config load REFUSES enabled: true with a null tariff because a null rate
|
|
# would record cost_usd = 0 instead of "unknown".
|
|
enabled: false
|
|
|
|
# Power sampler. Only "nvidia_smi" is supported today. It shells out rather
|
|
# than importing pynvml so the router has no new dependency and degrades
|
|
# cleanly on non-NVIDIA hosts.
|
|
meter: "nvidia_smi"
|
|
|
|
# How often to sample GPU power while a local call is running. 0.25s was the
|
|
# original llmrouter pinch sample rate and is fine as a starting point.
|
|
sample_interval_seconds: 0.25
|
|
|
|
# YOUR real per-kWh electricity rate. Set this before enabling metering.
|
|
#
|
|
# Example (do not use this number unless it matches your bill):
|
|
# tariff_usd_per_kwh: 0.12
|
|
#
|
|
# Set 2026-09-02 from the user's Great Lakes Energy bill. This is the
|
|
# MARGINAL rate -- what one more kWh actually costs -- not the all-in
|
|
# effective rate of 0.183 ($410.15 / 2,244 kWh). The difference is the
|
|
# $52.30/month of fixed charges, which are paid whether or not the GPU
|
|
# runs, so loading them onto compute would overstate local cost by ~15%
|
|
# and bias routing toward the cloud for a reason that is not real.
|
|
# 14.742c energy + 0.188c PSCR + 0.316c EO = 15.246c, x 1.04 MI sales
|
|
# tax = 15.856c/kWh.
|
|
tariff_usd_per_kwh:
|
|
|
|
# YOUR grid carbon intensity in grams of CO2 equivalent per kWh. Optional:
|
|
# null disables the carbon estimate but still records energy and cost.
|
|
#
|
|
# Example (do not use this number unless it matches your grid):
|
|
# grid_intensity_g_per_kwh: 475
|
|
grid_intensity_g_per_kwh:
|
|
|
|
logging:
|
|
# Whether to write a row to energy_observations for every completion. Off
|
|
# means no cost accounting, no reference sweep and no /outcome attribution,
|
|
# so leave it on unless you are debugging.
|
|
#
|
|
# There is no log_path. Nothing ever wrote a file -- the dispatcher prints to
|
|
# stderr and the systemd unit hands that to the journal (`journalctl --user
|
|
# -u llm-router -f`), so the setting named a destination that did not exist.
|
|
log_energy_observations: true
|
|
|
|
# Whether to write a row to route_decisions for every routing decision
|
|
# (route | dispatch | chat | passthrough | local-vision). Off leaves the
|
|
# monitoring TUI's decision history empty; it does not change what routes.
|
|
log_route_decisions: true
|
|
|
|
# debug | info | warning | error.
|
|
#
|
|
# info gives one line per request: what it was classified as, which model won,
|
|
# what it cost, how long each stage took. debug adds why -- every candidate
|
|
# that was dropped and by which filter, the ranking with scores, the
|
|
# classifier's raw reply. It logs no conversation text at any level; prompts
|
|
# here run 60k-150k tokens and the journal is on disk.
|
|
#
|
|
# LLM_ROUTER_LOG_LEVEL overrides this, so a running service can be turned up
|
|
# without editing a tracked file:
|
|
#
|
|
# systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
|
|
# systemctl --user restart llm-router
|
|
level: info
|
|
|
|
watchdog:
|
|
enabled: true
|
|
local_llm_enabled: true
|
|
model: null # defaults to verification.model
|
|
read_only_agents:
|
|
- explore
|
|
- librarian
|
|
- oracle
|
|
detector:
|
|
window: 60
|
|
dup_min: 0.25
|
|
top_min: 12
|
|
top_min_ro: 8
|
|
cum_min: 15
|
|
cover_min: 4.0
|
|
min_calls: 40
|
|
dashboard_base_url: "http://127.0.0.1:8080/admin" # base URL for alert links
|
|
# tick interval is 5 min (set in deploy/llm-router-watchdog.timer)
|
|
|
|
notifications:
|
|
channels:
|
|
- name: default
|
|
type: desktop
|
|
enabled: true
|
|
min_severity: warning
|
|
|