Files
6krrt/config/config.yaml
adlee-was-taken 1a6354deea refactor(session-cache): rename staleness_minutes to staleness_seconds
Change the unit of session_cache.staleness from minutes to seconds so it
can express finer-grained (sub-minute) staleness windows. This is a
straight rename, not an additive/compat knob — no deprecated alias, per
the project's convention of updating every consumer in the same change.

New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was
120 minutes). Default: 1200 seconds (was 20 minutes). The validator's
reasoning is unit-independent and carries over: the floor is deliberately
> 0 because 0 would make session_cache.get() miss every turn while
put() still writes and the classifier-failure cascade's stale_read
ignores staleness; the ceiling reasoning (unbounded window = never-expiring
cache, 7200s still >> 840s real max run) also carries over in seconds.

Every consumer updated in the same commit:
- src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX
= 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated
- src/dispatcher.py: drop the * 60 conversion (field is native seconds)
- src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST,
  _CONFIG_GET_ORDER, _runtime_state, error message template
- admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS
- config/config.yaml: staleness_seconds: 1200
- tests: test_admin_runtime/config/frontend/knob_coverage, plus stale
  comment in test_chat_completions
- docs: admin-portal.md, evaluation.md, README.md

config.local.yaml is gitignored and will be migrated separately.
2026-09-18 00:45:08 -04:00

1069 lines
55 KiB
YAML

# Local LLM router config
# All weights, thresholds, and provider settings live here so they can be
# tuned without touching code. Loaded/validated by config.py.
objective:
# Quality is the objective. Cost is a constraint and a tiebreak. Eco is
# logged per request but is NOT optimized here — that judgement is made
# outside this router.
#
# This replaced a weighted blend (cost 0.4 / eco 0.2 / proficiency 0.4).
# Measurement killed it: turning the cost weight from 0.4 to ZERO changed
# the winner in only 2 of 6 categories, so the blend was never steering on
# quality — while 60% of every decision adjudicated fractions of a cent
# (all real traffic to date totals $0.07).
# Proficiency differences smaller than this are treated as equal and the
# cheaper model wins. This is measurement noise, not preference: scores
# currently rest on 2-3 samples per category, so a 0.05 gap is
# indistinguishable from sampling variation and paying for it buys noise.
# Narrow it as samples accumulate.
# Above this pace ratio, the dashboard raises an alarm.
# 1.25 = burning 25% faster than the plan allows.
plan_pace_warn_ratio: 1.25
quality_tolerance: 0.1
# Cost is priced per-request from catalog prices, NOT from a benchmark
# sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper
# than deepseek-v4-flash; on a realistic 70k-token prompt deepseek is 5.0x
# cheaper. Attribution inverts with prompt size, so a fixed-shape benchmark
# cannot rank models for a workload of another shape. Catalog prices scaled
# to the actual request agree with the live measurement, cost nothing, and
# need no sweep.
# Share of prompt tokens served from the provider's prefix cache. Agent
# clients resend the whole conversation each turn, so most of it is a hit.
#
# Measured token-weighted across 50 sessions and 40.7M tokens (2026-08-23):
# 91.7% overall, 92.6% on the sessions above 400k tokens, which are the ones
# that carry the cost. The previous 0.84 came from 2.2M tokens of much
# earlier traffic.
#
# Raising it changed NO winner in any of the nine categories at 60k context
# — checked before editing. It is here because the number should be true,
# not because the routing needed it.
assumed_cache_rate: 0.917
# Completion length assumed when pricing a request. Real sessions here median
# around 200-400 completion tokens against enormous prompts.
assumed_completion_tokens: 500
# Per-request ceiling on measured ENERGY, in kWh. null disables it.
#
# Denominated in kWh rather than dollars. For scale: the reference task runs
# ~5e-06 kWh on the cheapest model and ~2.2e-04 on the most expensive.
# Overage is billed against the account's credit balance;
# plan_kwh_per_period gates nothing.
max_energy_per_request:
# The subscription's kWh allowance per billing period, for reporting burn in
# /health. This is a planning figure only: per-request traffic is never
# refused for exceeding plan_kwh_per_period — it gates nothing. Set to match
# your plan; null disables the report. NeuralWatt also returns
# allowance_remaining_usd per request, which is logged for /metrics, but that
# is a dollar figure while the plan is denominated in energy.
plan_kwh_per_period: 6.25
# Hours of recent balance history used to estimate the burn rate. The most
# recent monotonically-decreasing segment of allowance_remaining_usd values
# (segments split at each balance increase) is examined over this window.
quota_burn_window_hours: 24
# Projected hours of runway below which /metrics and dashboards emit a
# low-warning boolean; only fires when a burn rate estimate exists and the
# projected remainder is positive but short.
quota_runway_warning_hours: 6
# A burn estimate needs at least this many balance samples in the most recent
# monotonically-decreasing segment (post-top-up resets the segment); below
# which the estimate is None with an explanatory runway note.
quota_burn_min_segment_samples: 3
# ...and the segment must span at least this many hours, otherwise the
# estimate is None with an explanatory note (never a wild extrapolation).
quota_burn_min_segment_hours: 0.5
# The day-of-month your NeuralWatt subscription billing cycle resets. Set
# this to YOUR real billing day so the admin quota modal shows a genuine
# next-reset date instead of a misleading rolling-window start. null disables
# the feature and the modal shows "not configured". Valid range: 1-28 (skip
# 29-31 to avoid shorter-month edge cases).
billing_reset_day: 6
# Rejection-rate detection over route_decisions rows that selected no model
# (422 "no model satisfies the hard filters"). The signal is novelty or rate —
# NEVER mere presence: this deployment routinely has ~3 rejections/hour of
# ordinary over-large tier-3 requests that are behaving as designed (6 in the
# last 24h, 17 all-time at 2026-09-05), so a count-only tripwire would be
# permanently on.
#
# alert window (hours): rejections this recent are counted per
# (task_tier, digit-normalized reason) group.
rejection_warning_window_hours: 1
# baseline window (hours) BEFORE the alert window: groups absent here are
# "new". A new group warns from 2 occurrences — the 2026-09-04 vision
# incident produced exactly 2 and no rate threshold can sit below routine
# noise yet above that.
rejection_warning_baseline_hours: 24
# a group already present in the baseline warns only at this many
# occurrences inside the alert window: 6 = 2x the observed routine hourly
# peak, far below the dozens/hour a deprecation flare produces.
rejection_warning_min_count: 6
# How far back /metrics looks to decide which models the router has actually
# picked. Long on purpose: this answers "has this model EVER been chosen",
# and a model that only wins one category can go days between selections.
# Feeds two very different reports -- see metrics.selection_coverage.
selection_coverage_window_hours: 168
# Prefix-cache rate measured from energy_observations, and the expiry check
# on assumed_cache_rate above. That constant was measured ONCE, on
# 2026-08-23, and it is the highest-leverage term in routing.estimated_cost
# on a 100k-token prompt — a premise that large should not go unchecked just
# because it was true when it was written.
#
# The comparison is a DIVERGENCE from assumed_cache_rate, not an absolute
# floor. "Is the cache working" is the wrong question; a deployment whose
# real rate is 0.60 is not broken, it is mispriced, and a floor would say
# nothing about the number the cost model is actually built on.
#
# Trailing window, in hours. Long for the same reason
# selection_coverage_window_hours is: only a fraction of rows carry a
# reported cached count, so a 1h window is usually empty.
cache_rate_window_hours: 168
# Absolute divergence from assumed_cache_rate that warns. 0.10 sits well
# outside ordinary session-mix drift (measured 2026-09-13: 0.879 neuralwatt
# / 0.901 openrouter against the assumed 0.917) and well inside the 0.952 vs
# 0.478 same-model/switched gap the waves plan is chasing.
cache_rate_warn_margin: 0.10
# Minimum reported-cache observations before either warning fires, per
# aggregate and per (provider, model) group. Observations, not tokens: one
# 200k-token prompt outweighs a hundred ordinary turns, so a token floor
# would let a single request's luck read as a measurement.
cache_rate_warn_min_observations: 25
# --- Report-only measurement series. Neither is read by routing. ---------
#
# Cost-estimator calibration: routing.estimated_cost's prediction against
# the provider's own billed figure, per (provider, model), joined on
# request_id. Measured 2026-09-12 at ~100% telemetry coverage, the estimator
# is high by 1.6x-13x depending on the model.
#
# The scale error is NOT the finding — a uniform overestimate reorders
# nothing, because the ranking is a comparison and every candidate moves
# together. The SPREAD in that error is the finding: it priced
# deepseek/deepseek-v4-flash below qwen3.6-35b while the bill said the
# reverse. /metrics reports `spread` for exactly that reason.
#
# Nothing applies these factors. They are here so their stability can be
# judged first — this project has already mistaken one moment of a moving
# per-model quantity for a constant (see CLAUDE.md, "Attribution drifts
# across hours, so sampling must too").
#
# Trailing window, in hours. Long for the same reason
# cache_rate_window_hours is: only rows with both an estimate and a billed
# figure qualify, and a short window is usually too thin to read.
cost_calibration_window_hours: 168
# Joined observations before a (provider, model) factor is marked
# `sufficient` and allowed to set the reported spread. Groups below it are
# still listed with their counts — the count is itself information — but a
# ratio over three requests must not become the headline.
cost_calibration_min_observations: 25
# Router-observed latency: router_wall_seconds and router_ttft_seconds from
# energy_observations, p50 and p95 per (provider, model). These are the
# ROUTER's clock, not the provider's duration_seconds — which is why this
# can see OpenRouter at all, since OpenRouter reports no duration.
#
# It found z-ai/glm-5.3-flash, a model with "flash" in its name, at a p50 of
# 9.6s to first token against 1.6s for deepseek-v4-flash on NeuralWatt.
# Reported, not scored: latency is not an objective here, and one reading of
# a quantity that tracks pool load is not grounds to make it one.
latency_window_hours: 168
# Observations before a group's percentiles are marked `sufficient`. Applied
# SEPARATELY to the wall and TTFT counts, because TTFT is streaming-only by
# nature and a buffered deployment legitimately has fewer of them.
latency_min_observations: 25
# Gate: with this off, routing is byte-identical to today.
# Flip on only after the Wave 1 post-restart baseline day;
# see plans/token-waste-waves.md Wave 2 gate.
incumbent_cache_pricing: false
# Challenger cache-rate dial.
# - null (default): neutral, follows assumed_cache_rate → today's ranking exactly
# including the incumbent; this is the off position for tuning.
# - 0.0: challengers priced as fully cold prompts (maximum incumbent advantage).
# - Any value between is a partial cache-penalty — the whole wave dials
# from off to full here, no revert needed.
incumbent_challenger_cache_rate:
# Seconds between refreshes of the measured per-(provider, model) cache
# rates. The TTL is measured from the first call after the process started
# (time.monotonic), so a brief post-restart cold period is expected.
incumbent_rate_refresh_seconds: 300
# Minimum reported-cache observations before a per-(provider, model)
# cache rate is trusted for pricing decisions — the independent pricing
# floor. Tuning this does NOT move the cache-rate warning floor
# (cache_rate_warn_min_observations): the two knobs answer different
# questions with different failure costs.
#
# Shipped with a documented default (25) so the code always reads a
# concrete float when the penalty is engaged.
incumbent_rate_min_observations: 25
# Credit-aware routing attenuation (OFF BY DEFAULT).
# When enabled, a provider configured with a balance_url (e.g. OpenRouter's
# prepaid account balance) gets its *comparison cost* inflated inside the
# quality-first ranker as its account balance nears zero. Quality bands still
# win; this only shifts ties. Providers whose balance comes from per-request
# energy telemetry allowance_remaining_usd (e.g. NeuralWatt's overage-billed
# subscription) are ALWAYS multiplier 1.0 regardless of their reading, so a
# low soft_floor_usd never biases routing toward an attenuated provider just
# because a telemetry provider's allowance reads near zero under normal use.
#
# This knob is deliberately NOT exposed in the admin UI persisted-config
# allowlist (_CONFIG_ALLOWLIST in admin.py) or _ProviderUpdateBody; enabling
# or tuning it requires editing this file and restarting llm-router.service
# (the dispatcher's module-level cfg binds at import).
credit_attenuation:
enabled: false
soft_floor_usd: 5.0 # balance >= this -> multiplier 1.0
zero_floor_usd: 0.0 # balance <= this -> max_multiplier
max_multiplier: 5.0 # maximum cost-inflation at/below zero floor
refresh_seconds: 300 # cache duration for resolved multipliers
context:
safety_factor: 0.75 # fraction of advertised context treated as usable
default_output_reserve_tokens: 4096
# Ceiling on the output reserve, as a fraction of the usable window. A
# provider's advertised max_output_tokens is normally a small per-request
# cap, but OpenRouter reports max_completion_tokens -- "the most you may
# ASK for", 0.8-0.9 of context on a dozen rows. Subtracting that whole
# left 12 models at an effective context of 0, silently unroutable.
# Only the catalog-derived reserve is capped; a per_model_overrides
# reserve is a measurement and is used as written.
max_output_reserve_fraction: 0.5
# Per-model exceptions to the two settings above, for a row whose real
# limits you have measured. The global factor has to hold for the whole
# catalog, so it is deliberately pessimistic; a model you have actually
# pushed to its limit deserves its own number.
#
# Both keys are optional and each falls back to the global on its own.
# An override of 0 reserve tokens means zero, not "unset".
#
# per_model_overrides:
# qwen3.6-35b:
# safety_factor: 0.85
# output_reserve_tokens: 8192
#
# Takes effect on the next `python poller.py` -- effective_context_window is
# computed at poll time, not per request.
per_model_overrides: {}
tiers:
# Maps a tier number to a human label, purely for logging/dashboards.
1: "cheap / simple"
2: "mid / general"
3: "frontier / high-stakes"
tiering:
# Auto-tiering pass knobs. cheap_completion_max is the completion-cost
# (per 1M tokens) ceiling below which a model is eligible for tier 1.
# model_tiers overrides the heuristic per model_id and applies to ALL
# providers (limitation vs a (model_id, provider) key).
cheap_completion_max: 1.00
# Advertised context_window at or above which a model is NOT eligible for
# tier 1, whatever it costs. Tier is a capability FLOOR (routing drops any
# row with tier < required_tier), so tier 1 means "simple work only" — and
# deciding that on price alone excluded deepseek-v4-flash from every tier-2
# request purely for being $0.28/1M, despite a 1M window and 1.00 on all
# three coding categories.
#
# 512000 sits in the empty band between the catalog's 256K class (262128)
# and its 1M class (1048560) — a 2x margin either side, so it is not fitted
# to any one model. gemma-4-31b (256K) stays tier 1; the 1M rows do not.
tier1_context_max: 512000
model_tiers:
# This model lives on Ollama, not the NeuralWatt cloud catalog, so the
# tiering heuristic never sees it -- src/tier.py OVERWRITES models.tier on
# every run, so this pin is the ONLY thing that survives. If you swap the
# local dispatch model, change the KEY here too or it silently loses its
# tier. Pinned to tier 1 (cheap / simple) because it is targeted at
# lightweight summarization and diff-checking — precisely the use-cases
# tier 1 was designed for.
qwen2.5-coder-router:14b: 1
proficiency:
# Blending rule: leaderboard vs self-eval, once self-eval sample size
# crosses the threshold below. Below threshold, leaderboard score alone
# is used so thin self-eval data doesn't dominate.
self_eval_min_samples: 10
leaderboard_weight: 0.3
self_eval_weight: 0.7
# Prior strength (k) for the empirical-Bayes shrunken estimate: the peer
# prior contributes k pseudo-observations, pulling each noisy per-model
# score toward the global average. 20 pseudo-observations damps thin
# self-eval data (n < 100 per model) without erasing the per-model signal.
outcome_prior_strength: 20
categories:
- coding_general
- coding_refactor
- debugging
- docs_writing
- summarization
# file_summarization + diff_checking are served by the local dispatch model
# (see local_dispatch_models: section below).
# They are kept adjacent to summarization since this model is targeted
# at lightweight summarization and diff-checking tasks.
- file_summarization
- diff_checking
- translation
- reasoning_math
- tool_use_agentic
- general_chat
exploration:
# Epsilon-greedy exploration: on the epsilon share of requests, picks the
# hard-filter-eligible candidate with the fewest outcome samples (tie-break:
# lowest cost) instead of the highest-score model. Deliberately ON by default
# because the system is a ranking loop — without periodic exploration it
# converges on whatever happens to be sampled most, creating exposure bias.
# Adjust once real exploration traffic (was_exploration=True in
# route_decisions) shows how often the exploration path picks differently
# from the greedy path.
enabled: true
# Probability of exploration per request. 0.03 = ~3% of requests take an
# exploratory path, giving ~97% exploit on known winners while still
# occasionally sampling under-explored candidates.
epsilon: 0.03
# Cap: exploration candidate cost must be <= this multiple of the ranking
# winner's cost. 4.0 gives headroom — an under-sampled model can be more
# expensive than the greedy winner without eating the explore budget on
# wildly off-target picks.
max_cost_ratio: 4.0
# Only tiers 1 and 2 models are eligible for exploration. Tier 3 (frontier)
# is too expensive to spend on random sampling — save the explore budget for
# models where being wrong costs less.
max_tier: 2
escalation:
enabled: true
max_tier: 3
# Bump the tier when the classifier is unsure of its own call. DEFAULT OFF:
# this pays frontier prices on a hunch, before anything has gone wrong. The
# iteration budget below spends after a check has actually failed, which is
# strictly better on both mandates — the cheap attempt usually succeeds and
# costs nothing extra, and when it fails you have evidence.
preemptive_on_low_confidence: false
min_confidence_before_bump: 0.6
iteration:
# A tier is not only a capability floor, it is a budget for getting the
# answer right. These are corrective attempts AFTER a verification failure,
# not speculative retries.
#
# Retries are matched to the failure: a truncated answer gets a bigger token
# budget on the SAME model (a different one would also run out), while a
# malformed answer escalates to the next-best candidate (more tokens will
# not make unparseable output parse).
enabled: true
attempts_by_tier:
1: 0 # cheap/simple — one shot; iterating costs more than it is worth
2: 1
3: 2
# Interactive requests are capped below their tier's budget regardless of
# tier: every retry doubles time-to-answer, and in interactive use latency
# IS a quality loss. Batch work does not care.
max_attempts_interactive: 1
pinch:
# Relevance-based context pruning (ported from the MIT-licensed llmrouter's
# "pinch"). This is an OPTIONAL, pre-dispatch stage: when a conversation
# exceeds budget_tokens, the provider-bound messages are trimmed BEFORE any
# paid token is sent upstream. User/assistant/system messages are always
# kept verbatim; only old TOOL RESULTS are shortened or dropped, because
# they carry the bulk of a long agent session's tokens and are least needed
# in full by the time the next turn is answered.
#
# On by default. Tool results can only be dropped safely because tool
# outputs are idempotent enough for a placeholder; a wrong guess here loses
# context, so start conservative (large budget, small reduction) and watch
# route_decisions / pinch stats on real traffic before widening it.
enabled: true
budget_tokens: 50000
keep_last_turns: 4
max_summarize_chars: 4000
# keep_last_turns has no size limit inside it: an entire autonomous
# tool-call loop with no new user message can be one protected turn, and
# one outsized tool result inside it (a full verbose test run, a huge file
# read) ships verbatim regardless of size. Measured live 2026-09-06: a
# 324k-token conversation shrank only ~8% because nearly all of it sat
# inside the protected window. This closes that gap: any tool result
# inside the protected window over this many characters still gets the
# same head/tail elision candidates get. Deliberately a much higher bar
# than max_summarize_chars -- recent results are more likely to still
# matter -- so it only catches true outliers. Set to null to disable.
protected_max_chars: 20000
# Prefix-stability probe (Wave 1 item 1.3 of plans/token-waste-waves.md).
# The provider bills the longest byte-identical PREFIX of a prompt at the
# cached rate, so rewriting an early message re-bills everything after it.
# The uniform pruning path compresses a contiguous positional region, so an
# append cannot disturb it; the relevance path below compresses a prefix of
# a relevance-ORDERED list, so a growing deficit pulls in one more candidate
# each turn at an arbitrary message POSITION. Offline that rewrites 75% of a
# payload's tokens; the live cache rate on pruned turns is 0.924, which is
# not what that should look like. This settles it on real traffic.
#
# When on, each decision row records prefix_divergence_index,
# prefix_tokens_after_divergence and prefix_prev_message_count -- three
# integers. HASHES ONLY, and not even those: the per-message digests live in
# process memory for exactly one turn and never reach the database, so
# nothing reversible is stored and a restart costs one comparison per
# session. Measured on the live median payload (98k tokens in, 74k out):
# ~1.0 ms per turn, against a request path whose floor is a provider
# round-trip of 1.4-2.0 s, plus ~10 bytes per decision row.
prefix_probe: true
relevance:
# Off by default, matching every other new-and-unproven knob in this
# project — and specifically requires pinch.enabled too, since this has no
# effect otherwise. Ship it, watch route_decisions / pinch stats on real
# traffic, then decide the default.
enabled: true
# An EMBEDDING model, not a chat model — this must not point at
# classifier.model or verification.model. Pull one on the same Ollama:
# ollama pull nomic-embed-text
model: "nomic-embed-text"
base_url: "http://localhost:11434/v1"
timeout_seconds: 10
# Below this many trim-eligible candidates, skip the embedding call
# entirely and fall back to uniform compression — a network round trip
# to rank one candidate decides nothing.
min_candidates: 2
session_cache:
# In-memory per-session classification cache. Remembers the last
# task_category / task_tier decision for each session for a few minutes, so
# a long agent session skips the ~1-2s classifier round-trip on every turn.
# Capability flags (tools / images / json) are NEVER cached — they are read
# fresh from each request. Fallback classifications are NEVER cached. No
# persistence: a restart just reclassifies each session once.
#
# Off by default, matching every other new-and-unproven knob in this
# project: ship it, watch route_decisions.source="cached" on real traffic,
# then decide the right default.
enabled: true
staleness_seconds: 1200
circuit_breaker:
# Passive availability circuit breaker, on by default. When enabled, a model
# that returns 5xx is temporarily excluded from routing with exponential
# backoff; recovery is passive (the next real request becomes the probe once
# the cooldown passes). It has a low-risk failure mode even when wrong.
enabled: true
initial_cooldown_seconds: 30
max_cooldown_seconds: 600
backoff_multiplier: 2.0
routing:
# Access gating is prose-only in the NeuralWatt catalog ("Private preview
# (grant-gated)", "(Canary)"), so the poller parses it into access_level and
# routing excludes anything not listed here. Add 'preview'/'canary' only if
# the account actually holds the grant — otherwise dispatch earns a 403.
allowed_access_levels:
- public
# '-flex' rows are held server-side during peak until a capacity gap opens.
# That's correct for overnight/batch agent work and wrong for anything
# interactive, so a request has to opt in via latency_tolerance.
default_latency_tolerance: interactive # 'interactive' | 'batch'
# Operator's default stance on routing to '-flex' serving-class rows for
# requests that do not state one explicitly. A 4-position scale:
# no-flex never route to a flex row
# auto decide per request (the default; keeps existing behavior)
# prefer-flex flex first, standard as fallback
# force-flex flex only
# 'force-flex' is the dangerous global default: it bypasses the
# latency_tolerance: interactive hard filter for EVERY request, so even a
# request that opted into interactive routing would admit rows that are held
# server-side during peak. Use it only when you are certain the caller can
# tolerate flex latency globally.
default_flex_preference: auto
# Bare `auto` resolves to this profile. This is the profile name the router
# uses when the client does not specify one explicitly. It must name a
# built-in profile or a config-defined profile; the value is validated at
# load and will be editable from the admin Controls page in a later wave.
default_profile: "default"
# Minimum tool_use_agentic proficiency required of a model when the REQUEST
# carries tool definitions. A filter, not a weight, because it is a
# capability requirement rather than a preference.
#
# It does NOT ask whether the task is agentic — it asks whether the model
# can be trusted with tools that are on the table. The measured failure is
# the second one: deepseek-v4-flash scores 0.33 here, and the recorded case
# is a NON-agentic prompt ("it is 1:20pm, my meeting is at 3pm, how many
# minutes?") where it called two tools instead of subtracting. A model that
# over-reaches is a hazard wherever tools exist, not only where a classifier
# would say "agentic" — which it cannot do anyway: asked to identify six
# unambiguous tool-use prompts, the local models managed 2/6 and 1/6.
# Whether tools are present is stated in the request body. Read it.
#
# 0.5 sits in the empty band between the only two values the catalog
# currently holds (0.33 and 1.00), so it is not fitted to either. A model
# with NO measured tool score is unproven rather than proven bad, and is not
# dropped.
#
# CURRENTLY DISABLED (null) pending experimentation. The trade being
# measured: opencode sends `tools` on essentially every request, so with the
# filter on, deepseek-v4-flash is excluded from ordinary agent traffic and
# the ~7x cost advantage on coding routes goes unused. With it off, that
# advantage applies — and a model measured at 0.33 on tool use handles
# requests where tools are on the table.
#
# What would settle it is outcome data, not another benchmark: run with it
# off, let POST /outcome report real pass/fail, and compare
# tool_use_agentic proficiency for deepseek before and after. That is the
# one signal here that knows whether the work actually worked.
min_tool_proficiency:
tool_use_category: tool_use_agentic
# Request-side capability gates. These read the request body (image parts,
# response_format) and hard-restrict to models whose catalog row declares the
# capability. Unlike min_tool_proficiency these are NOT proficiency gates —
# a wrong guess is a guaranteed provider 400, so they gate by default and
# fail closed when the catalog flag is unknown.
require_vision: true
require_json_mode: true
# There is deliberately no flex discount knob. Whether a flex row is usable
# at all is a hard filter above (latency_tolerance), not a price adjustment
# -- being held server-side during peak is a latency property, and the
# catalog advertises flex and standard at the same token price anyway.
local_vision:
# Local Ollama vision fallback. Used ONLY when a request carries image
# parts and routing finds NO cloud model that supports_vision — the cloud
# catalog currently excludes the cost leader (deepseek) on vision, so a
# fallback is what keeps image requests working instead of 422ing.
# It speaks the OpenAI-compatible /v1 surface, so this is an Ollama endpoint
# and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host.
enabled: true
base_url: "http://localhost:11434/v1"
api_key_env:
# A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag.
# Measured live: the base tag comes up at Ollama's own default num_ctx
# (32768) and costs 9.4GB loaded — resident alongside the classifier's
# pre-fix 14GB, that left 1.9GB free on a 24GB card. num_ctx cannot be set
# per-request here: verified live that Ollama's OpenAI-compatible endpoint
# (0.22.0) silently ignores it under every field shape tried, so it has to
# be baked into the model tag itself:
# printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision
# ollama create qwen3-vl-router:4b -f Modelfile.vision
#
# Reduced 16384 -> 8192 on 2026-09-03 to buy back KV cache. Most of this
# model's footprint is KV, not weights: 3.3GB on disk but 6.9GB resident at
# 16384, and 5.6GB at 8192. That 1.3GB is what lets the dispatch model run
# at num_ctx 32768 (17GB) and still leave vision RESIDENT -- 23.2GB of 24GB
# together. Without it, loading vision EVICTS the dispatch model entirely
# and the next classification pays a ~6s cold reload on the latency floor.
#
# This is the one path that still has no measured ceiling: it gets the RAW
# message list (dispatcher.py calls it BEFORE pinch pruning), so 8192 is a
# judgement call, not a derived minimum. If image requests start failing on
# context, raise this FIRST and drop the dispatch model to num_ctx 16384 to
# pay for it. Budget guards (max_images, max_image_bytes) bound the image
# side but not the conversation text around it.
#
# Do NOT swap this for a 1B-class model to save memory. Tested 2026-09-03:
# moondream (1B, 2048 ctx) answered a real dashboard screenshot with "a
# spreadsheet ... possibly related to business decisions or financial
# analysis" -- it read no title, no column, no value, and confabulated a
# plausible description instead. qwen3-vl:4b read the page title, quoted the
# subtitle verbatim and recovered the row count. Screenshots are the actual
# workload here, and dense small text is exactly where tiny VLMs fail.
model: "qwen3-vl-router:4b"
timeout_seconds: 60
max_images: 4
max_image_bytes: 9437184
local_dispatch_models:
# Local Ollama models that can serve dispatch (not just classification/vision).
# Each entry is a model that the router can route traffic to — the same path
# that picks cloud models from the NeuralWatt catalog, but calling a local
# OpenAI-compatible endpoint instead.
#
# These models MUST ALSO appear under proficiency.categories in eligible_categories
# and MUST have a tier pin under tiering.model_tiers, because src/tier.py
# OVERWRITES models.tier on every run (the config override is the only pin).
#
# To add a model: copy this block and change the values. Keep the comment
# describing how to build the model tag (Modelfile recipe), because the num_ctx
# baked into the tag is what Ollama's /v1 endpoint actually uses — setting it
# per-request through the OpenAI API is not supported by Ollama.
- # DEFAULT. Measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5), and a
# sound starting point for any 24-32GB card -- the numbers below are real,
# not aspirational. It is a default rather than a recommendation only in
# the sense that YOUR hardware and YOUR taste in models should win:
# docs/local-models.md documents the measurement method so you can swap in
# whatever you actually want to run and defend the choice with your own
# numbers instead of inheriting these.
#
# Dormancy under the DEFAULT profile (measured 2026-09-02/03): quality-gated
# ranking — the local row must score within objective.quality_tolerance
# (0.10) of the cloud leader before its price advantage is even consulted.
# Today it does not: 0.767 vs 0.95 on file_summarization (gap 0.183), so
# this row is dormant under the default profile. That is by design, not a
# bug. Do NOT widen quality_tolerance. It fires as a fallback when the cloud
# is unavailable; see docs/routing.md § Local dispatch branch.
#
# qwen2.5-coder-router:14b serves BOTH local dispatch and classification,
# so only one model stays resident. Selected 2026-09-02/03 by measurement:
#
# file_summarization (n=6, judge-scored -- treat +-0.15 as a tie):
# deepseek-v4-flash (cloud) 0.95 <- local costs ~0.18 of quality
# qwen2.5-coder:32b 0.833 (22GB: evicts vision, 3.4s classify)
# qwen2.5-coder:14b Q8_0 0.817 (19GB, 1.5x slower, gain within noise)
# qwen2.5-coder:14b Q4_K_M 0.767 <- chosen
# nemotron-mini:4b 0.25 (the plan's original pick; FABRICATED
# on 3 of 6, and detected 0/4 bugs)
#
# classification was 11/14 with 0 hard failures for EVERY variant above --
# quantization and context size changed only speed, never accuracy.
# Q4_K_M @ 32k: 1.14s mean. Q8_0 @ 32k: 3.79s, because it needs 24GB and
# Ollama spills 17% to CPU (watch the "17%/83%" column in `ollama ps`).
#
# num_ctx 32768 costs ~4GB of KV cache over 16384 (13GB -> 17GB resident).
# That is affordable here ONLY because the vision model was retagged to
# num_ctx 8192; together they are 23.2GB of 24GB, with ~800MB spare. On a
# smaller card, drop this to 16384 first -- KV cache is the cheapest GB to
# reclaim, and the classifier never needs it (classifier.max_input_chars
# clamps to 8000 chars, ~2.7k tokens).
#
# Modelfile recipe (run on an Ollama host):
# ollama pull qwen2.5-coder:14b
# printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch
# ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch
model_id: "qwen2.5-coder-router:14b"
base_url: "http://localhost:11434/v1"
api_key_env:
timeout_seconds: 180
context_window: 32768
max_output_tokens: 2048
tier: 1
# Restricted to file_summarization and diff_checking. The router's hard
# filter (routing.py) will exclude this model for any other category. Do
# not widen this without measuring the new category first -- this model
# scored 0/4 at detecting bugs before the task set was fixed, and a
# confidently wrong answer is worse than an honest error.
eligible_categories:
- file_summarization
- diff_checking
verification:
# Structural checks (parse the code, never run it) are free, pure Python and
# always on — they need no model and run anywhere, down to an RPi.
# This section governs the LOCAL LLM check, which is not free.
#
# Set local_llm_enabled: false on a host with no usable local inference.
# Structural checking and POST /outcome both keep working; only the
# refusal/incoherence class of failure stops being caught.
local_llm_enabled: true
# This check speaks Ollama's NATIVE API (/api/chat with think=False), which
# no cloud provider offers, so it is configured SEPARATELY from the
# classifier rather than derived from it. Deriving it meant that pointing
# classification anywhere else sent these requests to <that host>/api/chat.
#
# Point it at the SAME Ollama as the classifier when that is across a VPN;
# leave the model null in that case, since the hosts then match.
base_url: "http://localhost:11434"
# null means "whatever classifier.model is", correct only while both run on
# the same local Ollama — config load REFUSES the null once the hosts differ,
# because the fallback would name a model this Ollama has never heard of and
# the verifier would fail silently. Stated explicitly here anyway, matching
# classifier.model exactly (see that field's comment) so both calls hit the
# SAME resident Ollama instance, loaded once at its Modelfile-tagged
# context size rather than two separately-sized copies.
model: "qwen2.5-coder-router:14b"
# Only check answers this large. Measured on real traffic: a local check
# costs ~15% of a median 193-token answer, so it would only pay if such
# answers failed more than ~15% of the time. At 1,500 completion tokens the
# break-even failure rate drops to ~1.9%, which is plausible. Below the
# threshold a check costs more than the risk it removes.
min_completion_tokens: 600
# The check runs AFTER the response has gone back to the client, so it never
# adds its ~6s to anyone's latency. It exists to learn which models fail on
# real work, not to gate answers.
timeout_seconds: 60
max_output_tokens: 1024
# How far back /outcome looks when a report carries no request_id and no
# matching directory. A test run follows the completion that caused it within
# seconds, so this is deliberately short: a wide window sweeps in sessions
# that finished long ago and makes every report look ambiguous. If more than
# one conversation was active inside it, the report is refused rather than
# guessed at.
outcome_attribution_window_seconds: 120
freshness:
stale_after_days: 3
# Router refuses to route to a model whose row is stale/deprecated,
# regardless of how good its score would otherwise be.
exclude_stale: true
exclude_deprecated: true
# Allowlisting a model does nothing until a poll ingests it, and the timer
# runs every 2 hours -- so a model added through /admin sat invisible until
# the next tick, looking like the allowlist had not worked.
#
# Debounce, not delay: each further edit restarts the clock, so adding five
# models one at a time costs one poll, not five. 0 disables the trigger and
# leaves the timer as the only path.
repoll_after_allowlist_change_seconds: 5.0
database:
path: "router.db"
classifier:
# The local LLM that classifies each task before a cloud model answers it.
# "Local" means your hardware, not necessarily this machine — the box with
# the GPU is usually not the laptop you are typing on. It speaks plain
# OpenAI-compatible chat completions, so point it wherever Ollama lives:
#
# same machine base_url: http://localhost:11434/v1
# over a VPN base_url: http://<vpn-ip>:11434/v1
# (serving host needs deploy/ollama-over-vpn.conf; Ollama
# binds loopback-only by default and will refuse)
#
# Any OpenAI-compatible endpoint works, so a cloud model can classify too —
# set api_key_env for one that checks a key. Worth knowing before assuming
# local is the cheap option: on five prompts NeuralWatt's deepseek-v4-flash
# classified in 1.02s mean against qwen3.5's 11.58s on an RTX 6000, agreed
# with the label 5/5 against 2/4, and cost $0.093 per thousand calls. Local
# inference is not free, it is unbilled.
provider: "ollama"
base_url: "http://localhost:11434/v1"
# Unset means unauthenticated, which is the Ollama case. Name the env var
# holding the key when the endpoint actually checks one.
api_key_env:
# A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag
# — must match a model `ollama list` reports. Ollama loads a model at its
# library Modelfile's default context unless told otherwise, and the base
# tag never was: measured live, mistral-nemo:12b came up at num_ctx=32768
# (14GB of a 24GB card) though max_input_chars + the system prompt +
# max_output_tokens need well under a quarter of that. num_ctx cannot be
# set per-request here: verified live that Ollama's OpenAI-compatible
# endpoint (0.22.0) silently ignores it under every field shape tried (a
# 200 comes back, the loaded context never changes), so it has to be baked
# into the tag itself:
# printf 'FROM mistral-nemo:12b\nPARAMETER num_ctx 8192\n' > Modelfile.router
# ollama create mistral-nemo-router:12b -f Modelfile.router
# 8192 leaves 2x headroom over the real requirement. verification.model
# points at this same tag (see its own comment) so classify and verify
# share one resident instance at one context size, rather than risking a
# reload thrash from two differently-sized copies of the same base model.
model: "qwen2.5-coder-router:14b"
# A cold Ollama took 43s to answer the first classification, which blew the
# old 30s ceiling and returned 503 to the client. Warm it is ~5s. The
# ceiling is for a cold model load, not the steady state.
timeout_seconds: 120
# Classification must be reproducible: at the default temperature the same
# task was classified tier 2 then tier 1 on consecutive calls, which routed
# it to two different models. Routing that changes under an identical
# prompt is untraceable.
temperature: 0
# Hard cap on the classifier's generation, to bound a failure mode that
# cascades on REASONING models: they emit a long chain of thought, blow past
# timeout_seconds, and Ollama keeps generating after the client gives up AND
# serializes per model, so one runaway request queues every later one behind
# it. Measured on qwen3.5, which failed this way on 4 of 19 calls — each one
# a 15s wait ending in a silent fallback.
#
# The current default (mistral-nemo) does not reason, so this does not bind
# for it. Kept anyway: it costs nothing when unused and is the only thing
# standing between a swapped-in reasoning model and that cascade. If you do
# swap one in, note 256 was too tight — the trace consumed the whole budget
# and the model was truncated before emitting any JSON.
max_output_tokens: 1024
# Where routing lands when the classifier times out, errors, or returns
# something unparseable. A local model being slow should degrade routing,
# not refuse the request — the caller is a coding agent that would rather
# have a mid-tier answer than a 502.
# Ceiling on the text handed to the classifier (head + tail, middle elided).
# It decides a category and a tier; it does not need the document, and
# feeding it one is harmful rather than merely wasteful. Measured on a ~20k
# token prompt: qwen3.5 spent 28.7s and returned empty (budget consumed by
# its reasoning trace), mistral-nemo spent 41.8s echoing the input back
# inside its JSON. Both land on source: "fallback" — the same answer an
# instant failure gives, after 30-40s of local inference.
#
# Nothing is lost: chat_completions measures the real conversation with
# estimate_prompt_tokens and takes the larger value, so required_context
# never depends on what the classifier saw. 0 disables clamping.
max_input_chars: 8000
# When the chat path supplies the previous turn as context (see the pinch /
# context notes), frame the classifier input as "Context: <prev> / Message:
# <current>" so a short follow-up inherits the prior turn's complexity
# instead of being classified in isolation as trivial.
context_framing: true
fallback_tier: 2
fallback_category: general_chat
# --- which implementation is PRIMARY -----------------------------------
# This is a peer concept to the fallback cascade below, not a replacement
# for it: whichever mode is primary, a failure still walks the SAME
# cascade (stale session -> session history -> cloud_fallback -> the
# fallback_tier/fallback_category above).
#
# local_llm (default) the model/base_url above, unchanged.
# cloud_llm a cloud model answers PRIMARY, not just as a fallback
# after local fails. Requires exactly one of
# cloud_primary or cloud_primary_auto below.
# local_encoder a small, non-generative classifier (see the encoder:
# block below) that cannot exhibit the runaway-reasoning
# failure mode documented above, at the cost of only
# producing task_category -- task_tier falls back to
# fallback_tier for this mode. Zero-shot, not trained on
# your traffic: this router never stores raw task text.
mode: local_llm
# Read only when mode: cloud_llm. Exactly one of these two:
# cloud_primary:
# base_url: https://api.neuralwatt.com/v1
# model: deepseek-v4-flash
# api_key_env: NEURALWATT_API_KEY
# timeout_seconds: 2
# max_output_tokens: 1024
# cloud_primary_auto: true # resolve the cheapest routable model, live
# Read only when mode: local_encoder. Every field has a default, so
# `encoder: {}` is enough to opt in.
# encoder:
# model: facebook/bart-large-mnli
# device: cpu
# confidence_threshold: 0.5
# --- fallback cascade ---------------------------------------------------
# When the local classifier fails, the router walks: stale session cache ->
# this session's history in route_decisions -> the optional cloud
# classifier below -> fallback_tier/fallback_category above. The first two
# steps are free and local.
#
# Global backoff after a classifier failure. This is what bounds cloud
# spend during a sustained outage: at most one cloud attempt per window
# across ALL requests, not one per request. It is also reused as the window
# in which an account-level provider refusal suppresses the cloud step,
# since an out-of-credit account makes that call a guaranteed waste.
cooldown_seconds: 30
# Degradation warning on /metrics. Silent below degraded_warn_min
# decisions in 24h, because a share computed over a handful of requests is
# noise and a flapping warning is one nobody reads.
degraded_warn_min: 20
degraded_warn_threshold: 0.5
# Optional. ABSENT BY DEFAULT, which is what makes the cascade cost
# nothing: with no cloud_fallback block the router degrades straight to the
# static guess. Uncomment and point it at any OpenAI-compatible endpoint to
# trade a little money for a real classification during a local outage.
# A configured api_key_env whose variable is missing degrades to the static
# guess rather than failing the request -- unlike the primary classifier,
# which raises, because this one is only reached when things already broke.
# cloud_fallback:
# base_url: https://api.neuralwatt.com/v1
# model: deepseek-v4-flash
# api_key_env: NEURALWATT_API_KEY
# timeout_seconds: 2
# max_output_tokens: 1024
response_format: "json" # ask Ollama to constrain output to valid JSON
# The labels a classifier may return. A SUBSET of proficiency.categories,
# validated as one at config load, and deliberately not the same list:
# proficiency.categories is the scoring axis (what a model is good at),
# this is the set a classifier is asked to choose between.
#
# tool_use_agentic is excluded. It describes what a turn mechanically DOES
# rather than what it is for, and an agent turn is always both -- measured
# 2026-09-14 on live traffic with the session cache off, 31 of 31
# consecutive turns classified tool_use_agentic and every one routed to the
# slowest model in the catalog (14.8s TTFT, p95 40s). Per-turn
# classification became accurate and routing got worse. The signal it was
# standing in for is already read exactly and for free from the request's
# own `tools` array by routing.min_tool_proficiency.
#
# WHAT THIS COSTS, so the next person does not discover it: with the
# classifier unable to emit tool_use_agentic, no new POST /outcome report
# attributes to that category, so its proficiency scores FREEZE at their
# current values (e.g. qwen3.6-35b at 0.902 over 154 samples). The tool
# filter keeps reading them; they simply stop accumulating. Accepted for
# now.
#
# Considered and REJECTED as the fix: attributing outcomes to
# tool_use_agentic whenever the request carried a `tools` array. opencode
# sends `tools` on essentially every request, so the score would converge
# on each model's overall pass rate and stop discriminating the exact thing
# the category exists to detect -- a model that OVER-reaches for tools on
# work that did not need them. A frozen honest score beats a live
# meaningless one.
#
# Omit this key entirely to mean "every proficiency category".
candidate_categories:
- coding_general
- coding_refactor
- debugging
- docs_writing
- summarization
- file_summarization
- diff_checking
- translation
- reasoning_math
- general_chat
# The dispatcher appends the authoritative category list from
# classifier.candidate_categories (above) to this prompt at call time. Do
# not enumerate the categories here as well — a hand-copied list drifts, and
# a category the model invents joins against nothing in the proficiency
# table.
system_prompt: |
You are a task router. Given a task description and any attached context,
respond with ONLY a JSON object with these fields:
{
"task_category": one of the allowed categories listed below,
"task_tier": integer 1-3, where 1 is cheap/simple, 2 is mid/general,
and 3 is frontier/high-stakes,
"required_context_tokens": integer estimate of prompt+context token count,
"confidence": float 0-1
}
dispatch_providers:
neuralwatt:
base_url: "https://api.neuralwatt.com/v1"
api_key_env: "NEURALWATT_API_KEY"
has_energy_telemetry: true
# false: NeuralWatt reports its bill in a top-level `cost` block (and as a
# `: cost {...}` SSE comment when streaming), so no request-side opt-in is
# needed and none is sent. See the openrouter entry below for what the
# flag does when it is on.
reports_cost_in_usage: false
enabled: true
openrouter:
base_url: "https://openrouter.ai/api/v1"
api_key_env: "OPENROUTER_API_KEY"
# Account-level balance poll; OpenRouter reports per-completion energy via
# a separate allowance_remaining field, so this URL is for the prepaid pool.
balance_url: "https://openrouter.ai/api/v1/credits"
has_energy_telemetry: false
# true: send OpenRouter's own `usage: {"include": true}` opt-in on every
# upstream request. `stream_options.include_usage` is the OpenAI spelling
# and gets token counts; this is the separate OpenRouter one, and without
# it the accounting block is simply absent -- no billed `usage.cost` and
# no `usage.prompt_tokens_details.cached_tokens`. Both ride on this one
# opt-in, so the two coverage figures move together.
#
# It does NOT bring generation timing: OpenRouter serves that only from
# its separate /api/v1/generation endpoint, so `duration_seconds` stays
# NULL on OpenRouter rows. Turn this off only for a provider that returns
# accounting unasked, or one that rejects the key.
reports_cost_in_usage: true
enabled: true
require_allowlist: true
local_compute:
# "Gaming mode", inverted: true means the router may use local hardware.
# Set it to false when you stop Ollama to give the GPU back to something
# else, and the router SKIPS every local call instead of discovering the
# outage one 120s timeout at a time: classifier, /health probe, local
# verification, the local-vision fallback, and local dispatch rows (which
# drop out of the candidate set rather than being picked and then failing).
#
# ONE flag that the call sites read, not a macro that rewrites
# verification.local_llm_enabled / local_vision.enabled / local_energy.enabled.
# Those keep their own meanings; this is an outer AND over all of them, so
# turning it back on restores exactly the state you left.
#
# Config load REFUSES false unless classifier.cloud_fallback is configured.
# Skipping the local classifier does not make classification remote, it makes
# it a static guess -- and that guess is recorded as general_chat, a fully
# scored category, so it looks like real classification afterwards.
# Uncommenting the cloud_fallback example above is the intended setup path.
enabled: true
local_energy:
# Local energy metering for the router's own Ollama calls (classifier,
# verifier, local-vision fallback). Cloud providers expose per-request energy,
# but these calls run on your own hardware and their electricity is real even
# though it never shows up on the NeuralWatt bill.
#
# OFF by default. Enable only after setting a real per-kWh rate below.
# Config load REFUSES enabled: true with a null tariff because a null rate
# would record cost_usd = 0 instead of "unknown".
enabled: false
# Power sampler. Only "nvidia_smi" is supported today. It shells out rather
# than importing pynvml so the router has no new dependency and degrades
# cleanly on non-NVIDIA hosts.
meter: "nvidia_smi"
# How often to sample GPU power while a local call is running. 0.25s was the
# original llmrouter pinch sample rate and is fine as a starting point.
sample_interval_seconds: 0.25
# YOUR real per-kWh electricity rate. Set this before enabling metering.
#
# Example (do not use this number unless it matches your bill):
# tariff_usd_per_kwh: 0.12
#
# Set 2026-09-02 from the user's Great Lakes Energy bill. This is the
# MARGINAL rate -- what one more kWh actually costs -- not the all-in
# effective rate of 0.183 ($410.15 / 2,244 kWh). The difference is the
# $52.30/month of fixed charges, which are paid whether or not the GPU
# runs, so loading them onto compute would overstate local cost by ~15%
# and bias routing toward the cloud for a reason that is not real.
# 14.742c energy + 0.188c PSCR + 0.316c EO = 15.246c, x 1.04 MI sales
# tax = 15.856c/kWh.
tariff_usd_per_kwh:
# YOUR grid carbon intensity in grams of CO2 equivalent per kWh. Optional:
# null disables the carbon estimate but still records energy and cost.
#
# Example (do not use this number unless it matches your grid):
# grid_intensity_g_per_kwh: 475
grid_intensity_g_per_kwh:
logging:
# Whether to write a row to energy_observations for every completion. Off
# means no cost accounting, no reference sweep and no /outcome attribution,
# so leave it on unless you are debugging.
#
# There is no log_path. Nothing ever wrote a file -- the dispatcher prints to
# stderr and the systemd unit hands that to the journal (`journalctl --user
# -u llm-router -f`), so the setting named a destination that did not exist.
log_energy_observations: true
# Whether to write a row to route_decisions for every routing decision
# (route | dispatch | chat | passthrough | local-vision). Off leaves the
# monitoring TUI's decision history empty; it does not change what routes.
log_route_decisions: true
# debug | info | warning | error.
#
# info gives one line per request: what it was classified as, which model won,
# what it cost, how long each stage took. debug adds why -- every candidate
# that was dropped and by which filter, the ranking with scores, the
# classifier's raw reply. It logs no conversation text at any level; prompts
# here run 60k-150k tokens and the journal is on disk.
#
# LLM_ROUTER_LOG_LEVEL overrides this, so a running service can be turned up
# without editing a tracked file:
#
# systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
# systemctl --user restart llm-router
level: info