Two measurement surfaces on /metrics. Neither is read by routing, and a test asserts that against the module source: a live incumbent-cache-pricing experiment is running, and a series that reached the ranker would confound it. cost_calibration -- routing.estimated_cost against the provider's own bill, per (provider, model), joined on request_id. The scale error is not the finding: a uniform overestimate reorders nothing, because the ranking is a comparison and every candidate moves together. The SPREAD does reorder, so that is the headline figure. Live over 168h, 4,152 joined requests: est/billed runs 1.62x (qwen/qwen3.6-35b-a3b, openrouter) to 12.86x (qwen3.6-35b-fast, neuralwatt) -- a 7.9x spread, wider than the 5x the brief was written against. The factors are REPORTED, not applied; whether they are stable enough to trust is the question this exists to answer, and this project has already mistaken one moment of a moving per-model quantity for a constant. The join needs the model as well as the request id. On the live database 5 of 4,143 rows pair a decision that selected qwen3.6-35b (neuralwatt) with a completion billed by qwen/qwen3.6-35b-a3b (openrouter) -- a cross-provider failover, and one model's estimate against another's bill. latency -- p50/p95 of router_wall_seconds and router_ttft_seconds, which landed recently and nothing read. These are the router's own clock, not the provider's duration_seconds, which is why OpenRouter is visible here at all: it reports no duration. z-ai/glm-5.3-flash, a model with "flash" in its name, measures p50 9.60s / p95 40.08s to first token against 1.58s / 3.09s for deepseek-v4-flash on NeuralWatt. Reported, never scored. The wall and TTFT sample counts are kept independent because TTFT is streaming-only by nature. No new warning class, deliberately. Every group is out of band on both series today, so a divergence warning would fire on all of them from the first run -- bare presence, which is the anti-pattern rejection_warnings exists to avoid -- and a warning is a form of trust these factors have not yet earned. Four knobs under objective, all report-only, all recorded in DELIBERATELY_NOT_IN_ADMIN with the reason: they change what /metrics shows, and not even what it warns about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
11 KiB
API surface reference. Back to README.
API Endpoints
The dispatcher binds 127.0.0.1:8080. No auth of its own — loopback
is the only thing standing between the open internet and your billing
allowance.
| Method | Path | Description |
|---|---|---|
GET |
/health |
Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set) |
GET |
/metrics |
Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth |
GET |
/events/decisions |
Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen |
POST |
/route |
Classify task, rank candidates, return selected model — no provider call, no cost |
POST |
/dispatch |
Same as /route, plus complete the provider call, stream response, log observation |
GET |
/v1/models |
OpenAI-compatible model list (router virtual models + catalog) |
POST |
/v1/chat/completions |
OpenAI-compatible completions — routes then proxies, streaming supported |
POST |
/outcome |
Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed |
GET |
/admin |
Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
/metrics returns a single JSON object with these top-level keys:
quota— per-provider billing shape fromquota_accounts():period(billing window with start/next_reset/elapsed_fraction/source),accounts[](list of per-provider dicts withprovider,shape,spend_usd, and type-specific blocks:planfor metered_plan,poolfor prepaid_credit,burnfor burn-rate metrics,creditwhen a balance URL was polled,energyfor kWh/calls),spend(aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), andalarm(plan-pace or stale-reading alert with kind/severity/headline). The oldby_provider/total_balance_usdshape and flat top-level keys were removed.coverage— routable-model counts (routable_modelscounts active models filtered onaccess_level IN (routing.allowed_access_levels), excluding any access-gated rows whoseaccess_levelis not in that list) with energy and proficiency data, plus aselectionsub-key reporting which active models (unfiltered by access-level) the router has actually picked over the configured window, and awarningsarray that now also includes entries for active models that can never enter the candidate set (structural-gate failures: zero context window or broken tier derived from the catalog, detected viarouting.rejection_reasonunder the most permissive request — one token, tier 1, batch latency). Existing warnings cover missing energy data, missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and classifier source degradation. Acachesub-key carries the prefix-cache series fromcache_rate_series():window_hours,assumed_cache_rate, aggregateobservations/prompt_tokens/cached_prompt_tokens/cache_rate, and aby_modellist of the same figures per(provider, model_id). The rate issum(cached_prompt_tokens) / sum(prompt_tokens)over rows the provider actually reported a count for (cached_tokens_source = 'reported'), excludingseed_referencesweeps. Two warnings read it:cache rate:when the aggregate diverges fromobjective.assumed_cache_rateby more thanobjective.cache_rate_warn_margin, andcache rate outlier:when one(provider, model_id)group does. Both needobjective.cache_rate_warn_min_observationsreported-cache rows first.recent_decisions— last 50 rows fromroute_decisionsper_model— per-model aggregates over the last 30 days ofenergy_observationsverdict_mix— counts by verification verdict over the last 7 daystop_proficiency— top models byblended_scoreforcoding_generalgenerated_at— ISO8601 timestamppinch— context-pruning aggregation (whenpinch.enabled):calls_30d,pruned_calls_30d, andshare_pruned(share of decisions pruned), plus median and totaltokens_savedover 30 days and estimated dollars saved at the blended rate — exposes no conversation textcost_calibration— fromcost_estimate_calibration(): how farrouting.estimated_costis from the provider's own bill, per(provider, model_id). Reported only — nothing applies it toestimated_costor to ranking.route_decisionsjoined toenergy_observationsonrequest_idand on the selected model and provider (the id alone pairs a failover's estimate with another model's bill), overobjective.cost_calibration_window_hours, excludingseed_referenceon both sides and rows with no positivecost_usdorest_cost_usd. Keys:window_hours,min_observations,observations,est_cost_usd,billed_cost_usd,overestimate_ratio(est/billed),correction_factor(billed/est — what would be multiplied in),spreadwithspread_low/spread_high, and aby_modellist of the same figures plusest_per_request_usd,billed_per_request_usdandsufficient.spreadis the figure that matters: a uniform scale error reorders nothing, while the spread in that error does. Groups belowobjective.cost_calibration_min_observationsare listed withsufficient: falseand excluded fromspread. No warning class reads it.latency— fromlatency_series(): p50 and p95 of the router-observedrouter_wall_secondsandrouter_ttft_secondsper(provider, model_id)overobjective.latency_window_hours, excludingseed_reference. A report, not an objective — nothing in the ranking reads it. These are the router's own clock, not the provider'sduration_seconds, which is why OpenRouter models are visible here at all (OpenRouter reports no duration). Keys:window_hours,min_observations, aggregatewall_observations/ttft_observations/wall_p50/wall_p95/ttft_p50/ttft_p95, and aby_modellist of the same pluswall_sufficient/ttft_sufficient. The wall and TTFT counts are independent because TTFT is streaming-only; a group can be sufficient on one and not the other. Degrades to an empty series on a database that predates the two columns. No warning class reads it.
It exposes no conversation text, prompts, or session_dir; it is bound to
loopback and unauthenticated exactly like /health.
Input to /route and /dispatch can include task_category, task_tier,
and required_context_tokens overrides — these skip the classifier, useful
for testing routing without the classifier in the loop.
Named routing profiles (auto:<profile>). The general form is
auto:<profile>; auto alone is shorthand for auto:default. The router
resolves the profile name against the built-in set and any custom entries in
config/config.yaml under profiles:. Profiles narrow the candidate set but
do not override the quality-first ranking objective.
| Profile | Effect |
|---|---|
auto:default |
Normal quality-first routing, -flex rows excluded |
auto:batch |
Admits -flex rows for overnight and async work |
auto:locality |
Restricts to the ollama-local provider only |
auto:onlycheaps |
Limits to models priced at ≤ $0.50 per 1M completion tokens |
auto:bigboybritches |
Restricts to tier-3 (frontier) models only |
From opencode: a profile is just another entry under
provider.llm-router.models in opencode.json — the object's key is the
literal model id opencode sends upstream, so auto:batch already works this
way (see the repo-local opencode.json). To make auto:locality,
auto:onlycheaps, or auto:bigboybritches selectable from opencode's model
picker instead of typing an override, add one entry per profile:
"auto:onlycheaps": {
"name": "auto (cost ceiling)",
"limit": { "context": 782324, "output": 16384 },
"modalities": { "input": ["text", "image"] }
}
limit/modalities can be copied from the auto entry unchanged — a
profile narrows which models are candidates, not the context window or
input types the router itself accepts. See clients.md for the
full opencode wiring.
An unknown profile name raises HTTP 422 and lists the valid names.
The profile field is also accepted on POST /route and POST /dispatch
to select a profile inline per request.
Ask for any real model id in /v1/chat/completions and it dispatches
directly, still logged — routing is transparent, not opaque.
Streaming (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt
emits energy and cost as SSE comment lines (: energy {...}) before
data: [DONE] — ordinary clients ignore comments, so the stream flows
untouched while the router scrapes telemetry on the way past. Without this,
streamed calls would log no energy at all.
Verification headers (non-streaming): When streaming is not used, the
structural verification verdict surfaces in the X-Router-Verification header
so a client can inspect it without parsing the response body. Valid values:
ok, truncated, malformed, unverifiable, none.
Capability 422s: When no model survives the hard filters, the 422 names the
active constraints. That now includes "vision-capable model" or
"json-mode-capable model" when the request carried images or a JSON-mode
response_format, alongside the existing context/tier/latency/tool reasons.
POST /outcome: everything else the router records is a proxy — structural
checks know whether code parses, the local LLM check guesses whether prose
looks right, neither knows whether the answer did the job. The client does,
because it ran the tests:
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
An unknown request_id returns 404 rather than being quietly accepted, so a
client whose reports go nowhere finds out. The lookup checks
energy_observations first (cloud), then local_energy_observations (local
rows carry the same request_id/session_dir attribution) — cloud wins on a
rare id collision.
Unlike the structural/local-LLM checks — which only ever record failures —
/outcome folds both directions into proficiency via feedback.py: a
false report counts against the model same as any other verification failure,
but a true report counts too. It's also the only quality signal that survives
streaming, since a retry can't reach a response whose bytes are already gone,
while a report arrives afterward and works either way.
Pinning and auto behavior for local models
model: "auto" will route eligible tasks to qwen2.5-coder-router:14b when the
classifier returns file_summarization or diff_checking and the local row
survives the filters; otherwise a cloud model is selected as usual. Pin any
local model_id (e.g. qwen2.5-coder-router:14b) and the request goes straight
to that Ollama tag. /v1/models lists local rows with owned_by: "ollama-local".