Files
6krrt/docs/api.md
adlee-was-taken a26765c6c0 feat(metrics): report the cost-estimator's error and the latency nobody could see
Two measurement surfaces on /metrics. Neither is read by routing, and a test
asserts that against the module source: a live incumbent-cache-pricing
experiment is running, and a series that reached the ranker would confound it.

cost_calibration -- routing.estimated_cost against the provider's own bill,
per (provider, model), joined on request_id. The scale error is not the
finding: a uniform overestimate reorders nothing, because the ranking is a
comparison and every candidate moves together. The SPREAD does reorder, so
that is the headline figure. Live over 168h, 4,152 joined requests: est/billed
runs 1.62x (qwen/qwen3.6-35b-a3b, openrouter) to 12.86x (qwen3.6-35b-fast,
neuralwatt) -- a 7.9x spread, wider than the 5x the brief was written against.
The factors are REPORTED, not applied; whether they are stable enough to trust
is the question this exists to answer, and this project has already mistaken
one moment of a moving per-model quantity for a constant.

The join needs the model as well as the request id. On the live database 5 of
4,143 rows pair a decision that selected qwen3.6-35b (neuralwatt) with a
completion billed by qwen/qwen3.6-35b-a3b (openrouter) -- a cross-provider
failover, and one model's estimate against another's bill.

latency -- p50/p95 of router_wall_seconds and router_ttft_seconds, which
landed recently and nothing read. These are the router's own clock, not the
provider's duration_seconds, which is why OpenRouter is visible here at all:
it reports no duration. z-ai/glm-5.3-flash, a model with "flash" in its name,
measures p50 9.60s / p95 40.08s to first token against 1.58s / 3.09s for
deepseek-v4-flash on NeuralWatt. Reported, never scored. The wall and TTFT
sample counts are kept independent because TTFT is streaming-only by nature.

No new warning class, deliberately. Every group is out of band on both series
today, so a divergence warning would fire on all of them from the first run --
bare presence, which is the anti-pattern rejection_warnings exists to avoid --
and a warning is a form of trust these factors have not yet earned.

Four knobs under objective, all report-only, all recorded in
DELIBERATELY_NOT_IN_ADMIN with the reason: they change what /metrics shows,
and not even what it warns about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-16 14:34:27 -04:00

11 KiB

API surface reference. Back to README.

API Endpoints

The dispatcher binds 127.0.0.1:8080. No auth of its own — loopback is the only thing standing between the open internet and your billing allowance.

Method Path Description
GET /health Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set)
GET /metrics Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth
GET /events/decisions Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen
POST /route Classify task, rank candidates, return selected model — no provider call, no cost
POST /dispatch Same as /route, plus complete the provider call, stream response, log observation
GET /v1/models OpenAI-compatible model list (router virtual models + catalog)
POST /v1/chat/completions OpenAI-compatible completions — routes then proxies, streaming supported
POST /outcome Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed
GET /admin Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits)

/metrics returns a single JSON object with these top-level keys:

  • quota — per-provider billing shape from quota_accounts(): period (billing window with start/next_reset/elapsed_fraction/source), accounts[] (list of per-provider dicts with provider, shape, spend_usd, and type-specific blocks: plan for metered_plan, pool for prepaid_credit, burn for burn-rate metrics, credit when a balance URL was polled, energy for kWh/calls), spend (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and alarm (plan-pace or stale-reading alert with kind/severity/headline). The old by_provider/total_balance_usd shape and flat top-level keys were removed.
  • coverage — routable-model counts (routable_models counts active models filtered on access_level IN (routing.allowed_access_levels), excluding any access-gated rows whose access_level is not in that list) with energy and proficiency data, plus a selection sub-key reporting which active models (unfiltered by access-level) the router has actually picked over the configured window, and a warnings array that now also includes entries for active models that can never enter the candidate set (structural-gate failures: zero context window or broken tier derived from the catalog, detected via routing.rejection_reason under the most permissive request — one token, tier 1, batch latency). Existing warnings cover missing energy data, missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and classifier source degradation. A cache sub-key carries the prefix-cache series from cache_rate_series(): window_hours, assumed_cache_rate, aggregate observations / prompt_tokens / cached_prompt_tokens / cache_rate, and a by_model list of the same figures per (provider, model_id). The rate is sum(cached_prompt_tokens) / sum(prompt_tokens) over rows the provider actually reported a count for (cached_tokens_source = 'reported'), excluding seed_reference sweeps. Two warnings read it: cache rate: when the aggregate diverges from objective.assumed_cache_rate by more than objective.cache_rate_warn_margin, and cache rate outlier: when one (provider, model_id) group does. Both need objective.cache_rate_warn_min_observations reported-cache rows first.
  • recent_decisions — last 50 rows from route_decisions
  • per_model — per-model aggregates over the last 30 days of energy_observations
  • verdict_mix — counts by verification verdict over the last 7 days
  • top_proficiency — top models by blended_score for coding_general
  • generated_at — ISO8601 timestamp
  • pinch — context-pruning aggregation (when pinch.enabled): calls_30d, pruned_calls_30d, and share_pruned (share of decisions pruned), plus median and total tokens_saved over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
  • cost_calibration — from cost_estimate_calibration(): how far routing.estimated_cost is from the provider's own bill, per (provider, model_id). Reported only — nothing applies it to estimated_cost or to ranking. route_decisions joined to energy_observations on request_id and on the selected model and provider (the id alone pairs a failover's estimate with another model's bill), over objective.cost_calibration_window_hours, excluding seed_reference on both sides and rows with no positive cost_usd or est_cost_usd. Keys: window_hours, min_observations, observations, est_cost_usd, billed_cost_usd, overestimate_ratio (est/billed), correction_factor (billed/est — what would be multiplied in), spread with spread_low / spread_high, and a by_model list of the same figures plus est_per_request_usd, billed_per_request_usd and sufficient. spread is the figure that matters: a uniform scale error reorders nothing, while the spread in that error does. Groups below objective.cost_calibration_min_observations are listed with sufficient: false and excluded from spread. No warning class reads it.
  • latency — from latency_series(): p50 and p95 of the router-observed router_wall_seconds and router_ttft_seconds per (provider, model_id) over objective.latency_window_hours, excluding seed_reference. A report, not an objective — nothing in the ranking reads it. These are the router's own clock, not the provider's duration_seconds, which is why OpenRouter models are visible here at all (OpenRouter reports no duration). Keys: window_hours, min_observations, aggregate wall_observations / ttft_observations / wall_p50 / wall_p95 / ttft_p50 / ttft_p95, and a by_model list of the same plus wall_sufficient / ttft_sufficient. The wall and TTFT counts are independent because TTFT is streaming-only; a group can be sufficient on one and not the other. Degrades to an empty series on a database that predates the two columns. No warning class reads it.

It exposes no conversation text, prompts, or session_dir; it is bound to loopback and unauthenticated exactly like /health.

Input to /route and /dispatch can include task_category, task_tier, and required_context_tokens overrides — these skip the classifier, useful for testing routing without the classifier in the loop.

Named routing profiles (auto:<profile>). The general form is auto:<profile>; auto alone is shorthand for auto:default. The router resolves the profile name against the built-in set and any custom entries in config/config.yaml under profiles:. Profiles narrow the candidate set but do not override the quality-first ranking objective.

Profile Effect
auto:default Normal quality-first routing, -flex rows excluded
auto:batch Admits -flex rows for overnight and async work
auto:locality Restricts to the ollama-local provider only
auto:onlycheaps Limits to models priced at ≤ $0.50 per 1M completion tokens
auto:bigboybritches Restricts to tier-3 (frontier) models only

From opencode: a profile is just another entry under provider.llm-router.models in opencode.json — the object's key is the literal model id opencode sends upstream, so auto:batch already works this way (see the repo-local opencode.json). To make auto:locality, auto:onlycheaps, or auto:bigboybritches selectable from opencode's model picker instead of typing an override, add one entry per profile:

"auto:onlycheaps": {
  "name": "auto (cost ceiling)",
  "limit": { "context": 782324, "output": 16384 },
  "modalities": { "input": ["text", "image"] }
}

limit/modalities can be copied from the auto entry unchanged — a profile narrows which models are candidates, not the context window or input types the router itself accepts. See clients.md for the full opencode wiring.

An unknown profile name raises HTTP 422 and lists the valid names.

The profile field is also accepted on POST /route and POST /dispatch to select a profile inline per request.

Ask for any real model id in /v1/chat/completions and it dispatches directly, still logged — routing is transparent, not opaque.

Streaming (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt emits energy and cost as SSE comment lines (: energy {...}) before data: [DONE] — ordinary clients ignore comments, so the stream flows untouched while the router scrapes telemetry on the way past. Without this, streamed calls would log no energy at all.

Verification headers (non-streaming): When streaming is not used, the structural verification verdict surfaces in the X-Router-Verification header so a client can inspect it without parsing the response body. Valid values: ok, truncated, malformed, unverifiable, none.

Capability 422s: When no model survives the hard filters, the 422 names the active constraints. That now includes "vision-capable model" or "json-mode-capable model" when the request carried images or a JSON-mode response_format, alongside the existing context/tier/latency/tool reasons.

POST /outcome: everything else the router records is a proxy — structural checks know whether code parses, the local LLM check guesses whether prose looks right, neither knows whether the answer did the job. The client does, because it ran the tests:

# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'

An unknown request_id returns 404 rather than being quietly accepted, so a client whose reports go nowhere finds out. The lookup checks energy_observations first (cloud), then local_energy_observations (local rows carry the same request_id/session_dir attribution) — cloud wins on a rare id collision.

Unlike the structural/local-LLM checks — which only ever record failures — /outcome folds both directions into proficiency via feedback.py: a false report counts against the model same as any other verification failure, but a true report counts too. It's also the only quality signal that survives streaming, since a retry can't reach a response whose bytes are already gone, while a report arrives afterward and works either way.

Pinning and auto behavior for local models

model: "auto" will route eligible tasks to qwen2.5-coder-router:14b when the classifier returns file_summarization or diff_checking and the local row survives the filters; otherwise a cloud model is selected as usual. Pin any local model_id (e.g. qwen2.5-coder-router:14b) and the request goes straight to that Ollama tag. /v1/models lists local rows with owned_by: "ollama-local".