Files
6krrt/docs/api.md
adlee-was-taken a26765c6c0 feat(metrics): report the cost-estimator's error and the latency nobody could see
Two measurement surfaces on /metrics. Neither is read by routing, and a test
asserts that against the module source: a live incumbent-cache-pricing
experiment is running, and a series that reached the ranker would confound it.

cost_calibration -- routing.estimated_cost against the provider's own bill,
per (provider, model), joined on request_id. The scale error is not the
finding: a uniform overestimate reorders nothing, because the ranking is a
comparison and every candidate moves together. The SPREAD does reorder, so
that is the headline figure. Live over 168h, 4,152 joined requests: est/billed
runs 1.62x (qwen/qwen3.6-35b-a3b, openrouter) to 12.86x (qwen3.6-35b-fast,
neuralwatt) -- a 7.9x spread, wider than the 5x the brief was written against.
The factors are REPORTED, not applied; whether they are stable enough to trust
is the question this exists to answer, and this project has already mistaken
one moment of a moving per-model quantity for a constant.

The join needs the model as well as the request id. On the live database 5 of
4,143 rows pair a decision that selected qwen3.6-35b (neuralwatt) with a
completion billed by qwen/qwen3.6-35b-a3b (openrouter) -- a cross-provider
failover, and one model's estimate against another's bill.

latency -- p50/p95 of router_wall_seconds and router_ttft_seconds, which
landed recently and nothing read. These are the router's own clock, not the
provider's duration_seconds, which is why OpenRouter is visible here at all:
it reports no duration. z-ai/glm-5.3-flash, a model with "flash" in its name,
measures p50 9.60s / p95 40.08s to first token against 1.58s / 3.09s for
deepseek-v4-flash on NeuralWatt. Reported, never scored. The wall and TTFT
sample counts are kept independent because TTFT is streaming-only by nature.

No new warning class, deliberately. Every group is out of band on both series
today, so a divergence warning would fire on all of them from the first run --
bare presence, which is the anti-pattern rejection_warnings exists to avoid --
and a warning is a form of trust these factors have not yet earned.

Four knobs under objective, all report-only, all recorded in
DELIBERATELY_NOT_IN_ADMIN with the reason: they change what /metrics shows,
and not even what it warns about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-16 14:34:27 -04:00

179 lines
11 KiB
Markdown

> API surface reference. Back to [README](../README.md).
## API Endpoints
The dispatcher binds `127.0.0.1:8080`. **No auth of its own** — loopback
is the only thing standing between the open internet and your billing
allowance.
| Method | Path | Description |
|---|---|---|
| `GET` | `/health` | Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set) |
| `GET` | `/metrics` | Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth |
| `GET` | `/events/decisions` | Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen |
| `POST` | `/route` | Classify task, rank candidates, return selected model — **no provider call, no cost** |
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
| `POST` | `/outcome` | Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed |
| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
`/metrics` returns a single JSON object with these top-level keys:
- `quota` — per-provider billing shape from `quota_accounts()`: `period` (billing window with start/next_reset/elapsed_fraction/source), `accounts[]` (list of per-provider dicts with `provider`, `shape`, `spend_usd`, and type-specific blocks: `plan` for metered_plan, `pool` for prepaid_credit, `burn` for burn-rate metrics, `credit` when a balance URL was polled, `energy` for kWh/calls), `spend` (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and `alarm` (plan-pace or stale-reading alert with kind/severity/headline). The old `by_provider`/`total_balance_usd` shape and flat top-level keys were removed.
- `coverage` — routable-model counts (`routable_models` counts active models filtered
on `access_level IN (routing.allowed_access_levels)`, excluding any access-gated rows
whose `access_level` is not in that list) with energy and proficiency data, plus a
`selection` sub-key reporting which active models (unfiltered by access-level) the
router has actually picked over the configured window, and a `warnings` array that
now also includes entries for active models that can never enter the candidate set
(structural-gate failures: zero context window or broken tier derived from the
catalog, detected via `routing.rejection_reason` under the most permissive request
— one token, tier 1, batch latency). Existing warnings cover missing energy data,
missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and
classifier source degradation.
A `cache` sub-key carries the prefix-cache series from `cache_rate_series()`:
`window_hours`, `assumed_cache_rate`, aggregate `observations` /
`prompt_tokens` / `cached_prompt_tokens` / `cache_rate`, and a `by_model` list
of the same figures per `(provider, model_id)`. The rate is
`sum(cached_prompt_tokens) / sum(prompt_tokens)` over rows the provider
actually reported a count for (`cached_tokens_source = 'reported'`), excluding
`seed_reference` sweeps. Two warnings read it: `cache rate:` when the
aggregate diverges from `objective.assumed_cache_rate` by more than
`objective.cache_rate_warn_margin`, and `cache rate outlier:` when one
`(provider, model_id)` group does. Both need
`objective.cache_rate_warn_min_observations` reported-cache rows first.
- `recent_decisions` — last 50 rows from `route_decisions`
- `per_model` — per-model aggregates over the last 30 days of `energy_observations`
- `verdict_mix` — counts by verification verdict over the last 7 days
- `top_proficiency` — top models by `blended_score` for `coding_general`
- `generated_at` — ISO8601 timestamp
- `pinch` — context-pruning aggregation (when `pinch.enabled`): `calls_30d`, `pruned_calls_30d`, and `share_pruned` (share of decisions pruned), plus median and total `tokens_saved` over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
- `cost_calibration` — from `cost_estimate_calibration()`: how far
`routing.estimated_cost` is from the provider's own bill, per
`(provider, model_id)`. **Reported only — nothing applies it to
`estimated_cost` or to ranking.** `route_decisions` joined to
`energy_observations` on `request_id` *and* on the selected model and
provider (the id alone pairs a failover's estimate with another model's
bill), over `objective.cost_calibration_window_hours`, excluding
`seed_reference` on both sides and rows with no positive `cost_usd` or
`est_cost_usd`. Keys: `window_hours`, `min_observations`, `observations`,
`est_cost_usd`, `billed_cost_usd`, `overestimate_ratio` (est/billed),
`correction_factor` (billed/est — what *would* be multiplied in), `spread`
with `spread_low` / `spread_high`, and a `by_model` list of the same figures
plus `est_per_request_usd`, `billed_per_request_usd` and `sufficient`.
`spread` is the figure that matters: a uniform scale error reorders nothing,
while the spread in that error does. Groups below
`objective.cost_calibration_min_observations` are listed with
`sufficient: false` and excluded from `spread`. No warning class reads it.
- `latency` — from `latency_series()`: p50 and p95 of the router-observed
`router_wall_seconds` and `router_ttft_seconds` per `(provider, model_id)`
over `objective.latency_window_hours`, excluding `seed_reference`. **A
report, not an objective — nothing in the ranking reads it.** These are the
router's own clock, not the provider's `duration_seconds`, which is why
OpenRouter models are visible here at all (OpenRouter reports no duration).
Keys: `window_hours`, `min_observations`, aggregate `wall_observations` /
`ttft_observations` / `wall_p50` / `wall_p95` / `ttft_p50` / `ttft_p95`, and
a `by_model` list of the same plus `wall_sufficient` / `ttft_sufficient`.
The wall and TTFT counts are independent because TTFT is streaming-only; a
group can be sufficient on one and not the other. Degrades to an empty
series on a database that predates the two columns. No warning class reads
it.
It exposes no conversation text, prompts, or `session_dir`; it is bound to
loopback and unauthenticated exactly like `/health`.
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`,
and `required_context_tokens` overrides — these skip the classifier, useful
for testing routing without the classifier in the loop.
Named routing profiles (`auto:<profile>`). The general form is
`auto:<profile>`; `auto` alone is shorthand for `auto:default`. The router
resolves the profile name against the built-in set and any custom entries in
`config/config.yaml` under `profiles:`. Profiles narrow the candidate set but
do not override the quality-first ranking objective.
| Profile | Effect |
|---|---|
| `auto:default` | Normal quality-first routing, `-flex` rows excluded |
| `auto:batch` | Admits `-flex` rows for overnight and async work |
| `auto:locality` | Restricts to the `ollama-local` provider only |
| `auto:onlycheaps` | Limits to models priced at ≤ $0.50 per 1M completion tokens |
| `auto:bigboybritches` | Restricts to tier-3 (frontier) models only |
**From opencode:** a profile is just another entry under
`provider.llm-router.models` in `opencode.json` — the object's key is the
literal model id opencode sends upstream, so `auto:batch` already works this
way (see the repo-local `opencode.json`). To make `auto:locality`,
`auto:onlycheaps`, or `auto:bigboybritches` selectable from opencode's model
picker instead of typing an override, add one entry per profile:
```json
"auto:onlycheaps": {
"name": "auto (cost ceiling)",
"limit": { "context": 782324, "output": 16384 },
"modalities": { "input": ["text", "image"] }
}
```
`limit`/`modalities` can be copied from the `auto` entry unchanged — a
profile narrows *which models* are candidates, not the context window or
input types the router itself accepts. See [clients.md](clients.md) for the
full opencode wiring.
An unknown profile name raises HTTP 422 and lists the valid names.
The `profile` field is also accepted on `POST /route` and `POST /dispatch`
to select a profile inline per request.
Ask for **any real model id** in `/v1/chat/completions` and it dispatches
directly, still logged — routing is transparent, not opaque.
**Streaming** (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt
emits energy and cost as SSE **comment** lines (`: energy {...}`) before
`data: [DONE]` — ordinary clients ignore comments, so the stream flows
untouched while the router scrapes telemetry on the way past. Without this,
streamed calls would log no energy at all.
**Verification headers** (non-streaming): When streaming is not used, the
structural verification verdict surfaces in the `X-Router-Verification` header
so a client can inspect it without parsing the response body. Valid values:
`ok`, `truncated`, `malformed`, `unverifiable`, `none`.
**Capability 422s**: When no model survives the hard filters, the 422 names the
active constraints. That now includes "vision-capable model" or
"json-mode-capable model" when the request carried images or a JSON-mode
`response_format`, alongside the existing context/tier/latency/tool reasons.
**`POST /outcome`**: everything else the router records is a proxy — structural
checks know whether code *parses*, the local LLM check guesses whether prose
*looks* right, neither knows whether the answer did the job. The client does,
because it ran the tests:
```bash
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
```
An unknown `request_id` returns `404` rather than being quietly accepted, so a
client whose reports go nowhere finds out. The lookup checks
`energy_observations` first (cloud), then `local_energy_observations` (local
rows carry the same `request_id`/`session_dir` attribution) — cloud wins on a
rare id collision.
Unlike the structural/local-LLM checks — which only ever record failures —
`/outcome` folds **both** directions into `proficiency` via `feedback.py`: a
`false` report counts against the model same as any other verification failure,
but a `true` report counts too. It's also the only quality signal that survives
streaming, since a retry can't reach a response whose bytes are already gone,
while a report arrives afterward and works either way.
## Pinning and auto behavior for local models
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
classifier returns `file_summarization` or `diff_checking` and the local row
survives the filters; otherwise a cloud model is selected as usual. Pin any
local `model_id` (e.g. `qwen2.5-coder-router:14b`) and the request goes straight
to that Ollama tag. `/v1/models` lists local rows with `owned_by: "ollama-local"`.