Two measurement surfaces on /metrics. Neither is read by routing, and a test asserts that against the module source: a live incumbent-cache-pricing experiment is running, and a series that reached the ranker would confound it. cost_calibration -- routing.estimated_cost against the provider's own bill, per (provider, model), joined on request_id. The scale error is not the finding: a uniform overestimate reorders nothing, because the ranking is a comparison and every candidate moves together. The SPREAD does reorder, so that is the headline figure. Live over 168h, 4,152 joined requests: est/billed runs 1.62x (qwen/qwen3.6-35b-a3b, openrouter) to 12.86x (qwen3.6-35b-fast, neuralwatt) -- a 7.9x spread, wider than the 5x the brief was written against. The factors are REPORTED, not applied; whether they are stable enough to trust is the question this exists to answer, and this project has already mistaken one moment of a moving per-model quantity for a constant. The join needs the model as well as the request id. On the live database 5 of 4,143 rows pair a decision that selected qwen3.6-35b (neuralwatt) with a completion billed by qwen/qwen3.6-35b-a3b (openrouter) -- a cross-provider failover, and one model's estimate against another's bill. latency -- p50/p95 of router_wall_seconds and router_ttft_seconds, which landed recently and nothing read. These are the router's own clock, not the provider's duration_seconds, which is why OpenRouter is visible here at all: it reports no duration. z-ai/glm-5.3-flash, a model with "flash" in its name, measures p50 9.60s / p95 40.08s to first token against 1.58s / 3.09s for deepseek-v4-flash on NeuralWatt. Reported, never scored. The wall and TTFT sample counts are kept independent because TTFT is streaming-only by nature. No new warning class, deliberately. Every group is out of band on both series today, so a divergence warning would fire on all of them from the first run -- bare presence, which is the anti-pattern rejection_warnings exists to avoid -- and a warning is a form of trust these factors have not yet earned. Four knobs under objective, all report-only, all recorded in DELIBERATELY_NOT_IN_ADMIN with the reason: they change what /metrics shows, and not even what it warns about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
179 lines
11 KiB
Markdown
179 lines
11 KiB
Markdown
> API surface reference. Back to [README](../README.md).
|
|
|
|
## API Endpoints
|
|
|
|
The dispatcher binds `127.0.0.1:8080`. **No auth of its own** — loopback
|
|
is the only thing standing between the open internet and your billing
|
|
allowance.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/health` | Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set) |
|
|
| `GET` | `/metrics` | Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth |
|
|
| `GET` | `/events/decisions` | Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen |
|
|
| `POST` | `/route` | Classify task, rank candidates, return selected model — **no provider call, no cost** |
|
|
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
|
|
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
|
|
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
|
|
| `POST` | `/outcome` | Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed |
|
|
| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
|
|
|
|
`/metrics` returns a single JSON object with these top-level keys:
|
|
|
|
- `quota` — per-provider billing shape from `quota_accounts()`: `period` (billing window with start/next_reset/elapsed_fraction/source), `accounts[]` (list of per-provider dicts with `provider`, `shape`, `spend_usd`, and type-specific blocks: `plan` for metered_plan, `pool` for prepaid_credit, `burn` for burn-rate metrics, `credit` when a balance URL was polled, `energy` for kWh/calls), `spend` (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and `alarm` (plan-pace or stale-reading alert with kind/severity/headline). The old `by_provider`/`total_balance_usd` shape and flat top-level keys were removed.
|
|
- `coverage` — routable-model counts (`routable_models` counts active models filtered
|
|
on `access_level IN (routing.allowed_access_levels)`, excluding any access-gated rows
|
|
whose `access_level` is not in that list) with energy and proficiency data, plus a
|
|
`selection` sub-key reporting which active models (unfiltered by access-level) the
|
|
router has actually picked over the configured window, and a `warnings` array that
|
|
now also includes entries for active models that can never enter the candidate set
|
|
(structural-gate failures: zero context window or broken tier derived from the
|
|
catalog, detected via `routing.rejection_reason` under the most permissive request
|
|
— one token, tier 1, batch latency). Existing warnings cover missing energy data,
|
|
missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and
|
|
classifier source degradation.
|
|
A `cache` sub-key carries the prefix-cache series from `cache_rate_series()`:
|
|
`window_hours`, `assumed_cache_rate`, aggregate `observations` /
|
|
`prompt_tokens` / `cached_prompt_tokens` / `cache_rate`, and a `by_model` list
|
|
of the same figures per `(provider, model_id)`. The rate is
|
|
`sum(cached_prompt_tokens) / sum(prompt_tokens)` over rows the provider
|
|
actually reported a count for (`cached_tokens_source = 'reported'`), excluding
|
|
`seed_reference` sweeps. Two warnings read it: `cache rate:` when the
|
|
aggregate diverges from `objective.assumed_cache_rate` by more than
|
|
`objective.cache_rate_warn_margin`, and `cache rate outlier:` when one
|
|
`(provider, model_id)` group does. Both need
|
|
`objective.cache_rate_warn_min_observations` reported-cache rows first.
|
|
- `recent_decisions` — last 50 rows from `route_decisions`
|
|
- `per_model` — per-model aggregates over the last 30 days of `energy_observations`
|
|
- `verdict_mix` — counts by verification verdict over the last 7 days
|
|
- `top_proficiency` — top models by `blended_score` for `coding_general`
|
|
- `generated_at` — ISO8601 timestamp
|
|
- `pinch` — context-pruning aggregation (when `pinch.enabled`): `calls_30d`, `pruned_calls_30d`, and `share_pruned` (share of decisions pruned), plus median and total `tokens_saved` over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
|
|
- `cost_calibration` — from `cost_estimate_calibration()`: how far
|
|
`routing.estimated_cost` is from the provider's own bill, per
|
|
`(provider, model_id)`. **Reported only — nothing applies it to
|
|
`estimated_cost` or to ranking.** `route_decisions` joined to
|
|
`energy_observations` on `request_id` *and* on the selected model and
|
|
provider (the id alone pairs a failover's estimate with another model's
|
|
bill), over `objective.cost_calibration_window_hours`, excluding
|
|
`seed_reference` on both sides and rows with no positive `cost_usd` or
|
|
`est_cost_usd`. Keys: `window_hours`, `min_observations`, `observations`,
|
|
`est_cost_usd`, `billed_cost_usd`, `overestimate_ratio` (est/billed),
|
|
`correction_factor` (billed/est — what *would* be multiplied in), `spread`
|
|
with `spread_low` / `spread_high`, and a `by_model` list of the same figures
|
|
plus `est_per_request_usd`, `billed_per_request_usd` and `sufficient`.
|
|
`spread` is the figure that matters: a uniform scale error reorders nothing,
|
|
while the spread in that error does. Groups below
|
|
`objective.cost_calibration_min_observations` are listed with
|
|
`sufficient: false` and excluded from `spread`. No warning class reads it.
|
|
- `latency` — from `latency_series()`: p50 and p95 of the router-observed
|
|
`router_wall_seconds` and `router_ttft_seconds` per `(provider, model_id)`
|
|
over `objective.latency_window_hours`, excluding `seed_reference`. **A
|
|
report, not an objective — nothing in the ranking reads it.** These are the
|
|
router's own clock, not the provider's `duration_seconds`, which is why
|
|
OpenRouter models are visible here at all (OpenRouter reports no duration).
|
|
Keys: `window_hours`, `min_observations`, aggregate `wall_observations` /
|
|
`ttft_observations` / `wall_p50` / `wall_p95` / `ttft_p50` / `ttft_p95`, and
|
|
a `by_model` list of the same plus `wall_sufficient` / `ttft_sufficient`.
|
|
The wall and TTFT counts are independent because TTFT is streaming-only; a
|
|
group can be sufficient on one and not the other. Degrades to an empty
|
|
series on a database that predates the two columns. No warning class reads
|
|
it.
|
|
|
|
It exposes no conversation text, prompts, or `session_dir`; it is bound to
|
|
loopback and unauthenticated exactly like `/health`.
|
|
|
|
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`,
|
|
and `required_context_tokens` overrides — these skip the classifier, useful
|
|
for testing routing without the classifier in the loop.
|
|
|
|
Named routing profiles (`auto:<profile>`). The general form is
|
|
`auto:<profile>`; `auto` alone is shorthand for `auto:default`. The router
|
|
resolves the profile name against the built-in set and any custom entries in
|
|
`config/config.yaml` under `profiles:`. Profiles narrow the candidate set but
|
|
do not override the quality-first ranking objective.
|
|
|
|
| Profile | Effect |
|
|
|---|---|
|
|
| `auto:default` | Normal quality-first routing, `-flex` rows excluded |
|
|
| `auto:batch` | Admits `-flex` rows for overnight and async work |
|
|
| `auto:locality` | Restricts to the `ollama-local` provider only |
|
|
| `auto:onlycheaps` | Limits to models priced at ≤ $0.50 per 1M completion tokens |
|
|
| `auto:bigboybritches` | Restricts to tier-3 (frontier) models only |
|
|
|
|
**From opencode:** a profile is just another entry under
|
|
`provider.llm-router.models` in `opencode.json` — the object's key is the
|
|
literal model id opencode sends upstream, so `auto:batch` already works this
|
|
way (see the repo-local `opencode.json`). To make `auto:locality`,
|
|
`auto:onlycheaps`, or `auto:bigboybritches` selectable from opencode's model
|
|
picker instead of typing an override, add one entry per profile:
|
|
|
|
```json
|
|
"auto:onlycheaps": {
|
|
"name": "auto (cost ceiling)",
|
|
"limit": { "context": 782324, "output": 16384 },
|
|
"modalities": { "input": ["text", "image"] }
|
|
}
|
|
```
|
|
|
|
`limit`/`modalities` can be copied from the `auto` entry unchanged — a
|
|
profile narrows *which models* are candidates, not the context window or
|
|
input types the router itself accepts. See [clients.md](clients.md) for the
|
|
full opencode wiring.
|
|
|
|
An unknown profile name raises HTTP 422 and lists the valid names.
|
|
|
|
The `profile` field is also accepted on `POST /route` and `POST /dispatch`
|
|
to select a profile inline per request.
|
|
|
|
Ask for **any real model id** in `/v1/chat/completions` and it dispatches
|
|
directly, still logged — routing is transparent, not opaque.
|
|
|
|
**Streaming** (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt
|
|
emits energy and cost as SSE **comment** lines (`: energy {...}`) before
|
|
`data: [DONE]` — ordinary clients ignore comments, so the stream flows
|
|
untouched while the router scrapes telemetry on the way past. Without this,
|
|
streamed calls would log no energy at all.
|
|
|
|
**Verification headers** (non-streaming): When streaming is not used, the
|
|
structural verification verdict surfaces in the `X-Router-Verification` header
|
|
so a client can inspect it without parsing the response body. Valid values:
|
|
`ok`, `truncated`, `malformed`, `unverifiable`, `none`.
|
|
|
|
**Capability 422s**: When no model survives the hard filters, the 422 names the
|
|
active constraints. That now includes "vision-capable model" or
|
|
"json-mode-capable model" when the request carried images or a JSON-mode
|
|
`response_format`, alongside the existing context/tier/latency/tool reasons.
|
|
|
|
**`POST /outcome`**: everything else the router records is a proxy — structural
|
|
checks know whether code *parses*, the local LLM check guesses whether prose
|
|
*looks* right, neither knows whether the answer did the job. The client does,
|
|
because it ran the tests:
|
|
|
|
```bash
|
|
# id comes from the completion body, or any stream chunk
|
|
curl -s localhost:8080/outcome -H 'content-type: application/json' \
|
|
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
|
|
```
|
|
|
|
An unknown `request_id` returns `404` rather than being quietly accepted, so a
|
|
client whose reports go nowhere finds out. The lookup checks
|
|
`energy_observations` first (cloud), then `local_energy_observations` (local
|
|
rows carry the same `request_id`/`session_dir` attribution) — cloud wins on a
|
|
rare id collision.
|
|
|
|
Unlike the structural/local-LLM checks — which only ever record failures —
|
|
`/outcome` folds **both** directions into `proficiency` via `feedback.py`: a
|
|
`false` report counts against the model same as any other verification failure,
|
|
but a `true` report counts too. It's also the only quality signal that survives
|
|
streaming, since a retry can't reach a response whose bytes are already gone,
|
|
while a report arrives afterward and works either way.
|
|
|
|
## Pinning and auto behavior for local models
|
|
|
|
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
|
|
classifier returns `file_summarization` or `diff_checking` and the local row
|
|
survives the filters; otherwise a cloud model is selected as usual. Pin any
|
|
local `model_id` (e.g. `qwen2.5-coder-router:14b`) and the request goes straight
|
|
to that Ollama tag. `/v1/models` lists local rows with `owned_by: "ollama-local"`.
|