Files
6krrt/docs/api.md
2026-09-26 16:50:24 -04:00

300 lines
17 KiB
Markdown

> API surface reference. Back to [README](../README.md).
## API Endpoints
The dispatcher binds `127.0.0.1:8080`. **No auth of its own** — loopback
is the only thing standing between the open internet and your billing
allowance.
| Method | Path | Description |
|---|---|---|
| `GET` | `/health` | Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set) |
| `GET` | `/metrics` | Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth |
| `GET` | `/events/decisions` | Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen |
| `POST` | `/route` | Classify task, rank candidates, return selected model — **no provider call, no cost** |
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
| `POST` | `/outcome` | Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed |
| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
`/metrics` returns a single JSON object with these top-level keys:
- `quota` — per-provider billing shape from `quota_accounts()`: `period` (billing window with start/next_reset/elapsed_fraction/source), `accounts[]` (list of per-provider dicts with `provider`, `shape`, `spend_usd`, and type-specific blocks: `plan` for metered_plan, `pool` for prepaid_credit, `burn` for burn-rate metrics, `credit` when a balance URL was polled, `energy` for kWh/calls), `spend` (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and `alarm` (plan-pace or stale-reading alert with kind/severity/headline). The old `by_provider`/`total_balance_usd` shape and flat top-level keys were removed.
- `coverage` — routable-model counts (`routable_models` counts active models filtered
on `access_level IN (routing.allowed_access_levels)`, excluding any access-gated rows
whose `access_level` is not in that list) with energy and proficiency data, plus a
`selection` sub-key reporting which active models (unfiltered by access-level) the
router has actually picked over the configured window, and a `warnings` array that
now also includes entries for active models that can never enter the candidate set
(structural-gate failures: zero context window or broken tier derived from the
catalog, detected via `routing.rejection_reason` under the most permissive request
— one token, tier 1, batch latency). Existing warnings cover missing energy data,
missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and
classifier source degradation.
A `cache` sub-key carries the prefix-cache series from `cache_rate_series()`:
`window_hours`, `assumed_cache_rate`, aggregate `observations` /
`prompt_tokens` / `cached_prompt_tokens` / `cache_rate`, and a `by_model` list
of the same figures per `(provider, model_id)`. The rate is
`sum(cached_prompt_tokens) / sum(prompt_tokens)` over rows the provider
actually reported a count for (`cached_tokens_source = 'reported'`), excluding
`seed_reference` sweeps. Two warnings read it: `cache rate:` when the
aggregate diverges from `objective.assumed_cache_rate` by more than
`objective.cache_rate_warn_margin`, and `cache rate outlier:` when one
`(provider, model_id)` group does. Both need
`objective.cache_rate_warn_min_observations` reported-cache rows first.
- `recent_decisions` — last 50 rows from `route_decisions`
- `per_model` — per-model aggregates over the last 30 days of `energy_observations`
- `verdict_mix` — counts by verification verdict over the last 7 days
- `top_proficiency` — top models by `blended_score` for `coding_general`
- `generated_at` — ISO8601 timestamp
- `pinch` — context-pruning aggregation (when `pinch.enabled`): `calls_30d`, `pruned_calls_30d`, and `share_pruned` (share of decisions pruned), plus median and total `tokens_saved` over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
- `cost_calibration` — from `cost_estimate_calibration()`: how far
`routing.estimated_cost` is from the provider's own bill, per
`(provider, model_id)`. **Reported only — nothing applies it to
`estimated_cost` or to ranking.** `route_decisions` joined to
`energy_observations` on `request_id` *and* on the selected model and
provider (the id alone pairs a failover's estimate with another model's
bill), over `objective.cost_calibration_window_hours`, excluding
`seed_reference` on both sides and rows with no positive `cost_usd` or
`est_cost_usd`. Keys: `window_hours`, `min_observations`, `observations`,
`est_cost_usd`, `billed_cost_usd`, `overestimate_ratio` (est/billed),
`correction_factor` (billed/est — what *would* be multiplied in), `spread`
with `spread_low` / `spread_high`, and a `by_model` list of the same figures
plus `est_per_request_usd`, `billed_per_request_usd` and `sufficient`.
`spread` is the figure that matters: a uniform scale error reorders nothing,
while the spread in that error does. Groups below
`objective.cost_calibration_min_observations` are listed with
`sufficient: false` and excluded from `spread`. No warning class reads it.
- `latency` — from `latency_series()`: p50 and p95 of the router-observed
`router_wall_seconds` and `router_ttft_seconds` per `(provider, model_id)`
over `objective.latency_window_hours`, excluding `seed_reference`. **A
report, not an objective — nothing in the ranking reads it.** These are the
router's own clock, not the provider's `duration_seconds`, which is why
OpenRouter models are visible here at all (OpenRouter reports no duration).
Keys: `window_hours`, `min_observations`, aggregate `wall_observations` /
`ttft_observations` / `wall_p50` / `wall_p95` / `ttft_p50` / `ttft_p95`, and
a `by_model` list of the same plus `wall_sufficient` / `ttft_sufficient`.
The wall and TTFT counts are independent because TTFT is streaming-only; a
group can be sufficient on one and not the other. Degrades to an empty
series on a database that predates the two columns. No warning class reads
it.
It exposes no conversation text, prompts, or `session_dir`; it is bound to
loopback and unauthenticated exactly like `/health`.
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`,
and `required_context_tokens` overrides — these skip the classifier, useful
for testing routing without the classifier in the loop.
Named routing profiles (`auto:<profile>`). The general form is
`auto:<profile>`; `auto` alone is shorthand for `auto:default`. The router
resolves the profile name against the built-in set and any custom entries in
`config/config.yaml` under `profiles:`. Profiles narrow the candidate set but
do not override the quality-first ranking objective.
| Profile | Effect |
|---|---|
| `auto:default` | Normal quality-first routing, `-flex` rows excluded |
| `auto:batch` | Admits `-flex` rows for overnight and async work |
| `auto:locality` | Restricts to the `ollama-local` provider only |
| `auto:onlycheaps` | Limits to models priced at ≤ $0.50 per 1M completion tokens |
| `auto:bigboybritches` | Restricts to tier-3 (frontier) models only |
**From opencode:** a profile is just another entry under
`provider.llm-router.models` in `opencode.json` — the object's key is the
literal model id opencode sends upstream, so `auto:batch` already works this
way (see the repo-local `opencode.json`). To make `auto:locality`,
`auto:onlycheaps`, or `auto:bigboybritches` selectable from opencode's model
picker instead of typing an override, add one entry per profile:
```json
"auto:onlycheaps": {
"name": "auto (cost ceiling)",
"limit": { "context": 782324, "output": 16384 },
"modalities": { "input": ["text", "image"] }
}
```
`limit`/`modalities` can be copied from the `auto` entry unchanged — a
profile narrows *which models* are candidates, not the context window or
input types the router itself accepts. See [clients.md](clients.md) for the
full opencode wiring.
An unknown profile name raises HTTP 422 and lists the valid names.
The `profile` field is also accepted on `POST /route` and `POST /dispatch`
to select a profile inline per request.
Ask for **any real model id** in `/v1/chat/completions` and it dispatches
directly, still logged — routing is transparent, not opaque.
**Streaming** (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt
emits energy and cost as SSE **comment** lines (`: energy {...}`) before
`data: [DONE]` — ordinary clients ignore comments, so the stream flows
untouched while the router scrapes telemetry on the way past. Without this,
streamed calls would log no energy at all.
**Verification headers** (non-streaming): When streaming is not used, the
structural verification verdict surfaces in the `X-Router-Verification` header
so a client can inspect it without parsing the response body. Valid values:
`ok`, `truncated`, `malformed`, `unverifiable`, `none`.
**Conversation identity headers** (optional): A client may attach three
headers to `POST /v1/chat/completions` so the router can group its decisions
by conversation and attribute an answer to the exact conversation that asked
for it. All three are optional; an absent or invalid header is treated as
absent (it is never an error).
| Header | Meaning | Length | Pattern |
|---|---|---|---|
| `X-Router-Conversation` | Id of ONE conversation (e.g. an opencode sessionID) | 1-128 chars | `^[A-Za-z0-9._:-]+$` |
| `X-Router-Agent` | Name of the agent making the request | 1-64 chars | `^[A-Za-z0-9._:-]+$` |
| `X-Router-Parent` | Parent conversation id, when this one is a sub-conversation | 1-128 chars | `^[A-Za-z0-9._:-]+$` |
Contract rules:
- Charset: letters, digits, `.`, `_`, `:`, `-` (the regex `^[A-Za-z0-9._:-]+$`);
anything else, or a length outside the table's range, makes the header
invalid and therefore absent.
- An absent or invalid header is treated as absent, never as an error: the
request is still routed normally.
- Namespacing: a valid `X-Router-Conversation` produces
`session_key = "c:" + conversation`, so client conversations are keyed
separately from the router's own hashed prompt fingerprints (which carry no
`c:` prefix). A valid `X-Router-Parent` produces `parent_key = "c:" + parent`.
Conversation ids never appear alongside message text: `session_key` stores
only the namespaced id, never the messages.
- When `X-Router-Conversation` is absent or invalid, the router falls back to
its hashed prompt fingerprint as `session_key` (see [routing.md](routing.md)
"incumbency and cache pricing").
These ids are what make `/outcome` attribution exact (below): an
`X-Router-Conversation` on the request sets the `session_key` row that a later
report can address by `conversation_id` without guessing `source` or ambiguity.
**Plugin loader contract:** the opencode plugin that stamps these headers
must export only functions — the opencode 1.18.x plugin loader rejects a
plugin outright if any export is not a function (so the plugin's
`parentCache` lives as a property on the factory rather than a named export).
See [clients.md](clients.md) for the full contract.
**Capability 422s**: When no model survives the hard filters, the 422 names the
active constraints. That now includes "vision-capable model" or
"json-mode-capable model" when the request carried images or a JSON-mode
`response_format`, alongside the existing context/tier/latency/tool reasons.
**`POST /outcome`**: everything else the router records is a proxy — structural
checks know whether code *parses*, the local LLM check guesses whether prose
*looks* right, neither knows whether the answer did the job. The client does,
because it ran the tests:
```bash
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
```
An unknown `request_id` returns `404` rather than being quietly accepted, so a
client whose reports go nowhere finds out. The lookup checks
`energy_observations` first (cloud), then `local_energy_observations` (local
rows carry the same `request_id`/`session_dir` attribution) — cloud wins on a
rare id collision.
`conversation_id` is an optional extra body field that makes attribution
exact: it names the conversation (the same `X-Router-Conversation` the client
sent on the request) so a report can land without a `request_id`, and without
guessing which source the row came from.
```bash
# attribute a report to a conversation, no request_id needed
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"conversation_id":"abc-123","ok":false,"detail":"tests failed"}'
```
Resolution order is strict:
1. `request_id` — exact id; unmatched returns `404`.
2. `conversation_id` (when present) — exact namespaced match on the
conversation's rows; unmatched returns `404` and never falls through.
3. `source` / unambiguous fallback — only when `conversation_id` is entirely
absent (no `source` given and exactly one plausible row).
When `conversation_id` is present but invalid (wrong charset or length) the
report is rejected with `422`, because a malformed id is a client bug, not an
absence.
Unlike the structural/local-LLM checks — which only ever record failures —
`/outcome` folds **both** directions into `proficiency` via `feedback.py`: a
`false` report counts against the model same as any other verification failure,
but a `true` report counts too. It's also the only quality signal that survives
streaming, since a retry can't reach a response whose bytes are already gone,
while a report arrives afterward and works either way.
## Admin watchdog and availability API
The admin portal's watchdog and model-availability operations sit under
`/admin` (loopback-only, like the rest of the admin API). The watchdog
endpoints back the dashboard's **Loops** card and the **Watchdog** control
card; full behaviour is described in [watchdog.md](watchdog.md) and
[admin-portal.md](admin-portal.md).
| Method | Path | Description |
|---|---|---|
| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count |
| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) |
| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings |
| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity |
| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels |
`GET /admin/api/watchdog/status` returns a single object:
- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the
first tick
- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or
`null` when there is no tick yet
- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null
`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows
where `resolved_at` is null, ordered by `opened_at` descending. Each element
is the full alert row.
`GET /admin/api/watchdog/channels` returns an array of
`watchdog_channel_settings` rows ordered by `channel_name`.
`POST /admin/api/watchdog/channels` upserts one channel. Body:
- `channel_name` (required) — the channel to create or update; absence is a
422
- `enabled` (optional boolean) — when omitted the existing value is kept
- `min_severity` (optional: `info`, `warning`, or `critical`) — any other
value is a 422
Returns `{ok: true}`.
`POST /admin/api/watchdog/test-alert` delivers a test alert through the
currently enabled channels. Body:
- `severity` (optional, default `warning`; `info`, `warning`, or `critical`)
- `channel_name` (optional) — restrict to a single channel
Returns `{sent: <n>, severity: <severity>}`, or `{sent: 0, message: "no
enabled channels"}` when nothing is enabled. An invalid `severity` is a 422.
**Model availability** — `POST
/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin
override for a model's availability. `availability` must be one of `active`,
`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an
operator stop that routing and pinned requests exclude
(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`.
`DELETE` on the same path removes the override. The `{model_id:path}`
converter accepts slash-bearing model ids (e.g. OpenRouter).
## Pinning and auto behavior for local models
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
classifier returns `file_summarization` or `diff_checking` and the local row
survives the filters; otherwise a cloud model is selected as usual. Pin any
local `model_id` (e.g. `qwen2.5-coder-router:14b`) and the request goes straight
to that Ollama tag. `/v1/models` lists local rows with `owned_by: "ollama-local"`.