300 lines
17 KiB
Markdown
300 lines
17 KiB
Markdown
> API surface reference. Back to [README](../README.md).
|
|
|
|
## API Endpoints
|
|
|
|
The dispatcher binds `127.0.0.1:8080`. **No auth of its own** — loopback
|
|
is the only thing standing between the open internet and your billing
|
|
allowance.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/health` | Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set) |
|
|
| `GET` | `/metrics` | Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth |
|
|
| `GET` | `/events/decisions` | Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen |
|
|
| `POST` | `/route` | Classify task, rank candidates, return selected model — **no provider call, no cost** |
|
|
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
|
|
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
|
|
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
|
|
| `POST` | `/outcome` | Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed |
|
|
| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
|
|
|
|
`/metrics` returns a single JSON object with these top-level keys:
|
|
|
|
- `quota` — per-provider billing shape from `quota_accounts()`: `period` (billing window with start/next_reset/elapsed_fraction/source), `accounts[]` (list of per-provider dicts with `provider`, `shape`, `spend_usd`, and type-specific blocks: `plan` for metered_plan, `pool` for prepaid_credit, `burn` for burn-rate metrics, `credit` when a balance URL was polled, `energy` for kWh/calls), `spend` (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and `alarm` (plan-pace or stale-reading alert with kind/severity/headline). The old `by_provider`/`total_balance_usd` shape and flat top-level keys were removed.
|
|
- `coverage` — routable-model counts (`routable_models` counts active models filtered
|
|
on `access_level IN (routing.allowed_access_levels)`, excluding any access-gated rows
|
|
whose `access_level` is not in that list) with energy and proficiency data, plus a
|
|
`selection` sub-key reporting which active models (unfiltered by access-level) the
|
|
router has actually picked over the configured window, and a `warnings` array that
|
|
now also includes entries for active models that can never enter the candidate set
|
|
(structural-gate failures: zero context window or broken tier derived from the
|
|
catalog, detected via `routing.rejection_reason` under the most permissive request
|
|
— one token, tier 1, batch latency). Existing warnings cover missing energy data,
|
|
missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and
|
|
classifier source degradation.
|
|
A `cache` sub-key carries the prefix-cache series from `cache_rate_series()`:
|
|
`window_hours`, `assumed_cache_rate`, aggregate `observations` /
|
|
`prompt_tokens` / `cached_prompt_tokens` / `cache_rate`, and a `by_model` list
|
|
of the same figures per `(provider, model_id)`. The rate is
|
|
`sum(cached_prompt_tokens) / sum(prompt_tokens)` over rows the provider
|
|
actually reported a count for (`cached_tokens_source = 'reported'`), excluding
|
|
`seed_reference` sweeps. Two warnings read it: `cache rate:` when the
|
|
aggregate diverges from `objective.assumed_cache_rate` by more than
|
|
`objective.cache_rate_warn_margin`, and `cache rate outlier:` when one
|
|
`(provider, model_id)` group does. Both need
|
|
`objective.cache_rate_warn_min_observations` reported-cache rows first.
|
|
- `recent_decisions` — last 50 rows from `route_decisions`
|
|
- `per_model` — per-model aggregates over the last 30 days of `energy_observations`
|
|
- `verdict_mix` — counts by verification verdict over the last 7 days
|
|
- `top_proficiency` — top models by `blended_score` for `coding_general`
|
|
- `generated_at` — ISO8601 timestamp
|
|
- `pinch` — context-pruning aggregation (when `pinch.enabled`): `calls_30d`, `pruned_calls_30d`, and `share_pruned` (share of decisions pruned), plus median and total `tokens_saved` over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
|
|
- `cost_calibration` — from `cost_estimate_calibration()`: how far
|
|
`routing.estimated_cost` is from the provider's own bill, per
|
|
`(provider, model_id)`. **Reported only — nothing applies it to
|
|
`estimated_cost` or to ranking.** `route_decisions` joined to
|
|
`energy_observations` on `request_id` *and* on the selected model and
|
|
provider (the id alone pairs a failover's estimate with another model's
|
|
bill), over `objective.cost_calibration_window_hours`, excluding
|
|
`seed_reference` on both sides and rows with no positive `cost_usd` or
|
|
`est_cost_usd`. Keys: `window_hours`, `min_observations`, `observations`,
|
|
`est_cost_usd`, `billed_cost_usd`, `overestimate_ratio` (est/billed),
|
|
`correction_factor` (billed/est — what *would* be multiplied in), `spread`
|
|
with `spread_low` / `spread_high`, and a `by_model` list of the same figures
|
|
plus `est_per_request_usd`, `billed_per_request_usd` and `sufficient`.
|
|
`spread` is the figure that matters: a uniform scale error reorders nothing,
|
|
while the spread in that error does. Groups below
|
|
`objective.cost_calibration_min_observations` are listed with
|
|
`sufficient: false` and excluded from `spread`. No warning class reads it.
|
|
- `latency` — from `latency_series()`: p50 and p95 of the router-observed
|
|
`router_wall_seconds` and `router_ttft_seconds` per `(provider, model_id)`
|
|
over `objective.latency_window_hours`, excluding `seed_reference`. **A
|
|
report, not an objective — nothing in the ranking reads it.** These are the
|
|
router's own clock, not the provider's `duration_seconds`, which is why
|
|
OpenRouter models are visible here at all (OpenRouter reports no duration).
|
|
Keys: `window_hours`, `min_observations`, aggregate `wall_observations` /
|
|
`ttft_observations` / `wall_p50` / `wall_p95` / `ttft_p50` / `ttft_p95`, and
|
|
a `by_model` list of the same plus `wall_sufficient` / `ttft_sufficient`.
|
|
The wall and TTFT counts are independent because TTFT is streaming-only; a
|
|
group can be sufficient on one and not the other. Degrades to an empty
|
|
series on a database that predates the two columns. No warning class reads
|
|
it.
|
|
|
|
It exposes no conversation text, prompts, or `session_dir`; it is bound to
|
|
loopback and unauthenticated exactly like `/health`.
|
|
|
|
Input to `/route` and `/dispatch` can include `task_category`, `task_tier`,
|
|
and `required_context_tokens` overrides — these skip the classifier, useful
|
|
for testing routing without the classifier in the loop.
|
|
|
|
Named routing profiles (`auto:<profile>`). The general form is
|
|
`auto:<profile>`; `auto` alone is shorthand for `auto:default`. The router
|
|
resolves the profile name against the built-in set and any custom entries in
|
|
`config/config.yaml` under `profiles:`. Profiles narrow the candidate set but
|
|
do not override the quality-first ranking objective.
|
|
|
|
| Profile | Effect |
|
|
|---|---|
|
|
| `auto:default` | Normal quality-first routing, `-flex` rows excluded |
|
|
| `auto:batch` | Admits `-flex` rows for overnight and async work |
|
|
| `auto:locality` | Restricts to the `ollama-local` provider only |
|
|
| `auto:onlycheaps` | Limits to models priced at ≤ $0.50 per 1M completion tokens |
|
|
| `auto:bigboybritches` | Restricts to tier-3 (frontier) models only |
|
|
|
|
**From opencode:** a profile is just another entry under
|
|
`provider.llm-router.models` in `opencode.json` — the object's key is the
|
|
literal model id opencode sends upstream, so `auto:batch` already works this
|
|
way (see the repo-local `opencode.json`). To make `auto:locality`,
|
|
`auto:onlycheaps`, or `auto:bigboybritches` selectable from opencode's model
|
|
picker instead of typing an override, add one entry per profile:
|
|
|
|
```json
|
|
"auto:onlycheaps": {
|
|
"name": "auto (cost ceiling)",
|
|
"limit": { "context": 782324, "output": 16384 },
|
|
"modalities": { "input": ["text", "image"] }
|
|
}
|
|
```
|
|
|
|
`limit`/`modalities` can be copied from the `auto` entry unchanged — a
|
|
profile narrows *which models* are candidates, not the context window or
|
|
input types the router itself accepts. See [clients.md](clients.md) for the
|
|
full opencode wiring.
|
|
|
|
An unknown profile name raises HTTP 422 and lists the valid names.
|
|
|
|
The `profile` field is also accepted on `POST /route` and `POST /dispatch`
|
|
to select a profile inline per request.
|
|
|
|
Ask for **any real model id** in `/v1/chat/completions` and it dispatches
|
|
directly, still logged — routing is transparent, not opaque.
|
|
|
|
**Streaming** (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt
|
|
emits energy and cost as SSE **comment** lines (`: energy {...}`) before
|
|
`data: [DONE]` — ordinary clients ignore comments, so the stream flows
|
|
untouched while the router scrapes telemetry on the way past. Without this,
|
|
streamed calls would log no energy at all.
|
|
|
|
**Verification headers** (non-streaming): When streaming is not used, the
|
|
structural verification verdict surfaces in the `X-Router-Verification` header
|
|
so a client can inspect it without parsing the response body. Valid values:
|
|
`ok`, `truncated`, `malformed`, `unverifiable`, `none`.
|
|
|
|
**Conversation identity headers** (optional): A client may attach three
|
|
headers to `POST /v1/chat/completions` so the router can group its decisions
|
|
by conversation and attribute an answer to the exact conversation that asked
|
|
for it. All three are optional; an absent or invalid header is treated as
|
|
absent (it is never an error).
|
|
|
|
| Header | Meaning | Length | Pattern |
|
|
|---|---|---|---|
|
|
| `X-Router-Conversation` | Id of ONE conversation (e.g. an opencode sessionID) | 1-128 chars | `^[A-Za-z0-9._:-]+$` |
|
|
| `X-Router-Agent` | Name of the agent making the request | 1-64 chars | `^[A-Za-z0-9._:-]+$` |
|
|
| `X-Router-Parent` | Parent conversation id, when this one is a sub-conversation | 1-128 chars | `^[A-Za-z0-9._:-]+$` |
|
|
|
|
Contract rules:
|
|
|
|
- Charset: letters, digits, `.`, `_`, `:`, `-` (the regex `^[A-Za-z0-9._:-]+$`);
|
|
anything else, or a length outside the table's range, makes the header
|
|
invalid and therefore absent.
|
|
- An absent or invalid header is treated as absent, never as an error: the
|
|
request is still routed normally.
|
|
- Namespacing: a valid `X-Router-Conversation` produces
|
|
`session_key = "c:" + conversation`, so client conversations are keyed
|
|
separately from the router's own hashed prompt fingerprints (which carry no
|
|
`c:` prefix). A valid `X-Router-Parent` produces `parent_key = "c:" + parent`.
|
|
Conversation ids never appear alongside message text: `session_key` stores
|
|
only the namespaced id, never the messages.
|
|
- When `X-Router-Conversation` is absent or invalid, the router falls back to
|
|
its hashed prompt fingerprint as `session_key` (see [routing.md](routing.md)
|
|
"incumbency and cache pricing").
|
|
|
|
These ids are what make `/outcome` attribution exact (below): an
|
|
`X-Router-Conversation` on the request sets the `session_key` row that a later
|
|
report can address by `conversation_id` without guessing `source` or ambiguity.
|
|
|
|
**Plugin loader contract:** the opencode plugin that stamps these headers
|
|
must export only functions — the opencode 1.18.x plugin loader rejects a
|
|
plugin outright if any export is not a function (so the plugin's
|
|
`parentCache` lives as a property on the factory rather than a named export).
|
|
See [clients.md](clients.md) for the full contract.
|
|
|
|
**Capability 422s**: When no model survives the hard filters, the 422 names the
|
|
active constraints. That now includes "vision-capable model" or
|
|
"json-mode-capable model" when the request carried images or a JSON-mode
|
|
`response_format`, alongside the existing context/tier/latency/tool reasons.
|
|
|
|
**`POST /outcome`**: everything else the router records is a proxy — structural
|
|
checks know whether code *parses*, the local LLM check guesses whether prose
|
|
*looks* right, neither knows whether the answer did the job. The client does,
|
|
because it ran the tests:
|
|
|
|
```bash
|
|
# id comes from the completion body, or any stream chunk
|
|
curl -s localhost:8080/outcome -H 'content-type: application/json' \
|
|
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
|
|
```
|
|
|
|
An unknown `request_id` returns `404` rather than being quietly accepted, so a
|
|
client whose reports go nowhere finds out. The lookup checks
|
|
`energy_observations` first (cloud), then `local_energy_observations` (local
|
|
rows carry the same `request_id`/`session_dir` attribution) — cloud wins on a
|
|
rare id collision.
|
|
|
|
`conversation_id` is an optional extra body field that makes attribution
|
|
exact: it names the conversation (the same `X-Router-Conversation` the client
|
|
sent on the request) so a report can land without a `request_id`, and without
|
|
guessing which source the row came from.
|
|
|
|
```bash
|
|
# attribute a report to a conversation, no request_id needed
|
|
curl -s localhost:8080/outcome -H 'content-type: application/json' \
|
|
-d '{"conversation_id":"abc-123","ok":false,"detail":"tests failed"}'
|
|
```
|
|
|
|
Resolution order is strict:
|
|
|
|
1. `request_id` — exact id; unmatched returns `404`.
|
|
2. `conversation_id` (when present) — exact namespaced match on the
|
|
conversation's rows; unmatched returns `404` and never falls through.
|
|
3. `source` / unambiguous fallback — only when `conversation_id` is entirely
|
|
absent (no `source` given and exactly one plausible row).
|
|
|
|
When `conversation_id` is present but invalid (wrong charset or length) the
|
|
report is rejected with `422`, because a malformed id is a client bug, not an
|
|
absence.
|
|
|
|
Unlike the structural/local-LLM checks — which only ever record failures —
|
|
`/outcome` folds **both** directions into `proficiency` via `feedback.py`: a
|
|
`false` report counts against the model same as any other verification failure,
|
|
but a `true` report counts too. It's also the only quality signal that survives
|
|
streaming, since a retry can't reach a response whose bytes are already gone,
|
|
while a report arrives afterward and works either way.
|
|
|
|
## Admin watchdog and availability API
|
|
|
|
The admin portal's watchdog and model-availability operations sit under
|
|
`/admin` (loopback-only, like the rest of the admin API). The watchdog
|
|
endpoints back the dashboard's **Loops** card and the **Watchdog** control
|
|
card; full behaviour is described in [watchdog.md](watchdog.md) and
|
|
[admin-portal.md](admin-portal.md).
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count |
|
|
| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) |
|
|
| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings |
|
|
| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity |
|
|
| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels |
|
|
|
|
`GET /admin/api/watchdog/status` returns a single object:
|
|
|
|
- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the
|
|
first tick
|
|
- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or
|
|
`null` when there is no tick yet
|
|
- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null
|
|
|
|
`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows
|
|
where `resolved_at` is null, ordered by `opened_at` descending. Each element
|
|
is the full alert row.
|
|
|
|
`GET /admin/api/watchdog/channels` returns an array of
|
|
`watchdog_channel_settings` rows ordered by `channel_name`.
|
|
|
|
`POST /admin/api/watchdog/channels` upserts one channel. Body:
|
|
|
|
- `channel_name` (required) — the channel to create or update; absence is a
|
|
422
|
|
- `enabled` (optional boolean) — when omitted the existing value is kept
|
|
- `min_severity` (optional: `info`, `warning`, or `critical`) — any other
|
|
value is a 422
|
|
|
|
Returns `{ok: true}`.
|
|
|
|
`POST /admin/api/watchdog/test-alert` delivers a test alert through the
|
|
currently enabled channels. Body:
|
|
|
|
- `severity` (optional, default `warning`; `info`, `warning`, or `critical`)
|
|
- `channel_name` (optional) — restrict to a single channel
|
|
|
|
Returns `{sent: <n>, severity: <severity>}`, or `{sent: 0, message: "no
|
|
enabled channels"}` when nothing is enabled. An invalid `severity` is a 422.
|
|
|
|
**Model availability** — `POST
|
|
/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin
|
|
override for a model's availability. `availability` must be one of `active`,
|
|
`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an
|
|
operator stop that routing and pinned requests exclude
|
|
(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`.
|
|
`DELETE` on the same path removes the override. The `{model_id:path}`
|
|
converter accepts slash-bearing model ids (e.g. OpenRouter).
|
|
|
|
## Pinning and auto behavior for local models
|
|
|
|
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
|
|
classifier returns `file_summarization` or `diff_checking` and the local row
|
|
survives the filters; otherwise a cloud model is selected as usual. Pin any
|
|
local `model_id` (e.g. `qwen2.5-coder-router:14b`) and the request goes straight
|
|
to that Ollama tag. `/v1/models` lists local rows with `owned_by: "ollama-local"`.
|