Files
6krrt/docs/api.md
2026-09-26 16:50:24 -04:00

17 KiB

API surface reference. Back to README.

API Endpoints

The dispatcher binds 127.0.0.1:8080. No auth of its own — loopback is the only thing standing between the open internet and your billing allowance.

Method Path Description
GET /health Catalog/reachability status, scoring coverage, warnings (including models that can never enter the candidate set)
GET /metrics Aggregated observability JSON: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency, cost-estimator calibration, router-observed latency; loopback-only, no auth
GET /events/decisions Server-Sent Events stream of routing decisions for the TUI's live feed; replays recent decisions, then streams new ones as they happen
POST /route Classify task, rank candidates, return selected model — no provider call, no cost
POST /dispatch Same as /route, plus complete the provider call, stream response, log observation
GET /v1/models OpenAI-compatible model list (router virtual models + catalog)
POST /v1/chat/completions OpenAI-compatible completions — routes then proxies, streaming supported
POST /outcome Client reports whether a completion actually worked — the only signal that knows the answer did the job, not just that it parsed
GET /admin Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits)

/metrics returns a single JSON object with these top-level keys:

  • quota — per-provider billing shape from quota_accounts(): period (billing window with start/next_reset/elapsed_fraction/source), accounts[] (list of per-provider dicts with provider, shape, spend_usd, and type-specific blocks: plan for metered_plan, pool for prepaid_credit, burn for burn-rate metrics, credit when a balance URL was polled, energy for kWh/calls), spend (aggregate spend with by_provider_usd, total_usd, estimated_usd, and estimate_ratio), and alarm (plan-pace or stale-reading alert with kind/severity/headline). The old by_provider/total_balance_usd shape and flat top-level keys were removed.
  • coverage — routable-model counts (routable_models counts active models filtered on access_level IN (routing.allowed_access_levels), excluding any access-gated rows whose access_level is not in that list) with energy and proficiency data, plus a selection sub-key reporting which active models (unfiltered by access-level) the router has actually picked over the configured window, and a warnings array that now also includes entries for active models that can never enter the candidate set (structural-gate failures: zero context window or broken tier derived from the catalog, detected via routing.rejection_reason under the most permissive request — one token, tier 1, batch latency). Existing warnings cover missing energy data, missing proficiency data, stale catalog age, capacity-demand ceiling breaches, and classifier source degradation. A cache sub-key carries the prefix-cache series from cache_rate_series(): window_hours, assumed_cache_rate, aggregate observations / prompt_tokens / cached_prompt_tokens / cache_rate, and a by_model list of the same figures per (provider, model_id). The rate is sum(cached_prompt_tokens) / sum(prompt_tokens) over rows the provider actually reported a count for (cached_tokens_source = 'reported'), excluding seed_reference sweeps. Two warnings read it: cache rate: when the aggregate diverges from objective.assumed_cache_rate by more than objective.cache_rate_warn_margin, and cache rate outlier: when one (provider, model_id) group does. Both need objective.cache_rate_warn_min_observations reported-cache rows first.
  • recent_decisions — last 50 rows from route_decisions
  • per_model — per-model aggregates over the last 30 days of energy_observations
  • verdict_mix — counts by verification verdict over the last 7 days
  • top_proficiency — top models by blended_score for coding_general
  • generated_at — ISO8601 timestamp
  • pinch — context-pruning aggregation (when pinch.enabled): calls_30d, pruned_calls_30d, and share_pruned (share of decisions pruned), plus median and total tokens_saved over 30 days and estimated dollars saved at the blended rate — exposes no conversation text
  • cost_calibration — from cost_estimate_calibration(): how far routing.estimated_cost is from the provider's own bill, per (provider, model_id). Reported only — nothing applies it to estimated_cost or to ranking. route_decisions joined to energy_observations on request_id and on the selected model and provider (the id alone pairs a failover's estimate with another model's bill), over objective.cost_calibration_window_hours, excluding seed_reference on both sides and rows with no positive cost_usd or est_cost_usd. Keys: window_hours, min_observations, observations, est_cost_usd, billed_cost_usd, overestimate_ratio (est/billed), correction_factor (billed/est — what would be multiplied in), spread with spread_low / spread_high, and a by_model list of the same figures plus est_per_request_usd, billed_per_request_usd and sufficient. spread is the figure that matters: a uniform scale error reorders nothing, while the spread in that error does. Groups below objective.cost_calibration_min_observations are listed with sufficient: false and excluded from spread. No warning class reads it.
  • latency — from latency_series(): p50 and p95 of the router-observed router_wall_seconds and router_ttft_seconds per (provider, model_id) over objective.latency_window_hours, excluding seed_reference. A report, not an objective — nothing in the ranking reads it. These are the router's own clock, not the provider's duration_seconds, which is why OpenRouter models are visible here at all (OpenRouter reports no duration). Keys: window_hours, min_observations, aggregate wall_observations / ttft_observations / wall_p50 / wall_p95 / ttft_p50 / ttft_p95, and a by_model list of the same plus wall_sufficient / ttft_sufficient. The wall and TTFT counts are independent because TTFT is streaming-only; a group can be sufficient on one and not the other. Degrades to an empty series on a database that predates the two columns. No warning class reads it.

It exposes no conversation text, prompts, or session_dir; it is bound to loopback and unauthenticated exactly like /health.

Input to /route and /dispatch can include task_category, task_tier, and required_context_tokens overrides — these skip the classifier, useful for testing routing without the classifier in the loop.

Named routing profiles (auto:<profile>). The general form is auto:<profile>; auto alone is shorthand for auto:default. The router resolves the profile name against the built-in set and any custom entries in config/config.yaml under profiles:. Profiles narrow the candidate set but do not override the quality-first ranking objective.

Profile Effect
auto:default Normal quality-first routing, -flex rows excluded
auto:batch Admits -flex rows for overnight and async work
auto:locality Restricts to the ollama-local provider only
auto:onlycheaps Limits to models priced at ≤ $0.50 per 1M completion tokens
auto:bigboybritches Restricts to tier-3 (frontier) models only

From opencode: a profile is just another entry under provider.llm-router.models in opencode.json — the object's key is the literal model id opencode sends upstream, so auto:batch already works this way (see the repo-local opencode.json). To make auto:locality, auto:onlycheaps, or auto:bigboybritches selectable from opencode's model picker instead of typing an override, add one entry per profile:

"auto:onlycheaps": {
  "name": "auto (cost ceiling)",
  "limit": { "context": 782324, "output": 16384 },
  "modalities": { "input": ["text", "image"] }
}

limit/modalities can be copied from the auto entry unchanged — a profile narrows which models are candidates, not the context window or input types the router itself accepts. See clients.md for the full opencode wiring.

An unknown profile name raises HTTP 422 and lists the valid names.

The profile field is also accepted on POST /route and POST /dispatch to select a profile inline per request.

Ask for any real model id in /v1/chat/completions and it dispatches directly, still logged — routing is transparent, not opaque.

Streaming (chunk-by-chunk proxy): Tokens render as they arrive. Neuralwatt emits energy and cost as SSE comment lines (: energy {...}) before data: [DONE] — ordinary clients ignore comments, so the stream flows untouched while the router scrapes telemetry on the way past. Without this, streamed calls would log no energy at all.

Verification headers (non-streaming): When streaming is not used, the structural verification verdict surfaces in the X-Router-Verification header so a client can inspect it without parsing the response body. Valid values: ok, truncated, malformed, unverifiable, none.

Conversation identity headers (optional): A client may attach three headers to POST /v1/chat/completions so the router can group its decisions by conversation and attribute an answer to the exact conversation that asked for it. All three are optional; an absent or invalid header is treated as absent (it is never an error).

Header Meaning Length Pattern
X-Router-Conversation Id of ONE conversation (e.g. an opencode sessionID) 1-128 chars ^[A-Za-z0-9._:-]+$
X-Router-Agent Name of the agent making the request 1-64 chars ^[A-Za-z0-9._:-]+$
X-Router-Parent Parent conversation id, when this one is a sub-conversation 1-128 chars ^[A-Za-z0-9._:-]+$

Contract rules:

  • Charset: letters, digits, ., _, :, - (the regex ^[A-Za-z0-9._:-]+$); anything else, or a length outside the table's range, makes the header invalid and therefore absent.
  • An absent or invalid header is treated as absent, never as an error: the request is still routed normally.
  • Namespacing: a valid X-Router-Conversation produces session_key = "c:" + conversation, so client conversations are keyed separately from the router's own hashed prompt fingerprints (which carry no c: prefix). A valid X-Router-Parent produces parent_key = "c:" + parent. Conversation ids never appear alongside message text: session_key stores only the namespaced id, never the messages.
  • When X-Router-Conversation is absent or invalid, the router falls back to its hashed prompt fingerprint as session_key (see routing.md "incumbency and cache pricing").

These ids are what make /outcome attribution exact (below): an X-Router-Conversation on the request sets the session_key row that a later report can address by conversation_id without guessing source or ambiguity.

Plugin loader contract: the opencode plugin that stamps these headers must export only functions — the opencode 1.18.x plugin loader rejects a plugin outright if any export is not a function (so the plugin's parentCache lives as a property on the factory rather than a named export). See clients.md for the full contract.

Capability 422s: When no model survives the hard filters, the 422 names the active constraints. That now includes "vision-capable model" or "json-mode-capable model" when the request carried images or a JSON-mode response_format, alongside the existing context/tier/latency/tool reasons.

POST /outcome: everything else the router records is a proxy — structural checks know whether code parses, the local LLM check guesses whether prose looks right, neither knows whether the answer did the job. The client does, because it ran the tests:

# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'

An unknown request_id returns 404 rather than being quietly accepted, so a client whose reports go nowhere finds out. The lookup checks energy_observations first (cloud), then local_energy_observations (local rows carry the same request_id/session_dir attribution) — cloud wins on a rare id collision.

conversation_id is an optional extra body field that makes attribution exact: it names the conversation (the same X-Router-Conversation the client sent on the request) so a report can land without a request_id, and without guessing which source the row came from.

# attribute a report to a conversation, no request_id needed
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"conversation_id":"abc-123","ok":false,"detail":"tests failed"}'

Resolution order is strict:

  1. request_id — exact id; unmatched returns 404.
  2. conversation_id (when present) — exact namespaced match on the conversation's rows; unmatched returns 404 and never falls through.
  3. source / unambiguous fallback — only when conversation_id is entirely absent (no source given and exactly one plausible row).

When conversation_id is present but invalid (wrong charset or length) the report is rejected with 422, because a malformed id is a client bug, not an absence.

Unlike the structural/local-LLM checks — which only ever record failures — /outcome folds both directions into proficiency via feedback.py: a false report counts against the model same as any other verification failure, but a true report counts too. It's also the only quality signal that survives streaming, since a retry can't reach a response whose bytes are already gone, while a report arrives afterward and works either way.

Admin watchdog and availability API

The admin portal's watchdog and model-availability operations sit under /admin (loopback-only, like the rest of the admin API). The watchdog endpoints back the dashboard's Loops card and the Watchdog control card; full behaviour is described in watchdog.md and admin-portal.md.

Method Path Description
GET /admin/api/watchdog/status Latest watchdog tick plus open-alert count
GET /admin/api/watchdog/loops Open alert rows (one per looping session)
GET /admin/api/watchdog/channels Notification-channel settings
POST /admin/api/watchdog/channels Create or update a channel's enabled/min-severity
POST /admin/api/watchdog/test-alert Deliver a test alert through the enabled channels

GET /admin/api/watchdog/status returns a single object:

  • last_tick — the most recent watchdog_ticks row, or null before the first tick
  • verdicts — {total, flagged} for that tick's watchdog_verdicts, or null when there is no tick yet
  • open_alerts — count of watchdog_alerts rows with resolved_at null

GET /admin/api/watchdog/loops returns an array of watchdog_alerts rows where resolved_at is null, ordered by opened_at descending. Each element is the full alert row.

GET /admin/api/watchdog/channels returns an array of watchdog_channel_settings rows ordered by channel_name.

POST /admin/api/watchdog/channels upserts one channel. Body:

  • channel_name (required) — the channel to create or update; absence is a 422
  • enabled (optional boolean) — when omitted the existing value is kept
  • min_severity (optional: info, warning, or critical) — any other value is a 422

Returns {ok: true}.

POST /admin/api/watchdog/test-alert delivers a test alert through the currently enabled channels. Body:

  • severity (optional, default warning; info, warning, or critical)
  • channel_name (optional) — restrict to a single channel

Returns {sent: <n>, severity: <severity>}, or {sent: 0, message: "no enabled channels"} when nothing is enabled. An invalid severity is a 422.

Model availability — POST /admin/api/models/{model_id:path}/{provider}/availability upserts an admin override for a model's availability. availability must be one of active, blocked, deprecated, or stale (anything else is a 422); blocked is an operator stop that routing and pinned requests exclude (_admin_excluded_models), distinct from the catalog meaning of deprecated. DELETE on the same path removes the override. The {model_id:path} converter accepts slash-bearing model ids (e.g. OpenRouter).

Pinning and auto behavior for local models

model: "auto" will route eligible tasks to qwen2.5-coder-router:14b when the classifier returns file_summarization or diff_checking and the local row survives the filters; otherwise a cloud model is selected as usual. Pin any local model_id (e.g. qwen2.5-coder-router:14b) and the request goes straight to that Ollama tag. /v1/models lists local rows with owned_by: "ollama-local".