diff --git a/README.md b/README.md index 8270867..d7e4b46 100644 --- a/README.md +++ b/README.md @@ -197,21 +197,51 @@ curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application -d '{"model":"auto","messages":[{"role":"user","content":"hello"}]}' ``` -- Ask for `model: "auto"` and the router picks per request. -- Ask for any real model id and the router dispatches directly, still logged. -- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines. - -### Route overnight/batch work through flex rows - -`-flex` rows are discounted asynchronous rows that may be held during peak. +Ask for a **virtual router model** and the router picks a candidate, subject +to the profile you specify. The general form is `auto:`; asking for +just `auto` is shorthand for `auto:default`. ```bash -curl -s -X POST localhost:8080/route -H 'content-type: application/json' \ - -d '{"task":"nightly code review","latency_tolerance":"batch"}' +# Normal interactive routing (same as "auto") +curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ + -d '{"model":"auto","messages":...}' + +# Restrict to local dispatch models (ollama-local provider only) +curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ + -d '{"model":"auto:locality","messages":...}' + +# Restrict to models priced at ≤ $0.50/1M completion tokens +curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ + -d '{"model":"auto:onlycheaps","messages":...}' + +# Restrict to frontier tier (tier 3) +curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ + -d '{"model":"auto:bigboybritches","messages":...}' + +# Overnight/async work via flex rows +curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application/json' \ + -d '{"model":"auto:batch","messages":...}' ``` -- `auto` excludes `-flex` rows; `auto:batch` admits them. -- `latency_tolerance: batch` on any request flips the router into the batch serving class. +| Profile | What it does | +|---|---| +| `auto` / `auto:default` | Normal quality-first routing; `-flex` rows excluded | +| `auto:batch` | Admits `-flex` rows that may be held during peak hours | +| `auto:locality` | Routes only to the `ollama-local` provider (local dispatch) | +| `auto:onlycheaps` | Limits to models at or below $0.50 per 1M completion tokens | +| `auto:bigboybritches` | Routes only to tier-3 (frontier) models | + +Profiles are **candidate-set filters** — they narrow which models the router +may pick from, but do not change the ranking objective (quality-first, cost as +tiebreak). An unknown profile raises HTTP 422. + +Operators can define custom profiles under `profiles:` in +`config/config.yaml`. Each entry accepts `provider`, `min_tier`/`max_tier`, +`max_cost_per_1m_completion`, `latency_tolerance`, and `allowed_model_ids` +as rewrite rules on top of the default profile. + +- Ask for any real model id and the router dispatches directly, still logged. +- Streaming is supported: tokens pass through as they arrive while the router scrapes energy/cost from provider SSE comment lines. ### Ask an image question diff --git a/admin/frontend/decisions.html b/admin/frontend/decisions.html index bbf59a0..016b5aa 100644 --- a/admin/frontend/decisions.html +++ b/admin/frontend/decisions.html @@ -205,6 +205,9 @@ header.navbar{padding-top:2px!important;padding-bottom:2px!important} +