285 lines
14 KiB
Markdown
285 lines
14 KiB
Markdown
# Deploying the router
|
|
|
|
The dispatcher runs continuously; the rest are one-shots on timers. The poller
|
|
is **not** optional — `freshness.stale_after_days` is 3 and
|
|
`freshness.exclude_stale` is true, so a catalog that goes unpolled for three
|
|
days marks every row stale and the router stops returning any candidate at
|
|
all.
|
|
|
|
| file | what it does |
|
|
|---|---|
|
|
| `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` |
|
|
| `llm-router-poller.service` | one-shot: `PYTHONPATH=src python -m poller` then `PYTHONPATH=src python -m tier` |
|
|
| `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h |
|
|
| `llm-router-seed.service` | one-shot: a small `PYTHONPATH=src python -m seed_energy` reference sweep |
|
|
| `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep |
|
|
| `llm-router-feedback.service` | one-shot: `PYTHONPATH=src python -m feedback` — folds client outcomes into `proficiency`. **Ships not enabled**, see below |
|
|
| `llm-router-feedback.timer` | every 12 h once you enable it |
|
|
|
|
These are **user** units — no root, and they run as you with your own
|
|
`$HOME`. The tradeoff is that a user service does not inherit your shell
|
|
environment, so the API key has to come from a file.
|
|
|
|
## Install
|
|
|
|
```bash
|
|
# 1. The key. User units don't see your shell env, so .env is required.
|
|
cd /path/to/this/repo
|
|
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
|
|
|
|
# 2. Install, pointing the units at wherever you actually cloned this.
|
|
# The shipped units say %h/llm-router; %h is systemd's expansion for your
|
|
# home directory, so only the part after it needs changing. Getting this
|
|
# wrong fails at start with status=200/CHDIR rather than anything obvious.
|
|
REPO=$(pwd)
|
|
mkdir -p ~/.config/systemd/user
|
|
for u in deploy/llm-router*.{service,timer}; do
|
|
sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")"
|
|
done
|
|
systemctl --user daemon-reload
|
|
# llm-router-feedback.timer is deliberately NOT in this line — the first fold
|
|
# is irreversible. See "The feedback fold, which ships switched off" below.
|
|
systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer
|
|
|
|
# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten
|
|
# by the same sed to your REPO) so the modules under `src/` are importable
|
|
# without an editable install. .env is read from the repo root and stays there.
|
|
|
|
# 3. Survive logout/reboot (user units stop with your session otherwise)
|
|
loginctl enable-linger "$USER"
|
|
|
|
# 4. Check. The health endpoint reports whether cost/eco/proficiency
|
|
# actually have data behind them, which is otherwise silent.
|
|
curl -s localhost:8080/health | python -m json.tool
|
|
systemctl --user list-timers 'llm-router*'
|
|
```
|
|
|
|
## Operating it
|
|
|
|
```bash
|
|
systemctl --user status llm-router.service
|
|
journalctl --user -u 'llm-router*' -f # everything, live (quote the glob)
|
|
journalctl --user -u llm-router -f -o cat # the request log, message only
|
|
journalctl --user -u llm-router -p warning # fallbacks, retries, refusals
|
|
journalctl --user -u llm-router-poller.service # catalog refreshes
|
|
systemctl --user restart llm-router.service # after editing config/config.yaml
|
|
systemctl --user start llm-router-poller.service # force a refresh now
|
|
```
|
|
|
|
`config/config.yaml` is read once at startup, so weight and threshold changes need a
|
|
restart. The catalog is read per-request, so a poller run takes effect
|
|
immediately.
|
|
|
|
## The feedback fold, which ships switched off
|
|
|
|
`POST /outcome` is the only ground truth this router has, and `feedback.py` is
|
|
what turns those reports into routing changes. It is the one loop here with no
|
|
timer, so until you enable one it runs only when someone remembers to run it —
|
|
on this deployment that meant a 404-sample backlog and a `proficiency` table
|
|
five days stale while thousands of decisions were routed off it.
|
|
|
|
The install loop above copies these two units in with the rest. **Neither is
|
|
enabled**, and the `.service` has no `[Install]` section at all, so it cannot be
|
|
enabled on its own — only the timer can:
|
|
|
|
```bash
|
|
systemctl --user start llm-router-feedback.service # fold once, now
|
|
systemctl --user enable --now llm-router-feedback.timer # and every 12 h after
|
|
```
|
|
|
|
That is deliberate, and it is the only unit in this directory treated this way.
|
|
**The fold is irreversible.** `add_outcome` accumulates into a running mean and
|
|
`recompute_category` re-derives every row in the category from the new peer
|
|
rate; neither keeps the pre-fold value anywhere, and `verifications.applied_at`
|
|
means a second run will not redo the work either. There is no undo and no
|
|
restore short of the backup timer. Turning that on for the first time against
|
|
an accumulated backlog is an operator decision, not a default.
|
|
|
|
So look first. The dry run copies the database into memory, runs the real
|
|
`add_outcome` against the copy, and reports the before/after per row. It opens
|
|
the source read-only and writes nothing to it:
|
|
|
|
```bash
|
|
PYTHONPATH=src .venv/bin/python -m feedback --dry-run # the live DB, safely
|
|
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --csv # per-row, for a spreadsheet
|
|
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --db /tmp/copy.db
|
|
```
|
|
|
|
Two kinds of movement come back, and the second is the one that surprises
|
|
people. **direct** rows have new outcomes of their own. **ripple** rows have
|
|
none and move anyway, because the whole category is re-derived against a peer
|
|
rate the new evidence just changed — on this deployment 404 samples across 17
|
|
pairs moved 17 rows directly and 363 by ripple. The table digests ripple per
|
|
category; `--csv` carries every row.
|
|
|
|
Cadence is 12 h rather than the poller's 2 h because the fold is exactly
|
|
additive: ten folds of five samples land on the same numbers as one fold of
|
|
fifty, so cadence cannot change *where* the scores end up, only how stale
|
|
routing's inputs get and how large each irreversible batch is. Twelve hours
|
|
caps staleness at half a day while keeping a run a reviewable batch.
|
|
|
|
The unit takes no `EnvironmentFile` and no `network-online.target`: the fold
|
|
makes no provider call and no HTTP request of any kind.
|
|
|
|
### Turning up the logs
|
|
|
|
`logging.level` in config/config.yaml is the documented setting, but flipping it means
|
|
editing a tracked file. For a running service use a drop-in instead:
|
|
|
|
```bash
|
|
systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
|
|
systemctl --user restart llm-router
|
|
```
|
|
|
|
`info` gives one `route` and one `dispatch` line per request — category, tier,
|
|
model chosen, cost, latency. `debug` adds every candidate that was dropped and
|
|
by which filter, plus the ranking with scores. No conversation text is logged
|
|
at any level.
|
|
|
|
Every line carries a trace id, and the `dispatch` line carries the provider's
|
|
completion id and the session fingerprint, both of which are columns in
|
|
`energy_observations`:
|
|
|
|
```bash
|
|
journalctl --user -u llm-router --grep ' id=r9116d9' # one request, all stages
|
|
journalctl --user -u llm-router --grep 'chatcmpl-abc123' # from a DB row back to its decision
|
|
```
|
|
|
|
Severity filtering works because the service prefixes its lines with journald
|
|
priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default).
|
|
A foreground `uvicorn` prints them clean, so the same binary is readable either
|
|
way.
|
|
|
|
**The oneshot units buffer.** `python -m poller` and `python -m seed_energy` print progress
|
|
with plain `print()`, and Python block-buffers stdout when it is not a
|
|
terminal, so their output arrives in one dump at exit rather than
|
|
progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you
|
|
want to watch a sweep as it runs.
|
|
|
|
## A note on the bind address
|
|
|
|
`--host 127.0.0.1` is deliberate. The service holds a billable API key and
|
|
has **no authentication of its own** — anything that reaches it can spend
|
|
your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for
|
|
the repo limits the blast radius on the filesystem, but nothing limits spend.
|
|
Putting this on a LAN address needs an auth layer first.
|
|
|
|
If you enable `classifier.mode: local_encoder`, note `HF_HOME` is redirected
|
|
into the repo (`%h/llm-router/.hf-cache`) for the same reason — Hugging
|
|
Face's default cache lives outside `ReadWritePaths` and the service will
|
|
crash-loop trying to download a model into a read-only home directory
|
|
otherwise. Run `pip install -r requirements-encoder.txt` before switching to
|
|
this mode; the startup check refuses to boot with a clear message if it's
|
|
missing, rather than failing opaquely on the first request, but a missing
|
|
model still needs the dependency installed first.
|
|
|
|
## Using an Ollama on another machine
|
|
|
|
The local LLM does the classifying; it does not have to be on the machine you
|
|
are typing on, and usually the GPU isn't. Most developers already have
|
|
WireGuard or a VPN back to a home lab, so the normal shape is router and
|
|
editor on the laptop, Ollama on the workstation.
|
|
|
|
On the **serving** host (the one with the GPU):
|
|
|
|
```bash
|
|
sudo mkdir -p /etc/systemd/system/ollama.service.d
|
|
sudo cp deploy/ollama-over-vpn.conf \
|
|
/etc/systemd/system/ollama.service.d/override.conf
|
|
# set OLLAMA_HOST to that host's own VPN address — `ip -4 -o addr show`
|
|
sudo nano /etc/systemd/system/ollama.service.d/override.conf
|
|
sudo systemctl daemon-reload && sudo systemctl restart ollama
|
|
```
|
|
|
|
On the **client** host, in `config/config.yaml`:
|
|
|
|
```yaml
|
|
classifier:
|
|
base_url: "http://<vpn-ip>:11434/v1"
|
|
verification:
|
|
base_url: "http://<vpn-ip>:11434" # same host, so `model` can stay null
|
|
```
|
|
|
|
Both must move together. `verification` speaks Ollama's native `/api/chat`
|
|
and used to derive its URL from the classifier's; it no longer does, so
|
|
pointing only the classifier across the tunnel leaves the verifier talking to
|
|
a `localhost` Ollama that may not exist. Config load refuses the combination
|
|
where `verification.model` is null and the two hosts differ, because that
|
|
failure is otherwise silent — the verifier 404s, catches it, and records no
|
|
sample while appearing to be enabled.
|
|
|
|
Bind Ollama to the **VPN address, not `0.0.0.0`**. It has no authentication of
|
|
any kind: anything that reaches the port can run inference, enumerate your
|
|
models and pull new ones. Same reasoning as the dispatcher's loopback bind
|
|
above.
|
|
|
|
A cloud endpoint works too — set `classifier.api_key_env` to the env var
|
|
holding its key. On a five-prompt comparison NeuralWatt's `deepseek-v4-flash`
|
|
classified in 1.02s mean against `qwen3.5`'s 11.58s on an RTX 6000. Local
|
|
inference is not free, it is unbilled.
|
|
|
|
## Pointing opencode at it
|
|
|
|
The repo-local `opencode.json` sets this up already, so running `opencode`
|
|
from inside a clone of this repo uses the router by default. To use it from
|
|
anywhere, merge the `provider.llm-router` block into
|
|
`~/.config/opencode/opencode.json` and set `"model": "llm-router/auto"`.
|
|
|
|
Two model names:
|
|
|
|
- `llm-router/auto` — normal routing; flex rows excluded, so nothing gets
|
|
held server-side during peak
|
|
- `llm-router/auto:batch` — admits flex rows, for overnight/async work
|
|
|
|
`limit.context` is declared as 782324, the largest effective window in the
|
|
routable catalog. The router hard-filters on the measured conversation size,
|
|
so a prompt too big for the smaller models simply won't be routed to them;
|
|
if it fits nothing, `/v1/chat/completions` returns a 422 naming the
|
|
constraint rather than truncating.
|
|
|
|
## opencode plugin
|
|
|
|
[`deploy/opencode-plugin/router-link.js`](opencode-plugin/router-link.js) is
|
|
an opencode plugin that does two things:
|
|
|
|
1. **Stamps conversation identity** on outbound requests to the router:
|
|
`X-Router-Conversation` (the session id), `X-Router-Agent` (the agent
|
|
name), and `X-Router-Parent` (the session's parent, if any). Only the
|
|
`llm-router` provider gets these headers — another provider (Anthropic,
|
|
OpenAI, …) would not understand them and they would just leak identity.
|
|
Agent names are slugged to the router's header charset (lowercase; runs of
|
|
characters outside `a-z0-9._:-` become `-`, trimmed, capped at 64), so a
|
|
multi-word name like `Sisyphus - ultraworker` arrives as
|
|
`sisyphus-ultraworker` instead of being dropped as invalid.
|
|
|
|
2. **Reports test outcomes** back to the router's `/outcome` endpoint so
|
|
`feedback.py` can fold client-side pass/fail into proficiency scoring.
|
|
(`tool.execute.after` hook — watches for `pytest`, `cargo test`, `go test`,
|
|
`tsc`, `ruff`, `eslint`, etc.)
|
|
|
|
### Install
|
|
|
|
```bash
|
|
# cp is aliased to -i in the default zsh; `command cp` bypasses the alias.
|
|
command cp -f deploy/opencode-plugin/router-link.js ~/.config/opencode/plugins/
|
|
rm ~/.config/opencode/plugins/router-outcome.js # <-- MUST remove this
|
|
```
|
|
|
|
**The old file must be removed** because both plugins hook
|
|
`tool.execute.after`. Keeping both would post every outcome twice, and
|
|
`feedback.py` would count each test result twice — once for the real result
|
|
and once for a duplicate that looks identical but belongs to a different
|
|
model in a different conversation. A test counted twice penalises the
|
|
incumbent model (it accrues extra failures) while never appearing in the
|
|
challenger's ledger, which is the worst attribution error in the system.
|
|
|
|
### Turn it off
|
|
|
|
Remove the file from `~/.config/opencode/plugins/` (or the per-project
|
|
`.opencode/plugins/`). Without a plugin, the router falls back to the
|
|
fingerprint-based heuristics it used before: the working directory from the
|
|
`/v1/chat/completions` request body's `source` field (set by the opencode
|
|
SDK), matched against the `session.directory` in the `/outcome` body. That
|
|
is less accurate (it cannot distinguish two sessions in the same directory),
|
|
but it is functional.
|