Files
6krrt/deploy
..

Deploying the router

The dispatcher runs continuously; the rest are one-shots on timers. The poller is not optional — freshness.stale_after_days is 3 and freshness.exclude_stale is true, so a catalog that goes unpolled for three days marks every row stale and the router stops returning any candidate at all.

file what it does
llm-router.service the FastAPI dispatcher, on 127.0.0.1:8080
llm-router-poller.service one-shot: PYTHONPATH=src python -m poller then PYTHONPATH=src python -m tier
llm-router-poller.timer fires the poller 2 min after boot, then every 2 h
llm-router-seed.service one-shot: a small PYTHONPATH=src python -m seed_energy reference sweep
llm-router-seed.timer every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep
llm-router-feedback.service one-shot: PYTHONPATH=src python -m feedback — folds client outcomes into proficiency. Ships not enabled, see below
llm-router-feedback.timer every 12 h once you enable it

These are user units — no root, and they run as you with your own $HOME. The tradeoff is that a user service does not inherit your shell environment, so the API key has to come from a file.

Install

# 1. The key. User units don't see your shell env, so .env is required.
cd /path/to/this/repo
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env

# 2. Install, pointing the units at wherever you actually cloned this.
#    The shipped units say %h/llm-router; %h is systemd's expansion for your
#    home directory, so only the part after it needs changing. Getting this
#    wrong fails at start with status=200/CHDIR rather than anything obvious.
REPO=$(pwd)
mkdir -p ~/.config/systemd/user
for u in deploy/llm-router*.{service,timer}; do
  sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
# llm-router-feedback.timer is deliberately NOT in this line — the first fold
# is irreversible. See "The feedback fold, which ships switched off" below.
systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer

# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten
# by the same sed to your REPO) so the modules under `src/` are importable
# without an editable install. .env is read from the repo root and stays there.

# 3. Survive logout/reboot (user units stop with your session otherwise)
loginctl enable-linger "$USER"

# 4. Check. The health endpoint reports whether cost/eco/proficiency
#    actually have data behind them, which is otherwise silent.
curl -s localhost:8080/health | python -m json.tool
systemctl --user list-timers 'llm-router*' 

Operating it

systemctl --user status llm-router.service
journalctl --user -u 'llm-router*' -f            # everything, live (quote the glob)
journalctl --user -u llm-router -f -o cat        # the request log, message only
journalctl --user -u llm-router -p warning       # fallbacks, retries, refusals
journalctl --user -u llm-router-poller.service   # catalog refreshes
systemctl --user restart llm-router.service      # after editing config/config.yaml
systemctl --user start llm-router-poller.service # force a refresh now

config/config.yaml is read once at startup, so weight and threshold changes need a restart. The catalog is read per-request, so a poller run takes effect immediately.

The feedback fold, which ships switched off

POST /outcome is the only ground truth this router has, and feedback.py is what turns those reports into routing changes. It is the one loop here with no timer, so until you enable one it runs only when someone remembers to run it — on this deployment that meant a 404-sample backlog and a proficiency table five days stale while thousands of decisions were routed off it.

The install loop above copies these two units in with the rest. Neither is enabled, and the .service has no [Install] section at all, so it cannot be enabled on its own — only the timer can:

systemctl --user start llm-router-feedback.service   # fold once, now
systemctl --user enable --now llm-router-feedback.timer   # and every 12 h after

That is deliberate, and it is the only unit in this directory treated this way. The fold is irreversible. add_outcome accumulates into a running mean and recompute_category re-derives every row in the category from the new peer rate; neither keeps the pre-fold value anywhere, and verifications.applied_at means a second run will not redo the work either. There is no undo and no restore short of the backup timer. Turning that on for the first time against an accumulated backlog is an operator decision, not a default.

So look first. The dry run copies the database into memory, runs the real add_outcome against the copy, and reports the before/after per row. It opens the source read-only and writes nothing to it:

PYTHONPATH=src .venv/bin/python -m feedback --dry-run           # the live DB, safely
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --csv     # per-row, for a spreadsheet
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --db /tmp/copy.db

Two kinds of movement come back, and the second is the one that surprises people. direct rows have new outcomes of their own. ripple rows have none and move anyway, because the whole category is re-derived against a peer rate the new evidence just changed — on this deployment 404 samples across 17 pairs moved 17 rows directly and 363 by ripple. The table digests ripple per category; --csv carries every row.

Cadence is 12 h rather than the poller's 2 h because the fold is exactly additive: ten folds of five samples land on the same numbers as one fold of fifty, so cadence cannot change where the scores end up, only how stale routing's inputs get and how large each irreversible batch is. Twelve hours caps staleness at half a day while keeping a run a reviewable batch.

The unit takes no EnvironmentFile and no network-online.target: the fold makes no provider call and no HTTP request of any kind.

Turning up the logs

logging.level in config/config.yaml is the documented setting, but flipping it means editing a tracked file. For a running service use a drop-in instead:

systemctl --user edit llm-router     # Environment="LLM_ROUTER_LOG_LEVEL=debug"
systemctl --user restart llm-router

info gives one route and one dispatch line per request — category, tier, model chosen, cost, latency. debug adds every candidate that was dropped and by which filter, plus the ranking with scores. No conversation text is logged at any level.

Every line carries a trace id, and the dispatch line carries the provider's completion id and the session fingerprint, both of which are columns in energy_observations:

journalctl --user -u llm-router --grep ' id=r9116d9'     # one request, all stages
journalctl --user -u llm-router --grep 'chatcmpl-abc123' # from a DB row back to its decision

Severity filtering works because the service prefixes its lines with journald priorities when systemd owns its stderr (SyslogLevelPrefix is on by default). A foreground uvicorn prints them clean, so the same binary is readable either way.

The oneshot units buffer. python -m poller and python -m seed_energy print progress with plain print(), and Python block-buffers stdout when it is not a terminal, so their output arrives in one dump at exit rather than progressively. Add Environment="PYTHONUNBUFFERED=1" to those units if you want to watch a sweep as it runs.

A note on the bind address

--host 127.0.0.1 is deliberate. The service holds a billable API key and has no authentication of its own — anything that reaches it can spend your allowance. ProtectHome=read-only plus a ReadWritePaths exception for the repo limits the blast radius on the filesystem, but nothing limits spend. Putting this on a LAN address needs an auth layer first.

If you enable classifier.mode: local_encoder, note HF_HOME is redirected into the repo (%h/llm-router/.hf-cache) for the same reason — Hugging Face's default cache lives outside ReadWritePaths and the service will crash-loop trying to download a model into a read-only home directory otherwise. Run pip install -r requirements-encoder.txt before switching to this mode; the startup check refuses to boot with a clear message if it's missing, rather than failing opaquely on the first request, but a missing model still needs the dependency installed first.

Using an Ollama on another machine

The local LLM does the classifying; it does not have to be on the machine you are typing on, and usually the GPU isn't. Most developers already have WireGuard or a VPN back to a home lab, so the normal shape is router and editor on the laptop, Ollama on the workstation.

On the serving host (the one with the GPU):

sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo cp deploy/ollama-over-vpn.conf \
        /etc/systemd/system/ollama.service.d/override.conf
# set OLLAMA_HOST to that host's own VPN address — `ip -4 -o addr show`
sudo nano /etc/systemd/system/ollama.service.d/override.conf
sudo systemctl daemon-reload && sudo systemctl restart ollama

On the client host, in config/config.yaml:

classifier:
  base_url: "http://<vpn-ip>:11434/v1"
verification:
  base_url: "http://<vpn-ip>:11434"    # same host, so `model` can stay null

Both must move together. verification speaks Ollama's native /api/chat and used to derive its URL from the classifier's; it no longer does, so pointing only the classifier across the tunnel leaves the verifier talking to a localhost Ollama that may not exist. Config load refuses the combination where verification.model is null and the two hosts differ, because that failure is otherwise silent — the verifier 404s, catches it, and records no sample while appearing to be enabled.

Bind Ollama to the VPN address, not 0.0.0.0. It has no authentication of any kind: anything that reaches the port can run inference, enumerate your models and pull new ones. Same reasoning as the dispatcher's loopback bind above.

A cloud endpoint works too — set classifier.api_key_env to the env var holding its key. On a five-prompt comparison NeuralWatt's deepseek-v4-flash classified in 1.02s mean against qwen3.5's 11.58s on an RTX 6000. Local inference is not free, it is unbilled.

Pointing opencode at it

The repo-local opencode.json sets this up already, so running opencode from inside a clone of this repo uses the router by default. To use it from anywhere, merge the provider.llm-router block into ~/.config/opencode/opencode.json and set "model": "llm-router/auto".

Two model names:

  • llm-router/auto — normal routing; flex rows excluded, so nothing gets held server-side during peak
  • llm-router/auto:batch — admits flex rows, for overnight/async work

limit.context is declared as 782324, the largest effective window in the routable catalog. The router hard-filters on the measured conversation size, so a prompt too big for the smaller models simply won't be routed to them; if it fits nothing, /v1/chat/completions returns a 422 naming the constraint rather than truncating.

opencode plugin

deploy/opencode-plugin/router-link.js is an opencode plugin that does two things:

  1. Stamps conversation identity on outbound requests to the router: X-Router-Conversation (the session id), X-Router-Agent (the agent name), and X-Router-Parent (the session's parent, if any). Only the llm-router provider gets these headers — another provider (Anthropic, OpenAI, …) would not understand them and they would just leak identity. Agent names are slugged to the router's header charset (lowercase; runs of characters outside a-z0-9._:- become -, trimmed, capped at 64), so a multi-word name like Sisyphus - ultraworker arrives as sisyphus-ultraworker instead of being dropped as invalid.

  2. Reports test outcomes back to the router's /outcome endpoint so feedback.py can fold client-side pass/fail into proficiency scoring. (tool.execute.after hook — watches for pytest, cargo test, go test, tsc, ruff, eslint, etc.)

Install

# cp is aliased to -i in the default zsh; `command cp` bypasses the alias.
command cp -f deploy/opencode-plugin/router-link.js ~/.config/opencode/plugins/
rm ~/.config/opencode/plugins/router-outcome.js   # <-- MUST remove this

The old file must be removed because both plugins hook tool.execute.after. Keeping both would post every outcome twice, and feedback.py would count each test result twice — once for the real result and once for a duplicate that looks identical but belongs to a different model in a different conversation. A test counted twice penalises the incumbent model (it accrues extra failures) while never appearing in the challenger's ledger, which is the worst attribution error in the system.

Turn it off

Remove the file from ~/.config/opencode/plugins/ (or the per-project .opencode/plugins/). Without a plugin, the router falls back to the fingerprint-based heuristics it used before: the working directory from the /v1/chat/completions request body's source field (set by the opencode SDK), matched against the session.directory in the /outcome body. That is less accurate (it cannot distinguish two sessions in the same directory), but it is functional.