# Deploying the router Three units. The dispatcher runs continuously; the poller runs on a timer and is **not** optional — `freshness.stale_after_days` is 3 and `freshness.exclude_stale` is true, so a catalog that goes unpolled for three days marks every row stale and the router stops returning any candidate at all. | file | what it does | |---|---| | `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` | | `llm-router-poller.service` | one-shot: `poller.py` then `tier.py` | | `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h | | `llm-router-seed.service` | one-shot: a small `seed_energy.py` reference sweep | | `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep | These are **user** units — no root, and they run as you with your own `$HOME`. The tradeoff is that a user service does not inherit your shell environment, so the API key has to come from a file. ## Install ```bash # 1. The key. User units don't see your shell env, so .env is required. cd /path/to/this/repo echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env # 2. Install, pointing the units at wherever you actually cloned this. # The shipped units say %h/llm-router; %h is systemd's expansion for your # home directory, so only the part after it needs changing. Getting this # wrong fails at start with status=200/CHDIR rather than anything obvious. REPO=$(pwd) mkdir -p ~/.config/systemd/user for u in deploy/llm-router*.{service,timer}; do sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")" done systemctl --user daemon-reload systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer # 3. Survive logout/reboot (user units stop with your session otherwise) loginctl enable-linger "$USER" # 4. Check. The health endpoint reports whether cost/eco/proficiency # actually have data behind them, which is otherwise silent. curl -s localhost:8080/health | python -m json.tool systemctl --user list-timers 'llm-router*' ``` ## Operating it ```bash systemctl --user status llm-router.service journalctl --user -u 'llm-router*' -f # everything, live (quote the glob) journalctl --user -u llm-router -f -o cat # the request log, message only journalctl --user -u llm-router -p warning # fallbacks, retries, refusals journalctl --user -u llm-router-poller.service # catalog refreshes systemctl --user restart llm-router.service # after editing config.yaml systemctl --user start llm-router-poller.service # force a refresh now ``` `config.yaml` is read once at startup, so weight and threshold changes need a restart. The catalog is read per-request, so a poller run takes effect immediately. ### Turning up the logs `logging.level` in config.yaml is the documented setting, but flipping it means editing a tracked file. For a running service use a drop-in instead: ```bash systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug" systemctl --user restart llm-router ``` `info` gives one `route` and one `dispatch` line per request — category, tier, model chosen, cost, latency. `debug` adds every candidate that was dropped and by which filter, plus the ranking with scores. No conversation text is logged at any level. Every line carries a trace id, and the `dispatch` line carries the provider's completion id and the session fingerprint, both of which are columns in `energy_observations`: ```bash journalctl --user -u llm-router --grep ' id=r9116d9' # one request, all stages journalctl --user -u llm-router --grep 'chatcmpl-abc123' # from a DB row back to its decision ``` Severity filtering works because the service prefixes its lines with journald priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default). A foreground `uvicorn` prints them clean, so the same binary is readable either way. **The oneshot units buffer.** `poller.py` and `seed_energy.py` print progress with plain `print()`, and Python block-buffers stdout when it is not a terminal, so their output arrives in one dump at exit rather than progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you want to watch a sweep as it runs. ## A note on the bind address `--host 127.0.0.1` is deliberate. The service holds a billable API key and has **no authentication of its own** — anything that reaches it can spend your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for the repo limits the blast radius on the filesystem, but nothing limits spend. Putting this on a LAN address needs an auth layer first. ## Using an Ollama on another machine The local LLM does the classifying; it does not have to be on the machine you are typing on, and usually the GPU isn't. Most developers already have WireGuard or a VPN back to a home lab, so the normal shape is router and editor on the laptop, Ollama on the workstation. On the **serving** host (the one with the GPU): ```bash sudo mkdir -p /etc/systemd/system/ollama.service.d sudo cp deploy/ollama-over-vpn.conf \ /etc/systemd/system/ollama.service.d/override.conf # set OLLAMA_HOST to that host's own VPN address — `ip -4 -o addr show` sudo nano /etc/systemd/system/ollama.service.d/override.conf sudo systemctl daemon-reload && sudo systemctl restart ollama ``` On the **client** host, in `config.yaml`: ```yaml classifier: base_url: "http://:11434/v1" verification: base_url: "http://:11434" # same host, so `model` can stay null ``` Both must move together. `verification` speaks Ollama's native `/api/chat` and used to derive its URL from the classifier's; it no longer does, so pointing only the classifier across the tunnel leaves the verifier talking to a `localhost` Ollama that may not exist. Config load refuses the combination where `verification.model` is null and the two hosts differ, because that failure is otherwise silent — the verifier 404s, catches it, and records no sample while appearing to be enabled. Bind Ollama to the **VPN address, not `0.0.0.0`**. It has no authentication of any kind: anything that reaches the port can run inference, enumerate your models and pull new ones. Same reasoning as the dispatcher's loopback bind above. A cloud endpoint works too — set `classifier.api_key_env` to the env var holding its key. On a five-prompt comparison NeuralWatt's `deepseek-v4-flash` classified in 1.02s mean against `qwen3.5`'s 11.58s on an RTX 6000. Local inference is not free, it is unbilled. ## Pointing opencode at it The repo-local `opencode.json` sets this up already, so running `opencode` from inside a clone of this repo uses the router by default. To use it from anywhere, merge the `provider.llm-router` block into `~/.config/opencode/opencode.json` and set `"model": "llm-router/auto"`. Two model names: - `llm-router/auto` — normal routing; flex rows excluded, so nothing gets held server-side during peak - `llm-router/auto:batch` — admits flex rows, for overnight/async work `limit.context` is declared as 782324, the largest effective window in the routable catalog. The router hard-filters on the measured conversation size, so a prompt too big for the smaller models simply won't be routed to them; if it fits nothing, `/v1/chat/completions` returns a 422 naming the constraint rather than truncating.