Files
6krrt/deploy/README.md
adlee-was-taken c639d60859 feat(feedback): a dry run that projects the fold, and the timer it never had
`POST /outcome` is the only ground truth this router has, and feedback.py is
the one loop with no timer -- poller, seed sweep, backup and offsite all have
one. Live state on 2026-09-15: `proficiency` last written 2026-09-10, with 404
unapplied attributable outcomes and thousands of decisions routed off the stale
scores in between.

The fold is irreversible. add_outcome accumulates into a running mean and
recompute_category re-derives every row in the category from the new peer rate;
neither keeps the pre-fold value anywhere, and verifications.applied_at means a
second run will not redo the work either. So the deliverable is a preview plus
units that ship unstarted, not an automatic fold.

`--dry-run` now projects instead of describing. feedback_preview copies the
database into memory, runs the REAL add_outcome against the copy, and diffs the
two proficiency snapshots. It does not re-implement the empirical-Bayes
conversion -- a second implementation would drift, and a confidently wrong
forecast of an irreversible action is the worst failure available here.

Two kinds of movement come out, and the second is the surprise: `direct` rows
carry new outcomes of their own; `ripple` rows carry none and move anyway,
because the whole category is re-derived against a peer rate the new evidence
just changed. On the live backlog 404 samples across 17 pairs move 17 rows
directly and 363 by ripple, so ripple is digested per category and `--csv`
carries every row.

The units are named and hardened like the poller/seed pair, and take no
EnvironmentFile and no network-online.target because the fold makes no HTTP
request of any kind. The .service has no [Install] section, so it cannot be
enabled on its own -- "not enabled by default" is structural rather than a
README promise. 12h cadence because the fold is exactly additive: ten folds of
five land on the same numbers as one fold of fifty, so cadence caps staleness
and batch size and cannot change where the scores end up.

Three corrections to what was in the tree:

- The timer carried `Persistent=true`, which systemd.timer(5) says "only has an
  effect on timers configured with OnCalendar=". This timer is monotonic, so
  the line bought nothing; OnBootSec is the real catch-up. The sibling poller
  and seed timers carry the same inert line -- noted, not fixed in passing.
- `--csv` without `--dry-run` was accepted and discarded, and the run it was
  silently dropped from is the irreversible one. It is now refused.
- preview() copied the whole 41 MB database into memory just to print coverage,
  then project() copied it again. Coverage only reads, so it takes a mode=ro
  handle instead.

Tests: 2068 -> 2087. The load-bearing one asserts a dry run leaves the database
byte-identical -- proficiency rows, verifications.applied_at, and the file's
sha256 -- while still reporting the 8-sample delta it would apply. Verified
non-vacuous by handing copy_database the real connection and watching it fail.
A second test folds for real onto an identical copy and demands the projection
match every score, which is what stops the preview drifting from the store.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-15 22:34:16 -04:00

239 lines
11 KiB
Markdown

# Deploying the router
The dispatcher runs continuously; the rest are one-shots on timers. The poller
is **not** optional — `freshness.stale_after_days` is 3 and
`freshness.exclude_stale` is true, so a catalog that goes unpolled for three
days marks every row stale and the router stops returning any candidate at
all.
| file | what it does |
|---|---|
| `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` |
| `llm-router-poller.service` | one-shot: `PYTHONPATH=src python -m poller` then `PYTHONPATH=src python -m tier` |
| `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h |
| `llm-router-seed.service` | one-shot: a small `PYTHONPATH=src python -m seed_energy` reference sweep |
| `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep |
| `llm-router-feedback.service` | one-shot: `PYTHONPATH=src python -m feedback` — folds client outcomes into `proficiency`. **Ships not enabled**, see below |
| `llm-router-feedback.timer` | every 12 h once you enable it |
These are **user** units — no root, and they run as you with your own
`$HOME`. The tradeoff is that a user service does not inherit your shell
environment, so the API key has to come from a file.
## Install
```bash
# 1. The key. User units don't see your shell env, so .env is required.
cd /path/to/this/repo
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
# 2. Install, pointing the units at wherever you actually cloned this.
# The shipped units say %h/llm-router; %h is systemd's expansion for your
# home directory, so only the part after it needs changing. Getting this
# wrong fails at start with status=200/CHDIR rather than anything obvious.
REPO=$(pwd)
mkdir -p ~/.config/systemd/user
for u in deploy/llm-router*.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
# llm-router-feedback.timer is deliberately NOT in this line — the first fold
# is irreversible. See "The feedback fold, which ships switched off" below.
systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer
# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten
# by the same sed to your REPO) so the modules under `src/` are importable
# without an editable install. .env is read from the repo root and stays there.
# 3. Survive logout/reboot (user units stop with your session otherwise)
loginctl enable-linger "$USER"
# 4. Check. The health endpoint reports whether cost/eco/proficiency
# actually have data behind them, which is otherwise silent.
curl -s localhost:8080/health | python -m json.tool
systemctl --user list-timers 'llm-router*'
```
## Operating it
```bash
systemctl --user status llm-router.service
journalctl --user -u 'llm-router*' -f # everything, live (quote the glob)
journalctl --user -u llm-router -f -o cat # the request log, message only
journalctl --user -u llm-router -p warning # fallbacks, retries, refusals
journalctl --user -u llm-router-poller.service # catalog refreshes
systemctl --user restart llm-router.service # after editing config/config.yaml
systemctl --user start llm-router-poller.service # force a refresh now
```
`config/config.yaml` is read once at startup, so weight and threshold changes need a
restart. The catalog is read per-request, so a poller run takes effect
immediately.
## The feedback fold, which ships switched off
`POST /outcome` is the only ground truth this router has, and `feedback.py` is
what turns those reports into routing changes. It is the one loop here with no
timer, so until you enable one it runs only when someone remembers to run it —
on this deployment that meant a 404-sample backlog and a `proficiency` table
five days stale while thousands of decisions were routed off it.
The install loop above copies these two units in with the rest. **Neither is
enabled**, and the `.service` has no `[Install]` section at all, so it cannot be
enabled on its own — only the timer can:
```bash
systemctl --user start llm-router-feedback.service # fold once, now
systemctl --user enable --now llm-router-feedback.timer # and every 12 h after
```
That is deliberate, and it is the only unit in this directory treated this way.
**The fold is irreversible.** `add_outcome` accumulates into a running mean and
`recompute_category` re-derives every row in the category from the new peer
rate; neither keeps the pre-fold value anywhere, and `verifications.applied_at`
means a second run will not redo the work either. There is no undo and no
restore short of the backup timer. Turning that on for the first time against
an accumulated backlog is an operator decision, not a default.
So look first. The dry run copies the database into memory, runs the real
`add_outcome` against the copy, and reports the before/after per row. It opens
the source read-only and writes nothing to it:
```bash
PYTHONPATH=src .venv/bin/python -m feedback --dry-run # the live DB, safely
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --csv # per-row, for a spreadsheet
PYTHONPATH=src .venv/bin/python -m feedback --dry-run --db /tmp/copy.db
```
Two kinds of movement come back, and the second is the one that surprises
people. **direct** rows have new outcomes of their own. **ripple** rows have
none and move anyway, because the whole category is re-derived against a peer
rate the new evidence just changed — on this deployment 404 samples across 17
pairs moved 17 rows directly and 363 by ripple. The table digests ripple per
category; `--csv` carries every row.
Cadence is 12 h rather than the poller's 2 h because the fold is exactly
additive: ten folds of five samples land on the same numbers as one fold of
fifty, so cadence cannot change *where* the scores end up, only how stale
routing's inputs get and how large each irreversible batch is. Twelve hours
caps staleness at half a day while keeping a run a reviewable batch.
The unit takes no `EnvironmentFile` and no `network-online.target`: the fold
makes no provider call and no HTTP request of any kind.
### Turning up the logs
`logging.level` in config/config.yaml is the documented setting, but flipping it means
editing a tracked file. For a running service use a drop-in instead:
```bash
systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
systemctl --user restart llm-router
```
`info` gives one `route` and one `dispatch` line per request — category, tier,
model chosen, cost, latency. `debug` adds every candidate that was dropped and
by which filter, plus the ranking with scores. No conversation text is logged
at any level.
Every line carries a trace id, and the `dispatch` line carries the provider's
completion id and the session fingerprint, both of which are columns in
`energy_observations`:
```bash
journalctl --user -u llm-router --grep ' id=r9116d9' # one request, all stages
journalctl --user -u llm-router --grep 'chatcmpl-abc123' # from a DB row back to its decision
```
Severity filtering works because the service prefixes its lines with journald
priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default).
A foreground `uvicorn` prints them clean, so the same binary is readable either
way.
**The oneshot units buffer.** `python -m poller` and `python -m seed_energy` print progress
with plain `print()`, and Python block-buffers stdout when it is not a
terminal, so their output arrives in one dump at exit rather than
progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you
want to watch a sweep as it runs.
## A note on the bind address
`--host 127.0.0.1` is deliberate. The service holds a billable API key and
has **no authentication of its own** — anything that reaches it can spend
your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for
the repo limits the blast radius on the filesystem, but nothing limits spend.
Putting this on a LAN address needs an auth layer first.
If you enable `classifier.mode: local_encoder`, note `HF_HOME` is redirected
into the repo (`%h/llm-router/.hf-cache`) for the same reason — Hugging
Face's default cache lives outside `ReadWritePaths` and the service will
crash-loop trying to download a model into a read-only home directory
otherwise. Run `pip install -r requirements-encoder.txt` before switching to
this mode; the startup check refuses to boot with a clear message if it's
missing, rather than failing opaquely on the first request, but a missing
model still needs the dependency installed first.
## Using an Ollama on another machine
The local LLM does the classifying; it does not have to be on the machine you
are typing on, and usually the GPU isn't. Most developers already have
WireGuard or a VPN back to a home lab, so the normal shape is router and
editor on the laptop, Ollama on the workstation.
On the **serving** host (the one with the GPU):
```bash
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo cp deploy/ollama-over-vpn.conf \
/etc/systemd/system/ollama.service.d/override.conf
# set OLLAMA_HOST to that host's own VPN address — `ip -4 -o addr show`
sudo nano /etc/systemd/system/ollama.service.d/override.conf
sudo systemctl daemon-reload && sudo systemctl restart ollama
```
On the **client** host, in `config/config.yaml`:
```yaml
classifier:
base_url: "http://<vpn-ip>:11434/v1"
verification:
base_url: "http://<vpn-ip>:11434" # same host, so `model` can stay null
```
Both must move together. `verification` speaks Ollama's native `/api/chat`
and used to derive its URL from the classifier's; it no longer does, so
pointing only the classifier across the tunnel leaves the verifier talking to
a `localhost` Ollama that may not exist. Config load refuses the combination
where `verification.model` is null and the two hosts differ, because that
failure is otherwise silent — the verifier 404s, catches it, and records no
sample while appearing to be enabled.
Bind Ollama to the **VPN address, not `0.0.0.0`**. It has no authentication of
any kind: anything that reaches the port can run inference, enumerate your
models and pull new ones. Same reasoning as the dispatcher's loopback bind
above.
A cloud endpoint works too — set `classifier.api_key_env` to the env var
holding its key. On a five-prompt comparison NeuralWatt's `deepseek-v4-flash`
classified in 1.02s mean against `qwen3.5`'s 11.58s on an RTX 6000. Local
inference is not free, it is unbilled.
## Pointing opencode at it
The repo-local `opencode.json` sets this up already, so running `opencode`
from inside a clone of this repo uses the router by default. To use it from
anywhere, merge the `provider.llm-router` block into
`~/.config/opencode/opencode.json` and set `"model": "llm-router/auto"`.
Two model names:
- `llm-router/auto` — normal routing; flex rows excluded, so nothing gets
held server-side during peak
- `llm-router/auto:batch` — admits flex rows, for overnight/async work
`limit.context` is declared as 782324, the largest effective window in the
routable catalog. The router hard-filters on the measured conversation size,
so a prompt too big for the smaller models simply won't be routed to them;
if it fits nothing, `/v1/chat/completions` returns a 422 naming the
constraint rather than truncating.