# Deploy separation: production runs from its own clone, not the dev checkout Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run Date: 2026-09-26 Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`, one uncommitted edit or one agent session away from the dev checkout. ## Current state (measured 2026-09-26) - **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout: - `llm-router.service` - `llm-router-poller`, `-seed`, `-backup`, `-offsite` - `llm-router-baseline-report` - **What they take from it:** `WorkingDirectory`, `PYTHONPATH`, `EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`. `ReadWritePaths` covers the whole source tree, so the live service can write anywhere in the repo. - **Everything else follows the working directory,** because the code resolves paths against it: `config/config.yaml`, the `config/config.local.yaml` overlay, and `database.path: "router.db"`. Moving the working directory moves all of it, with no code change. - **The repo's own unit templates** (`deploy/*.service`) already target a separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split was the original design. The installed units drifted onto the dev checkout, and `~/llm-router` does not exist. - **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env` holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`. ## Target layout ``` ~/.local/share/6krrt/app production clone of the Gitea repo, detached at a merged origin/main commit. Nobody edits here; no agent works here; nothing here ever runs `git clean`. router.db live DB (gitignored; checkout never touches it) config/config.local.yaml live overlay (gitignored) .env live keys, mode 600 (gitignored) .venv/ production venv, built from the clone's requirements ~/.cache/6krrt-hf shared HF model cache (moved once from ~/Sources/6krrt/.hf-cache; production and the dev sandbox both point HF_HOME here) ~/Sources/6krrt dev checkout: source only. No live DB, no production .env. The 8081 sandbox uses DB copies. ``` Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to `%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory. The backup and offsite scripts already take `REPO` from the environment; set it to the app directory. ## `scripts/deploy.sh` (to build; not needed for the first cutover) The only way code reaches production after the cutover: 1. **Show the change.** `git -C $APP fetch origin`, then `git log --oneline HEAD..origin/main`. 2. **Preflight, without touching the running service:** - load `origin/main`'s config against the real overlay (catches a #7-style rename) - check the requirements are satisfied in `$APP/.venv` - `import dispatcher` from a temp export of `origin/main` 3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`, `uv pip install` only if the requirements changed, restart, and poll `/health` for 30 s. 4. **Roll back on failure.** Check out the recorded SHA, restart, and alert. 5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health` if that endpoint gains a field, so "what is running" has one answer. ## Cutover runbook Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run the commands from a shell, not from an agent session. ```sh APP=~/.local/share/6krrt/app DEV=~/Sources/6krrt # 0. Build the clone and venv while production keeps running. mkdir -p ~/.local/share/6krrt git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP" git -C "$APP" checkout --detach origin/main (cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \ -r requirements.txt -r requirements-encoder.txt) # 1. Share the HF cache (one move; the dev sandbox follows in step 6). mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf # 2. Stop production and every timer that touches the DB. systemctl --user stop llm-router-poller.timer llm-router-seed.timer \ llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer systemctl --user stop llm-router.service # 3. Copy state; originals stay put until step 7. sqlite3 "$DEV/router.db" ".backup '$APP/router.db'" command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml" command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env" sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok # 4. Repoint the units (backups first), then reload. cd ~/.config/systemd/user mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service systemctl --user daemon-reload # 5. Start and verify. systemctl --user start llm-router.service sleep 3; curl -s localhost:8080/health systemctl --user show llm-router.service -p WorkingDirectory curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded systemctl --user start llm-router-poller.timer llm-router-seed.timer \ llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer # 6. Point the dev sandbox at the shared HF cache and the new live DB. # (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups") # 7. Only after a day of healthy running: retire the dev copies. mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)" ``` **Rollback, before step 7:** restore the unit backups from `~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service. The dev copies of the DB, overlay and `.env` were never moved, so production is back where it started. Any decisions logged after the cutover stay in the app-directory DB. ## Follow-ups that must land with or right after the cutover - **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to `~/.local/share/6krrt/app/router.db`, or agents will measure a stale file. - **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy instructions change to the app directory's DB. - **`deploy/*.service` templates:** make the repo's templates say `%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this layout, and fix `-backup` and `-offsite` to match. - **`deploy/README.md`:** replace the install section with this layout and `deploy.sh`. - **The admin portal's "Restart Service" trigger** keeps working (it calls `systemctl`), but any trigger that runs a script by repo-relative path should be checked against the new working directory. ## Dry run See the result appended below. **Result, 2026-09-26 04:3x, against `origin/main` `827c408`:** - **Clone from Gitea and detached checkout:** OK. - **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv does not affect downtime, and step 0 can run hours ahead. - **State:** the DB copied by sqlite backup with the source opened `mode=ro`, and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys for the dry run. - **Served on 8081** from the clone. `/proc//cwd` confirmed the clone directory, with `HF_HOME` on the dev cache read-only: - `/health` 200; `/admin/controls` 200 - `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the copied overlay loaded - the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7 days - `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash` - **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted. **Not exercised by the dry run:** the systemd unit edits (step 4); the timers under the new `ReadWritePaths`; the HF cache move (step 1); and real provider keys. The unit edit is the step to watch during the cutover. `systemctl --user show llm-router.service -p WorkingDirectory` in step 5 confirms it took.