Files
6krrt/plans/deploy-separation.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

168 lines
8.3 KiB
Markdown

# Deploy separation: production runs from its own clone, not the dev checkout
Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run
Date: 2026-09-26
Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`,
one uncommitted edit or one agent session away from the dev checkout.
## Current state (measured 2026-09-26)
- **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout:
- `llm-router.service`
- `llm-router-poller`, `-seed`, `-backup`, `-offsite`
- `llm-router-baseline-report`
- **What they take from it:** `WorkingDirectory`, `PYTHONPATH`,
`EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`.
`ReadWritePaths` covers the whole source tree, so the live service can write
anywhere in the repo.
- **Everything else follows the working directory,** because the code resolves
paths against it: `config/config.yaml`, the `config/config.local.yaml`
overlay, and `database.path: "router.db"`. Moving the working directory moves
all of it, with no code change.
- **The repo's own unit templates** (`deploy/*.service`) already target a
separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split
was the original design. The installed units drifted onto the dev checkout,
and `~/llm-router` does not exist.
- **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env`
holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`.
## Target layout
```
~/.local/share/6krrt/app production clone of the Gitea repo, detached at a
merged origin/main commit. Nobody edits here; no
agent works here; nothing here ever runs `git clean`.
router.db live DB (gitignored; checkout never touches it)
config/config.local.yaml live overlay (gitignored)
.env live keys, mode 600 (gitignored)
.venv/ production venv, built from the clone's requirements
~/.cache/6krrt-hf shared HF model cache (moved once from
~/Sources/6krrt/.hf-cache; production and the dev
sandbox both point HF_HOME here)
~/Sources/6krrt dev checkout: source only. No live DB, no
production .env. The 8081 sandbox uses DB copies.
```
Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to
`%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory.
The backup and offsite scripts already take `REPO` from the environment; set it
to the app directory.
## `scripts/deploy.sh` (to build; not needed for the first cutover)
The only way code reaches production after the cutover:
1. **Show the change.** `git -C $APP fetch origin`, then
`git log --oneline HEAD..origin/main`.
2. **Preflight, without touching the running service:**
- load `origin/main`'s config against the real overlay (catches a #7-style
rename)
- check the requirements are satisfied in `$APP/.venv`
- `import dispatcher` from a temp export of `origin/main`
3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`,
`uv pip install` only if the requirements changed, restart, and poll
`/health` for 30 s.
4. **Roll back on failure.** Check out the recorded SHA, restart, and alert.
5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health`
if that endpoint gains a field, so "what is running" has one answer.
## Cutover runbook
Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run
the commands from a shell, not from an agent session.
```sh
APP=~/.local/share/6krrt/app
DEV=~/Sources/6krrt
# 0. Build the clone and venv while production keeps running.
mkdir -p ~/.local/share/6krrt
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
git -C "$APP" checkout --detach origin/main
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
-r requirements.txt -r requirements-encoder.txt)
# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf
# 2. Stop production and every timer that touches the DB.
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
systemctl --user stop llm-router.service
# 3. Copy state; originals stay put until step 7.
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok
# 4. Repoint the units (backups first), then reload.
cd ~/.config/systemd/user
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
systemctl --user daemon-reload
# 5. Start and verify.
systemctl --user start llm-router.service
sleep 3; curl -s localhost:8080/health
systemctl --user show llm-router.service -p WorkingDirectory
curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
# 6. Point the dev sandbox at the shared HF cache and the new live DB.
# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")
# 7. Only after a day of healthy running: retire the dev copies.
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"
```
**Rollback, before step 7:** restore the unit backups from
`~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service.
The dev copies of the DB, overlay and `.env` were never moved, so production is
back where it started. Any decisions logged after the cutover stay in the
app-directory DB.
## Follow-ups that must land with or right after the cutover
- **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory
name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to
`~/.local/share/6krrt/app/router.db`, or agents will measure a stale file.
- **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy
instructions change to the app directory's DB.
- **`deploy/*.service` templates:** make the repo's templates say
`%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this
layout, and fix `-backup` and `-offsite` to match.
- **`deploy/README.md`:** replace the install section with this layout and
`deploy.sh`.
- **The admin portal's "Restart Service" trigger** keeps working (it calls
`systemctl`), but any trigger that runs a script by repo-relative path should
be checked against the new working directory.
## Dry run
See the result appended below.
**Result, 2026-09-26 04:3x, against `origin/main` `827c408`:**
- **Clone from Gitea and detached checkout:** OK.
- **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s
cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv
does not affect downtime, and step 0 can run hours ahead.
- **State:** the DB copied by sqlite backup with the source opened `mode=ro`,
and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys
for the dry run.
- **Served on 8081** from the clone. `/proc/<pid>/cwd` confirmed the clone
directory, with `HF_HOME` on the dev cache read-only:
- `/health` 200; `/admin/controls` 200
- `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the
copied overlay loaded
- the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7
days
- `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash`
- **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted.
**Not exercised by the dry run:** the systemd unit edits (step 4); the timers
under the new `ReadWritePaths`; the HF cache move (step 1); and real provider
keys. The unit edit is the step to watch during the cutover. `systemctl --user
show llm-router.service -p WorkingDirectory` in step 5 confirms it took.