tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
168 lines
8.3 KiB
Markdown
168 lines
8.3 KiB
Markdown
# Deploy separation: production runs from its own clone, not the dev checkout
|
|
|
|
Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run
|
|
Date: 2026-09-26
|
|
Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`,
|
|
one uncommitted edit or one agent session away from the dev checkout.
|
|
|
|
## Current state (measured 2026-09-26)
|
|
|
|
- **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout:
|
|
- `llm-router.service`
|
|
- `llm-router-poller`, `-seed`, `-backup`, `-offsite`
|
|
- `llm-router-baseline-report`
|
|
- **What they take from it:** `WorkingDirectory`, `PYTHONPATH`,
|
|
`EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`.
|
|
`ReadWritePaths` covers the whole source tree, so the live service can write
|
|
anywhere in the repo.
|
|
- **Everything else follows the working directory,** because the code resolves
|
|
paths against it: `config/config.yaml`, the `config/config.local.yaml`
|
|
overlay, and `database.path: "router.db"`. Moving the working directory moves
|
|
all of it, with no code change.
|
|
- **The repo's own unit templates** (`deploy/*.service`) already target a
|
|
separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split
|
|
was the original design. The installed units drifted onto the dev checkout,
|
|
and `~/llm-router` does not exist.
|
|
- **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env`
|
|
holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`.
|
|
|
|
## Target layout
|
|
|
|
```
|
|
~/.local/share/6krrt/app production clone of the Gitea repo, detached at a
|
|
merged origin/main commit. Nobody edits here; no
|
|
agent works here; nothing here ever runs `git clean`.
|
|
router.db live DB (gitignored; checkout never touches it)
|
|
config/config.local.yaml live overlay (gitignored)
|
|
.env live keys, mode 600 (gitignored)
|
|
.venv/ production venv, built from the clone's requirements
|
|
~/.cache/6krrt-hf shared HF model cache (moved once from
|
|
~/Sources/6krrt/.hf-cache; production and the dev
|
|
sandbox both point HF_HOME here)
|
|
~/Sources/6krrt dev checkout: source only. No live DB, no
|
|
production .env. The 8081 sandbox uses DB copies.
|
|
```
|
|
|
|
Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to
|
|
`%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory.
|
|
The backup and offsite scripts already take `REPO` from the environment; set it
|
|
to the app directory.
|
|
|
|
## `scripts/deploy.sh` (to build; not needed for the first cutover)
|
|
|
|
The only way code reaches production after the cutover:
|
|
1. **Show the change.** `git -C $APP fetch origin`, then
|
|
`git log --oneline HEAD..origin/main`.
|
|
2. **Preflight, without touching the running service:**
|
|
- load `origin/main`'s config against the real overlay (catches a #7-style
|
|
rename)
|
|
- check the requirements are satisfied in `$APP/.venv`
|
|
- `import dispatcher` from a temp export of `origin/main`
|
|
3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`,
|
|
`uv pip install` only if the requirements changed, restart, and poll
|
|
`/health` for 30 s.
|
|
4. **Roll back on failure.** Check out the recorded SHA, restart, and alert.
|
|
5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health`
|
|
if that endpoint gains a field, so "what is running" has one answer.
|
|
|
|
## Cutover runbook
|
|
|
|
Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run
|
|
the commands from a shell, not from an agent session.
|
|
|
|
```sh
|
|
APP=~/.local/share/6krrt/app
|
|
DEV=~/Sources/6krrt
|
|
|
|
# 0. Build the clone and venv while production keeps running.
|
|
mkdir -p ~/.local/share/6krrt
|
|
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
|
|
git -C "$APP" checkout --detach origin/main
|
|
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
|
|
-r requirements.txt -r requirements-encoder.txt)
|
|
|
|
# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
|
|
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf
|
|
|
|
# 2. Stop production and every timer that touches the DB.
|
|
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
|
|
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
|
|
systemctl --user stop llm-router.service
|
|
|
|
# 3. Copy state; originals stay put until step 7.
|
|
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
|
|
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
|
|
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
|
|
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok
|
|
|
|
# 4. Repoint the units (backups first), then reload.
|
|
cd ~/.config/systemd/user
|
|
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
|
|
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
|
|
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
|
|
systemctl --user daemon-reload
|
|
|
|
# 5. Start and verify.
|
|
systemctl --user start llm-router.service
|
|
sleep 3; curl -s localhost:8080/health
|
|
systemctl --user show llm-router.service -p WorkingDirectory
|
|
curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded
|
|
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
|
|
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
|
|
|
|
# 6. Point the dev sandbox at the shared HF cache and the new live DB.
|
|
# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")
|
|
|
|
# 7. Only after a day of healthy running: retire the dev copies.
|
|
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"
|
|
```
|
|
|
|
**Rollback, before step 7:** restore the unit backups from
|
|
`~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service.
|
|
The dev copies of the DB, overlay and `.env` were never moved, so production is
|
|
back where it started. Any decisions logged after the cutover stay in the
|
|
app-directory DB.
|
|
|
|
## Follow-ups that must land with or right after the cutover
|
|
|
|
- **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory
|
|
name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to
|
|
`~/.local/share/6krrt/app/router.db`, or agents will measure a stale file.
|
|
- **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy
|
|
instructions change to the app directory's DB.
|
|
- **`deploy/*.service` templates:** make the repo's templates say
|
|
`%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this
|
|
layout, and fix `-backup` and `-offsite` to match.
|
|
- **`deploy/README.md`:** replace the install section with this layout and
|
|
`deploy.sh`.
|
|
- **The admin portal's "Restart Service" trigger** keeps working (it calls
|
|
`systemctl`), but any trigger that runs a script by repo-relative path should
|
|
be checked against the new working directory.
|
|
|
|
## Dry run
|
|
|
|
See the result appended below.
|
|
|
|
**Result, 2026-09-26 04:3x, against `origin/main` `827c408`:**
|
|
- **Clone from Gitea and detached checkout:** OK.
|
|
- **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s
|
|
cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv
|
|
does not affect downtime, and step 0 can run hours ahead.
|
|
- **State:** the DB copied by sqlite backup with the source opened `mode=ro`,
|
|
and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys
|
|
for the dry run.
|
|
- **Served on 8081** from the clone. `/proc/<pid>/cwd` confirmed the clone
|
|
directory, with `HF_HOME` on the dev cache read-only:
|
|
- `/health` 200; `/admin/controls` 200
|
|
- `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the
|
|
copied overlay loaded
|
|
- the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7
|
|
days
|
|
- `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash`
|
|
- **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted.
|
|
|
|
**Not exercised by the dry run:** the systemd unit edits (step 4); the timers
|
|
under the new `ReadWritePaths`; the HF cache move (step 1); and real provider
|
|
keys. The unit edit is the step to watch during the cutover. `systemctl --user
|
|
show llm-router.service -p WorkingDirectory` in step 5 confirms it took.
|