Files
6krrt/plans/deploy-separation.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

8.3 KiB

Deploy separation: production runs from its own clone, not the dev checkout

Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run Date: 2026-09-26 Why: incident #8, and #4, #5 and #7 before it. Production has been one git pull, one uncommitted edit or one agent session away from the dev checkout.

Current state (measured 2026-09-26)

  • Every installed user unit runs from ~/Sources/6krrt, the dev checkout:
    • llm-router.service
    • llm-router-poller, -seed, -backup, -offsite
    • llm-router-baseline-report
  • What they take from it: WorkingDirectory, PYTHONPATH, EnvironmentFile=.env, .venv/bin/*, and HF_HOME=.hf-cache. ReadWritePaths covers the whole source tree, so the live service can write anywhere in the repo.
  • Everything else follows the working directory, because the code resolves paths against it: config/config.yaml, the config/config.local.yaml overlay, and database.path: "router.db". Moving the working directory moves all of it, with no code change.
  • The repo's own unit templates (deploy/*.service) already target a separate %h/llm-router directory, except -backup and -offsite. The split was the original design. The installed units drifted onto the dev checkout, and ~/llm-router does not exist.
  • Sizes: router.db 48 MB, .hf-cache 2.8 GB, .venv 7.2 GB. .env holds NEURALWATT_API_KEY and OPENROUTER_API_KEY.

Target layout

~/.local/share/6krrt/app      production clone of the Gitea repo, detached at a
                              merged origin/main commit. Nobody edits here; no
                              agent works here; nothing here ever runs `git clean`.
  router.db                   live DB (gitignored; checkout never touches it)
  config/config.local.yaml    live overlay (gitignored)
  .env                        live keys, mode 600 (gitignored)
  .venv/                      production venv, built from the clone's requirements
~/.cache/6krrt-hf             shared HF model cache (moved once from
                              ~/Sources/6krrt/.hf-cache; production and the dev
                              sandbox both point HF_HOME here)
~/Sources/6krrt               dev checkout: source only. No live DB, no
                              production .env. The 8081 sandbox uses DB copies.

Units: every llm-router* unit's paths change from %h/Sources/6krrt to %h/.local/share/6krrt/app, and ReadWritePaths narrows to that directory. The backup and offsite scripts already take REPO from the environment; set it to the app directory.

scripts/deploy.sh (to build; not needed for the first cutover)

The only way code reaches production after the cutover:

  1. Show the change. git -C $APP fetch origin, then git log --oneline HEAD..origin/main.
  2. Preflight, without touching the running service:
    • load origin/main's config against the real overlay (catches a #7-style rename)
    • check the requirements are satisfied in $APP/.venv
    • import dispatcher from a temp export of origin/main
  3. Cut over. Record the current SHA, git checkout --detach origin/main, uv pip install only if the requirements changed, restart, and poll /health for 30 s.
  4. Roll back on failure. Check out the recorded SHA, restart, and alert.
  5. Record the deployed SHA in $APP/.deployed, and expose it on /health if that endpoint gains a field, so "what is running" has one answer.

Cutover runbook

Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run the commands from a shell, not from an agent session.

APP=~/.local/share/6krrt/app
DEV=~/Sources/6krrt

# 0. Build the clone and venv while production keeps running.
mkdir -p ~/.local/share/6krrt
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
git -C "$APP" checkout --detach origin/main
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
   -r requirements.txt -r requirements-encoder.txt)

# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf

# 2. Stop production and every timer that touches the DB.
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
  llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
systemctl --user stop llm-router.service

# 3. Copy state; originals stay put until step 7.
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;'     # must print: ok

# 4. Repoint the units (backups first), then reload.
cd ~/.config/systemd/user
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
systemctl --user daemon-reload

# 5. Start and verify.
systemctl --user start llm-router.service
sleep 3; curl -s localhost:8080/health
systemctl --user show llm-router.service -p WorkingDirectory
curl -s localhost:8080/admin/api/runtime | head -c 200     # pinch persisted false = overlay loaded
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
  llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer

# 6. Point the dev sandbox at the shared HF cache and the new live DB.
#    (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")

# 7. Only after a day of healthy running: retire the dev copies.
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"

Rollback, before step 7: restore the unit backups from ~/.local/share/6krrt/unit-backup/, run daemon-reload, start the service. The dev copies of the DB, overlay and .env were never moved, so production is back where it started. Any decisions logged after the cutover stay in the app-directory DB.

Follow-ups that must land with or right after the cutover

  • CLAUDE.md North Star #3 and the "Measure the live router.db" memory name /home/alee/Sources/6krrt/router.db as the live DB. Update both to ~/.local/share/6krrt/app/router.db, or agents will measure a stale file.
  • scripts/sandbox.sh: HF_HOME goes to ~/.cache/6krrt-hf. Its DB copy instructions change to the app directory's DB.
  • deploy/*.service templates: make the repo's templates say %h/.local/share/6krrt/app so reinstalling from the repo reproduces this layout, and fix -backup and -offsite to match.
  • deploy/README.md: replace the install section with this layout and deploy.sh.
  • The admin portal's "Restart Service" trigger keeps working (it calls systemctl), but any trigger that runs a script by repo-relative path should be checked against the new working directory.

Dry run

See the result appended below.

Result, 2026-09-26 04:3x, against origin/main 827c408:

  • Clone from Gitea and detached checkout: OK.
  • Fresh venv, uv venv plus both requirements files: 2 s (all from uv's cache); torch 2.9.1+cu128 and transformers import. Building the venv does not affect downtime, and step 0 can run hours ahead.
  • State: the DB copied by sqlite backup with the source opened mode=ro, and integrity_check = ok; overlay copied; .env replaced with dummy keys for the dry run.
  • Served on 8081 from the clone. /proc/<pid>/cwd confirmed the clone directory, with HF_HOME on the dev cache read-only:
    • /health 200; /admin/controls 200
    • /admin/api/runtime showed pinch_enabled persisted: False, so the copied overlay loaded
    • the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7 days
    • /route considered 41 candidates and picked deepseek/deepseek-v4-flash
  • Stopped by the recorded PID; 8081 free; the dummy .env deleted.

Not exercised by the dry run: the systemd unit edits (step 4); the timers under the new ReadWritePaths; the HF cache move (step 1); and real provider keys. The unit edit is the step to watch during the cutover. systemctl --user show llm-router.service -p WorkingDirectory in step 5 confirms it took.