tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
8.3 KiB
Deploy separation: production runs from its own clone, not the dev checkout
Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run
Date: 2026-09-26
Why: incident #8, and #4, #5 and #7 before it. Production has been one git pull,
one uncommitted edit or one agent session away from the dev checkout.
Current state (measured 2026-09-26)
- Every installed user unit runs from
~/Sources/6krrt, the dev checkout:llm-router.servicellm-router-poller,-seed,-backup,-offsitellm-router-baseline-report
- What they take from it:
WorkingDirectory,PYTHONPATH,EnvironmentFile=.env,.venv/bin/*, andHF_HOME=.hf-cache.ReadWritePathscovers the whole source tree, so the live service can write anywhere in the repo. - Everything else follows the working directory, because the code resolves
paths against it:
config/config.yaml, theconfig/config.local.yamloverlay, anddatabase.path: "router.db". Moving the working directory moves all of it, with no code change. - The repo's own unit templates (
deploy/*.service) already target a separate%h/llm-routerdirectory, except-backupand-offsite. The split was the original design. The installed units drifted onto the dev checkout, and~/llm-routerdoes not exist. - Sizes:
router.db48 MB,.hf-cache2.8 GB,.venv7.2 GB..envholdsNEURALWATT_API_KEYandOPENROUTER_API_KEY.
Target layout
~/.local/share/6krrt/app production clone of the Gitea repo, detached at a
merged origin/main commit. Nobody edits here; no
agent works here; nothing here ever runs `git clean`.
router.db live DB (gitignored; checkout never touches it)
config/config.local.yaml live overlay (gitignored)
.env live keys, mode 600 (gitignored)
.venv/ production venv, built from the clone's requirements
~/.cache/6krrt-hf shared HF model cache (moved once from
~/Sources/6krrt/.hf-cache; production and the dev
sandbox both point HF_HOME here)
~/Sources/6krrt dev checkout: source only. No live DB, no
production .env. The 8081 sandbox uses DB copies.
Units: every llm-router* unit's paths change from %h/Sources/6krrt to
%h/.local/share/6krrt/app, and ReadWritePaths narrows to that directory.
The backup and offsite scripts already take REPO from the environment; set it
to the app directory.
scripts/deploy.sh (to build; not needed for the first cutover)
The only way code reaches production after the cutover:
- Show the change.
git -C $APP fetch origin, thengit log --oneline HEAD..origin/main. - Preflight, without touching the running service:
- load
origin/main's config against the real overlay (catches a #7-style rename) - check the requirements are satisfied in
$APP/.venv import dispatcherfrom a temp export oforigin/main
- load
- Cut over. Record the current SHA,
git checkout --detach origin/main,uv pip installonly if the requirements changed, restart, and poll/healthfor 30 s. - Roll back on failure. Check out the recorded SHA, restart, and alert.
- Record the deployed SHA in
$APP/.deployed, and expose it on/healthif that endpoint gains a field, so "what is running" has one answer.
Cutover runbook
Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run the commands from a shell, not from an agent session.
APP=~/.local/share/6krrt/app
DEV=~/Sources/6krrt
# 0. Build the clone and venv while production keeps running.
mkdir -p ~/.local/share/6krrt
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
git -C "$APP" checkout --detach origin/main
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
-r requirements.txt -r requirements-encoder.txt)
# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf
# 2. Stop production and every timer that touches the DB.
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
systemctl --user stop llm-router.service
# 3. Copy state; originals stay put until step 7.
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok
# 4. Repoint the units (backups first), then reload.
cd ~/.config/systemd/user
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
systemctl --user daemon-reload
# 5. Start and verify.
systemctl --user start llm-router.service
sleep 3; curl -s localhost:8080/health
systemctl --user show llm-router.service -p WorkingDirectory
curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
# 6. Point the dev sandbox at the shared HF cache and the new live DB.
# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")
# 7. Only after a day of healthy running: retire the dev copies.
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"
Rollback, before step 7: restore the unit backups from
~/.local/share/6krrt/unit-backup/, run daemon-reload, start the service.
The dev copies of the DB, overlay and .env were never moved, so production is
back where it started. Any decisions logged after the cutover stay in the
app-directory DB.
Follow-ups that must land with or right after the cutover
CLAUDE.mdNorth Star #3 and the "Measure the liverouter.db" memory name/home/alee/Sources/6krrt/router.dbas the live DB. Update both to~/.local/share/6krrt/app/router.db, or agents will measure a stale file.scripts/sandbox.sh:HF_HOMEgoes to~/.cache/6krrt-hf. Its DB copy instructions change to the app directory's DB.deploy/*.servicetemplates: make the repo's templates say%h/.local/share/6krrt/appso reinstalling from the repo reproduces this layout, and fix-backupand-offsiteto match.deploy/README.md: replace the install section with this layout anddeploy.sh.- The admin portal's "Restart Service" trigger keeps working (it calls
systemctl), but any trigger that runs a script by repo-relative path should be checked against the new working directory.
Dry run
See the result appended below.
Result, 2026-09-26 04:3x, against origin/main 827c408:
- Clone from Gitea and detached checkout: OK.
- Fresh venv,
uv venvplus both requirements files: 2 s (all fromuv's cache);torch 2.9.1+cu128andtransformersimport. Building the venv does not affect downtime, and step 0 can run hours ahead. - State: the DB copied by sqlite backup with the source opened
mode=ro, andintegrity_check= ok; overlay copied;.envreplaced with dummy keys for the dry run. - Served on 8081 from the clone.
/proc/<pid>/cwdconfirmed the clone directory, withHF_HOMEon the dev cache read-only:/health200;/admin/controls200/admin/api/runtimeshowedpinch_enabledpersisted: False, so the copied overlay loaded- the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7 days
/routeconsidered 41 candidates and pickeddeepseek/deepseek-v4-flash
- Stopped by the recorded PID; 8081 free; the dummy
.envdeleted.
Not exercised by the dry run: the systemd unit edits (step 4); the timers
under the new ReadWritePaths; the HF cache move (step 1); and real provider
keys. The unit edit is the step to watch during the cutover. systemctl --user show llm-router.service -p WorkingDirectory in step 5 confirms it took.