Two incidents from the last day's deploy cycle, same family as #4 -- a file
the live service depends on was wrong in a way no tracked diff can show.
#6: PR #89/#90 shipped admin knobs (frontend + backend together, correctly),
but the running process only picks up the backend half on restart while the
frontend half is served fresh every request -- so a fresh page load against
an unrestarted process silently omitted the new controls, with /health
staying green throughout.
#7: the session_cache.staleness_minutes -> staleness_seconds rename updated
every tracked consumer in the same change, per this project's no-alias
convention, but config/config.local.yaml (gitignored, machine-local) still
held the old key. StrictModel's extra_forbidden rejected it on every startup
attempt, crash-looping production for ~7 restart cycles until the overlay was
migrated by hand.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0161UAZyxtvmKaVzasJsnqJc