Files
6krrt/deploy/llm-router-feedback.timer
adlee-was-taken c639d60859 feat(feedback): a dry run that projects the fold, and the timer it never had
`POST /outcome` is the only ground truth this router has, and feedback.py is
the one loop with no timer -- poller, seed sweep, backup and offsite all have
one. Live state on 2026-09-15: `proficiency` last written 2026-09-10, with 404
unapplied attributable outcomes and thousands of decisions routed off the stale
scores in between.

The fold is irreversible. add_outcome accumulates into a running mean and
recompute_category re-derives every row in the category from the new peer rate;
neither keeps the pre-fold value anywhere, and verifications.applied_at means a
second run will not redo the work either. So the deliverable is a preview plus
units that ship unstarted, not an automatic fold.

`--dry-run` now projects instead of describing. feedback_preview copies the
database into memory, runs the REAL add_outcome against the copy, and diffs the
two proficiency snapshots. It does not re-implement the empirical-Bayes
conversion -- a second implementation would drift, and a confidently wrong
forecast of an irreversible action is the worst failure available here.

Two kinds of movement come out, and the second is the surprise: `direct` rows
carry new outcomes of their own; `ripple` rows carry none and move anyway,
because the whole category is re-derived against a peer rate the new evidence
just changed. On the live backlog 404 samples across 17 pairs move 17 rows
directly and 363 by ripple, so ripple is digested per category and `--csv`
carries every row.

The units are named and hardened like the poller/seed pair, and take no
EnvironmentFile and no network-online.target because the fold makes no HTTP
request of any kind. The .service has no [Install] section, so it cannot be
enabled on its own -- "not enabled by default" is structural rather than a
README promise. 12h cadence because the fold is exactly additive: ten folds of
five land on the same numbers as one fold of fifty, so cadence caps staleness
and batch size and cannot change where the scores end up.

Three corrections to what was in the tree:

- The timer carried `Persistent=true`, which systemd.timer(5) says "only has an
  effect on timers configured with OnCalendar=". This timer is monotonic, so
  the line bought nothing; OnBootSec is the real catch-up. The sibling poller
  and seed timers carry the same inert line -- noted, not fixed in passing.
- `--csv` without `--dry-run` was accepted and discarded, and the run it was
  silently dropped from is the irreversible one. It is now refused.
- preview() copied the whole 41 MB database into memory just to print coverage,
  then project() copied it again. Coverage only reads, so it takes a mode=ro
  handle instead.

Tests: 2068 -> 2087. The load-bearing one asserts a dry run leaves the database
byte-identical -- proficiency rows, verifications.applied_at, and the file's
sha256 -- while still reporting the 8-sample delta it would apply. Verified
non-vacuous by handing copy_database the real connection and watching it fail.
A second test folds for real onto an identical copy and demands the projection
match every score, which is what stops the preview drifting from the store.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-15 22:34:16 -04:00

1.5 KiB