Files
6krrt/tests/test_prefix_probe.py
adlee-was-taken fb511cd78b feat(telemetry): a prefix-stability probe that stores no prefix
Wave 1 item 1.3 of plans/token-waste-waves.md.

The provider bills the longest byte-identical PREFIX of a prompt at the
cached rate, so rewriting an early message re-bills everything after it.
context_prune has two paths and they differ exactly there: the uniform path
compresses a contiguous positional region, so an append cannot disturb it,
while the relevance path compresses a prefix of a relevance-ORDERED list
until a growing target_save is covered, so one more candidate crosses each
turn at an arbitrary message POSITION. Offline that rewrites 75% of a
payload's tokens. Live, the cache rate on pruned turns is 0.924, which is
not what that should look like. This is the instrument that settles which
reading is right, and it is deliberately the only thing in this commit --
no behavior change.

WHAT IS STORED, AND WHY THAT SHAPE. Three integers per decision row:
prefix_divergence_index, prefix_tokens_after_divergence and
prefix_prev_message_count. The third is not redundant and is the reason the
other two can be read at all: a conversation that only grew diverges at
exactly the previous turn's message count, and that is the GOOD case even
though tokens_after is non-zero there. Anything lower is rewritten history.
All three are NULL together when there was no previous turn -- a zero would
read as total cache loss at message 0, which is a measurement nobody made.

The natural shape was a per-message hash array on the row. That was rejected.
This router never stores raw task text anywhere -- local_encoder is zero-shot
for exactly that reason -- and a per-message digest list is also a per-message
LENGTH vector, which is the closest thing to a content side-channel available
here. So the digests live in process memory for exactly one turn, long enough
to compare the next turn against them, and never reach the database. A restart
costs one comparison per session; that is the whole price. The store is
bounded (32 sessions) because an entry is O(messages), unlike session_cache's
fixed-size dataclass, and no digest is ever returned to a caller, so there is
no path by which one gets persisted by accident.

COST, MEASURED, at the live median payload shape (98k tokens in, 74k out,
57 messages): 1.03 ms per turn. 2.70 ms at 336k/177k, 7.41 ms at 1.0M/468k.
Against a request path whose floor is a provider round-trip of 1.4-2.0 s that
is ~0.06%, and json.dumps is the bulk of it, not the hashing. Gated anyway on
pinch.prefix_probe, defaulted ON: a probe that is off measures nothing, and
Wave 3 is waiting on what this says. It is NOT gated on whether pruning
actually fired -- an under-budget turn is the cheapest one to fingerprint and
is the baseline the pruned turns are read against.

Both prune call sites feed it, routed and passthrough, and observe() is
called exactly once per request: it remembers this turn as a side effect, so
a second call would compare a turn against itself and report a perfect prefix
that nothing measured.

Reads tolerate a database that never ran the ALTER. metrics is imported by
admin, which can open one, and the live router.db is exactly that until its
next restart -- so recent_decisions selects the columns only when a PRAGMA
probe finds them and backfills the keys as NULL otherwise, the same shape
6f9f663 used. Confirmed against the live DB read-only: zero probe columns
present, three rows back, three NULL fields, no OperationalError.

Registered in both drift guards, and the schema-drift registry's FULL_ROW
carries a destructive divergence (19 of 82 against 81 prior messages) rather
than a placeholder. ROUTE_DECISIONS_COLUMNS was checked against the live
schema first, per its own history of drifting; it was correct, and gained
three entries.

16 new tests. The load-bearing one replays the exact scenario the direct
investigation used -- ten turns, one tool result appended, a LITERALLY
identical relevance order on both turns so embedding jitter cannot be the
explanation -- and pins both halves through the probe rather than by hand:
relevance diverges at message 19 of 82 with 75% of tokens after it, uniform
diverges only at the appended message with 2%. The asymmetry is asserted as
its own test, because the asymmetry is the finding. Also pinned: nothing
recoverable is retained (no distinctive substring of the payload appears in
the store, every remembered value is a short hex digest or an int), the store
is bounded and evicts oldest-first, an unserializable message never breaks a
dispatch, a shrinking payload reads as destructive, key order is not a
divergence, and the knob off writes three NULLs and fingerprints nothing.

1951 -> 1967 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 11:37:59 -04:00

13 KiB