Files
6krrt/tests/test_tui_schema_drift.py
adlee-was-taken fb511cd78b feat(telemetry): a prefix-stability probe that stores no prefix
Wave 1 item 1.3 of plans/token-waste-waves.md.

The provider bills the longest byte-identical PREFIX of a prompt at the
cached rate, so rewriting an early message re-bills everything after it.
context_prune has two paths and they differ exactly there: the uniform path
compresses a contiguous positional region, so an append cannot disturb it,
while the relevance path compresses a prefix of a relevance-ORDERED list
until a growing target_save is covered, so one more candidate crosses each
turn at an arbitrary message POSITION. Offline that rewrites 75% of a
payload's tokens. Live, the cache rate on pruned turns is 0.924, which is
not what that should look like. This is the instrument that settles which
reading is right, and it is deliberately the only thing in this commit --
no behavior change.

WHAT IS STORED, AND WHY THAT SHAPE. Three integers per decision row:
prefix_divergence_index, prefix_tokens_after_divergence and
prefix_prev_message_count. The third is not redundant and is the reason the
other two can be read at all: a conversation that only grew diverges at
exactly the previous turn's message count, and that is the GOOD case even
though tokens_after is non-zero there. Anything lower is rewritten history.
All three are NULL together when there was no previous turn -- a zero would
read as total cache loss at message 0, which is a measurement nobody made.

The natural shape was a per-message hash array on the row. That was rejected.
This router never stores raw task text anywhere -- local_encoder is zero-shot
for exactly that reason -- and a per-message digest list is also a per-message
LENGTH vector, which is the closest thing to a content side-channel available
here. So the digests live in process memory for exactly one turn, long enough
to compare the next turn against them, and never reach the database. A restart
costs one comparison per session; that is the whole price. The store is
bounded (32 sessions) because an entry is O(messages), unlike session_cache's
fixed-size dataclass, and no digest is ever returned to a caller, so there is
no path by which one gets persisted by accident.

COST, MEASURED, at the live median payload shape (98k tokens in, 74k out,
57 messages): 1.03 ms per turn. 2.70 ms at 336k/177k, 7.41 ms at 1.0M/468k.
Against a request path whose floor is a provider round-trip of 1.4-2.0 s that
is ~0.06%, and json.dumps is the bulk of it, not the hashing. Gated anyway on
pinch.prefix_probe, defaulted ON: a probe that is off measures nothing, and
Wave 3 is waiting on what this says. It is NOT gated on whether pruning
actually fired -- an under-budget turn is the cheapest one to fingerprint and
is the baseline the pruned turns are read against.

Both prune call sites feed it, routed and passthrough, and observe() is
called exactly once per request: it remembers this turn as a side effect, so
a second call would compare a turn against itself and report a perfect prefix
that nothing measured.

Reads tolerate a database that never ran the ALTER. metrics is imported by
admin, which can open one, and the live router.db is exactly that until its
next restart -- so recent_decisions selects the columns only when a PRAGMA
probe finds them and backfills the keys as NULL otherwise, the same shape
6f9f663 used. Confirmed against the live DB read-only: zero probe columns
present, three rows back, three NULL fields, no OperationalError.

Registered in both drift guards, and the schema-drift registry's FULL_ROW
carries a destructive divergence (19 of 82 against 81 prior messages) rather
than a placeholder. ROUTE_DECISIONS_COLUMNS was checked against the live
schema first, per its own history of drifting; it was correct, and gained
three entries.

16 new tests. The load-bearing one replays the exact scenario the direct
investigation used -- ten turns, one tool result appended, a LITERALLY
identical relevance order on both turns so embedding jitter cannot be the
explanation -- and pins both halves through the probe rather than by hand:
relevance diverges at message 19 of 82 with 75% of tokens after it, uniform
diverges only at the appended message with 2%. The asymmetry is asserted as
its own test, because the asymmetry is the finding. Also pinned: nothing
recoverable is retained (no distinctive substring of the payload appears in
the store, every remembered value is a short hex digest or an int), the store
is bounded and evicts oldest-first, an unserializable message never breaks a
dispatch, a shrinking payload reads as destructive, key order is not a
divergence, and the knob off writes three NULLs and fingerprints nothing.

1951 -> 1967 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 11:37:59 -04:00

212 lines
8.5 KiB
Python

"""Schema-drift guard: every route_decisions column reaches the TUI on purpose.
Five-plus columns shipped to ``route_decisions`` across recent merges without
any of them reaching the dashboard, and ``tests/test_route_decisions.py``'s own
hand-maintained column list had ALREADY drifted (missing ``request_id``) — the
drift recurs, so this file is the enforcement, not the catch-up.
Two registries, not a magic diff:
- ``SCHEMA_TO_MODEL_KEYS`` — schema column -> the ``tui_model.decision_row``
key that surfaces it. This is the recorded DECISION about surfacing. Three
keys are intentionally renamed (``task_category`` -> ``category``,
``task_tier`` -> ``tier``, ``selected_model`` -> ``selected``), which is
exactly why the mapping is explicit and never derived by string identity.
- ``UNSURFACED_COLUMNS`` — columns deliberately NOT surfaced, each with a
one-line reason. Empty today.
Every new schema column must land in one of the two registries or
``test_schema_columns_all_have_surfacing_decisions`` fails naming the column;
every registry entry must actually round-trip through ``decision_row`` and the
``/metrics`` SELECT or the other two tests fail naming the key.
Offline: temp SQLite created from ``config/schema.sql``, never the live
router.db (pattern borrowed from tests/test_route_decisions.py).
"""
import sqlite3
from pathlib import Path
import metrics
import tui_model
ROOT = Path(__file__).resolve().parent.parent
SCHEMA_SQL = (ROOT / "config" / "schema.sql").read_text()
# route_decisions column -> the tui_model.decision_row key that surfaces it.
# Schema order (config/schema.sql route_decisions declaration). Add a new
# schema column here when the TUI should surface it, or to
# UNSURFACED_COLUMNS when it deliberately should not.
SCHEMA_TO_MODEL_KEYS = {
"id": "id",
"observed_at": "observed_at",
"kind": "kind",
"task_category": "category",
"task_tier": "tier",
"required_context_tokens": "required_context_tokens",
"confidence": "confidence",
"classifier_ms": "classifier_ms",
"classification_source": "classification_source",
"latency_tolerance": "latency_tolerance",
"candidates_considered": "candidates_considered",
"selected_model": "selected",
"selected_provider": "selected_provider",
"runner_up_models": "runner_up_models",
"est_cost_usd": "est_cost_usd",
"est_proficiency": "est_proficiency",
"rejected_reason": "rejected_reason",
"session_key": "session_key",
"tools": "tools",
"images": "images",
"json_mode": "json_mode",
"streamed": "streamed",
"flex_preference": "flex_preference",
"flex_swapped": "flex_swapped",
"flex_forced": "flex_forced",
"request_id": "request_id",
"exploration": "exploration",
"pinch_original_tokens": "pinch_original_tokens",
"pinch_final_tokens": "pinch_final_tokens",
"profile": "profile",
# Prefix-stability probe (pinch.prefix_probe). Surfaced through the detail
# popup, which renders the whole decision_row dict, rather than as new
# table columns: the three numbers only mean anything read together.
"prefix_divergence_index": "prefix_divergence_index",
"prefix_tokens_after_divergence": "prefix_tokens_after_divergence",
"prefix_prev_message_count": "prefix_prev_message_count",
}
# Columns deliberately NOT surfaced anywhere in the TUI get recorded here with a
# one-line reason (callers must keep the comment). Empty today: every column is
# surfaced via decision_row. A new schema column that lands in NEITHER registry
# fails the drift test.
UNSURFACED_COLUMNS: set[str] = set()
# Every column seeded non-NULL: the single source of truth for the INSERT and
# the round-trip assertions, so the tests prove the column flows, not that
# None flows.
FULL_ROW = {
"id": 1,
"observed_at": "2026-09-05T12:34:56+00:00",
"kind": "chat",
"task_category": "coding_general",
"task_tier": 2,
"required_context_tokens": 123456,
"confidence": 0.91,
"classifier_ms": 1800,
"classification_source": "classifier",
"latency_tolerance": "interactive",
"candidates_considered": 7,
"selected_model": "fixture-model",
"selected_provider": "neuralwatt",
"runner_up_models": '[{"model_id": "runner", "provider": "neuralwatt"}]',
"est_cost_usd": 0.00123,
"est_proficiency": 0.88,
# Schema-legal even for a selected row (no CHECK constraint); non-NULL so the
# round-trip assertion proves the column flows, not that None flows.
"rejected_reason": "fixture rejected reason",
"session_key": "fixture-session-key",
"tools": 0,
"images": 0,
"json_mode": 1,
"streamed": 1,
"flex_preference": "auto",
"flex_swapped": 0,
"flex_forced": 0,
"request_id": "chatcmpl-fixture-42",
"exploration": 1,
"pinch_original_tokens": 120000,
"pinch_final_tokens": 96122,
"profile": "default",
# A destructive divergence: index 19 sits below the 81 messages the
# previous turn had, so the payload was rewritten rather than appended to.
"prefix_divergence_index": 19,
"prefix_tokens_after_divergence": 72091,
"prefix_prev_message_count": 81,
}
def _db_from_schema(tmp_path, row_factory=None) -> sqlite3.Connection:
"""A throwaway DB from config/schema.sql — never the live router.db.
``row_factory`` defaults to None (plain tuples) like the PRAGMA test
needs; the metrics round-trip test passes ``sqlite3.Row`` because
``metrics.recent_decisions`` builds its rows via ``dict(row)``.
"""
conn = sqlite3.connect(tmp_path / "schema-drift.db")
conn.executescript(SCHEMA_SQL)
conn.row_factory = row_factory
return conn
def test_schema_columns_all_have_surfacing_decisions(tmp_path):
"""Every route_decisions column is either surfaced or consciously not."""
conn = _db_from_schema(tmp_path)
schema_cols = {
row[1] for row in conn.execute("PRAGMA table_info(route_decisions)")
}
conn.close()
registered = set(SCHEMA_TO_MODEL_KEYS) | UNSURFACED_COLUMNS
missing = schema_cols - set(SCHEMA_TO_MODEL_KEYS) - UNSURFACED_COLUMNS
extra = registered - schema_cols
assert not missing, (
f"route_decisions column(s) added without a surfacing decision: "
f"{sorted(missing)}. Add each to SCHEMA_TO_MODEL_KEYS if the TUI "
f"should surface it (project it in tui_model.decision_row), or to "
f"UNSURFACED_COLUMNS with a reason comment."
)
assert not extra, (
f"registry registers column(s) that no longer exist in "
f"route_decisions: {sorted(extra)}."
)
def test_decision_row_projects_every_tracked_column():
"""decision_row surfaces every registered column's value faithfully."""
projected = tui_model.decision_row(dict(FULL_ROW))
expected_keys = set(SCHEMA_TO_MODEL_KEYS.values())
assert set(projected) == expected_keys, (
"tui_model.decision_row no longer matches the surfacing registry "
f"(symmetric difference: {sorted(expected_keys ^ set(projected))}).\n"
f" registry values: {sorted(expected_keys)}\n"
f" decision_row keys: {sorted(projected)}"
)
for col, key in SCHEMA_TO_MODEL_KEYS.items():
assert projected[key] == FULL_ROW[col], (
f"route_decisions column {col!r} must surface as decision_row "
f"key {key!r}: got {projected[key]!r}, want {FULL_ROW[col]!r}"
)
def test_metrics_recent_decisions_selects_every_tracked_column(tmp_path):
"""The /metrics path carries every registered column end-to-end.
``request_id`` reaches the TUI ONLY through this SELECT (SSE publishes
before the backfill), so a narrowed SELECT silently erases forensics
fields a few seconds after a live decision arrives.
"""
conn = _db_from_schema(tmp_path, row_factory=sqlite3.Row)
cols = ",".join(FULL_ROW.keys())
placeholders = ",".join("?" * len(FULL_ROW))
conn.execute(
f"INSERT INTO route_decisions ({cols}) VALUES ({placeholders})",
tuple(FULL_ROW.values()),
)
conn.commit()
rows = metrics.recent_decisions(conn, limit=5)
conn.close()
assert rows, "the seeded FULL_ROW must come back from recent_decisions"
projected = tui_model.decision_row(rows[0])
unreachable = []
for col, key in SCHEMA_TO_MODEL_KEYS.items():
if projected.get(key) != FULL_ROW[col]:
unreachable.append(
f"route_decisions column {col!r} does not reach the TUI "
f"through /metrics — is it selected by "
f"metrics.recent_decisions()?"
)
assert not unreachable, "\n".join(unreachable)