feat(tui): catch the dashboard up to the schema #33

Merged
alee merged 4 commits from feat/tui-overhaul into main 2026-09-05 15:36:21 +00:00
Owner

Five columns shipped to route_decisions across recent merges and none
reached the dashboard. The decision table had no timestamp at all.

The load-bearing part is the drift test, not the columns

tests/test_tui_schema_drift.py diffs PRAGMA table_info(route_decisions)
against two explicit registries — SCHEMA_TO_MODEL_KEYS (the surfacing
decision) and UNSURFACED_COLUMNS (deliberately not shown, empty today).
A new column must land in one or the other or the test fails naming the
column
, in both directions.

Live proof the drift is real: tests/test_route_decisions.py's own
hand-maintained column list had already drifted, missing request_id. Fixed
here.

Catch-up

  • Decision table: time first (short HH:MM:SS local — the date is almost
    always today and would crowd out selected), profile after ctx, and
    flex becomes flags, composing the flex label with an E marker for
    epsilon-greedy picks. An exploratory pick is a random sample, not the
    ranking's judgement, and must be distinguishable.
  • Detail popup gains profile, exploration, pinch_original_tokens,
    pinch_final_tokens, request_id, session_key.
  • profile renders blank rather than the literal "None", which in a dense
    table reads as a real profile name.

Quota panel (unblocked by #30)

Leads with balance $X left · burn $Y/h · runway ~Nh; the plan-percentage bar
is demoted. Wired outside the plan-configured gate on purpose — the credit
balance is meaningful whether or not a kWh plan is set.

burn/runway are legitimately None when the segment is too short to
extrapolate. On the live deployment that is the current state, so it is
the first thing an operator sees, not an edge case: the lead says burn n/a
and #quota-note carries runway_note verbatim. Tests assert it is never
blank, never a bare 0, never None.

Warning-class tripwire

tests/test_tui_warnings.py registers all 10 classes (including PR #31's
capability sub-ceiling and both rejection prefixes) and asserts both
directions: every registered class fires against a fixture that provokes all
ten, and every emitted warning is registered — the second failing with the
RAW warning text, because the point is that nobody knew the class existed.

That is the 2026-09-04 vision-ceiling failure mode: computed correctly, never
displayed.

Verification

  • Full suite 1229 passed, with local_energy.enabled both true and false.
  • Textual boundary grep empty — textual still imported only by TUI modules.
  • Root config/config.local.yaml byte-identical (cc3d4c2a…).
  • dispatcher.py, routing.py, admin.py, tui_screens.py, tui_sse.py
    and config/config.yaml untouched.
  • Evidence under .omo/evidence/tui-overhaul/.

Known interaction

classifier-fallback-cascade adds a degraded-source warning class. When it
lands, test_every_emitted_warning_belongs_to_a_registered_class will fail
until a matcher is added — that is the tripwire working. The fix is one
matcher entry, not disabling the test.

Five columns shipped to `route_decisions` across recent merges and none reached the dashboard. The decision table had no timestamp at all. ## The load-bearing part is the drift test, not the columns `tests/test_tui_schema_drift.py` diffs `PRAGMA table_info(route_decisions)` against two explicit registries — `SCHEMA_TO_MODEL_KEYS` (the surfacing decision) and `UNSURFACED_COLUMNS` (deliberately not shown, empty today). A new column must land in one or the other or the test fails **naming the column**, in both directions. Live proof the drift is real: `tests/test_route_decisions.py`'s own hand-maintained column list had already drifted, missing `request_id`. Fixed here. ## Catch-up - Decision table: `time` first (short `HH:MM:SS` local — the date is almost always today and would crowd out `selected`), `profile` after `ctx`, and `flex` becomes `flags`, composing the flex label with an `E` marker for epsilon-greedy picks. An exploratory pick is a random sample, not the ranking's judgement, and must be distinguishable. - Detail popup gains `profile`, `exploration`, `pinch_original_tokens`, `pinch_final_tokens`, `request_id`, `session_key`. - `profile` renders blank rather than the literal `"None"`, which in a dense table reads as a real profile name. ## Quota panel (unblocked by #30) Leads with `balance $X left · burn $Y/h · runway ~Nh`; the plan-percentage bar is demoted. Wired outside the plan-configured gate on purpose — the credit balance is meaningful whether or not a kWh plan is set. `burn`/`runway` are legitimately `None` when the segment is too short to extrapolate. **On the live deployment that is the current state**, so it is the first thing an operator sees, not an edge case: the lead says `burn n/a` and `#quota-note` carries `runway_note` verbatim. Tests assert it is never blank, never a bare `0`, never `None`. ## Warning-class tripwire `tests/test_tui_warnings.py` registers all 10 classes (including PR #31's capability sub-ceiling and both rejection prefixes) and asserts both directions: every registered class fires against a fixture that provokes all ten, and every emitted warning is registered — the second failing with the RAW warning text, because the point is that nobody knew the class existed. That is the 2026-09-04 vision-ceiling failure mode: computed correctly, never displayed. ## Verification - Full suite **1229 passed**, with `local_energy.enabled` both true and false. - Textual boundary grep empty — `textual` still imported only by TUI modules. - Root `config/config.local.yaml` byte-identical (`cc3d4c2a…`). - `dispatcher.py`, `routing.py`, `admin.py`, `tui_screens.py`, `tui_sse.py` and `config/config.yaml` untouched. - Evidence under `.omo/evidence/tui-overhaul/`. ## Known interaction `classifier-fallback-cascade` adds a degraded-source warning class. When it lands, `test_every_emitted_warning_belongs_to_a_registered_class` will fail until a matcher is added — that is the tripwire working. The fix is one matcher entry, not disabling the test.
alee added 4 commits 2026-09-05 06:28:34 +00:00
The decision feed had no timestamp at all and the eight columns predated
five schema additions. Adds `time` first (short HH:MM:SS local -- the date
is almost always today and would crowd out `selected`, the column operators
actually read) and `profile` after `ctx`, since a decision cannot be read
without knowing which profile filtered the candidate set.

`flex` becomes `flags` and now composes the flex label with an `E` marker
for epsilon-greedy picks. One cell rather than two columns: the terminal is
not wide, and an exploratory pick must be visually distinguishable or an
operator reads a deliberate random sample as the router's judgement.

`profile` renders `or ""` so an absent value is blank rather than the
literal "None", which in a dense table reads as a real profile name.

Placeholder row widened to arity 10 -- a short placeholder raises inside
Textual at render time rather than looking wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
The panel led with a progress bar against plan_kwh, which is misleading:
usage routinely exceeds 100% and nothing fails, because overage bills
against a credit balance. PR #30 made balance, burn rate and runway
available on /metrics; this surfaces them.

Balance and runway now lead, in a bold #quota-lead line, with the plan
percentage bar demoted below. Wired outside the plan-configured gate on
purpose -- the credit balance is meaningful whether or not a kWh plan is
set, and it is the figure that actually stops traffic.

burn and runway are legitimately None when the sample segment is too short
to extrapolate, which is PR #30's guard against a wild rate right after a
top-up. On the live deployment that is the CURRENT state, so it is the
first thing an operator sees rather than an edge case. The lead renders
"burn n/a" and #quota-note carries runway_note verbatim -- never blank,
never a bare 0, never the literal None. Tests assert all three.

Formatting is 2dp rather than _fmt_usd's 6: money an operator acts on, not
per-request micro-money.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
The 2026-09-04 vision-ceiling incident was invisible for ~19 hours because
a whole warning family was computed correctly and never displayed. The fix
there was surfacing, not computing -- so the guard has to be a test that
fails when a class exists but nothing renders it.

Two directions, both load-bearing. Every REGISTERED class must fire against
a fixture that provokes all ten, catching the fixture rotting or a message
shape drifting under a matcher. And every EMITTED warning must be
registered, catching a new class added to scoring_coverage that nothing
displays -- this one fails with the RAW warning text, because the whole
point is that nobody knew the class existed.

No assertions on warning count: the numbers move as fixture semantics
evolve, and what matters is class coverage, not arity.

Three fixes to the fixture the plan specified, found by running it:
proficiency.last_updated is NOT NULL; a seed_reference row is identified by
task_category rather than a `kind` column that does not exist; and
scoring_coverage reads `.value` off default_flex_preference, which is an
enum in real config, so a bare string is not a faithful stand-in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
alee merged commit b5c39ad87f into main 2026-09-05 15:36:21 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#33