Files
6krrt/plans/tui-overhaul.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

127 lines
6.0 KiB
Markdown

# TUI overhaul: catch the dashboard up to the schema
Status: done -- tui.py plus the tui_model.py split
**Status: FINAL — decision-complete.** Written 2026-09-04 against `main` at
`a0c7e8e`.
## Why now
The TUI (`src/tui.py`, `tui_model.py`, `tui_screens.py`, `tui_sse.py` — ~1,000
lines total) has drifted behind the data it displays. Five columns were added
to `route_decisions` across this session's merges and none reached the
dashboard, and the decision table has never had a timestamp.
The admin portal got a Profile column in PR #25. The TUI did not. It is the
same data.
## 1. The decision table is missing a timestamp
Declared columns (`src/tui.py:314-316`):
```
"id", "kind", "category", "tier", "ctx", "selected", "est $", "flex"
```
`observed_at` **is already carried** by `decision_row` in `tui_model.py` — it
is fetched and then never rendered. So this is a display change, not a data
change.
Add a `time` column. Render it **short** (`HH:MM:SS`), not the full ISO
timestamp: the rows are dense, the date is almost always today, and the full
form would crowd out `selected`, which is the column people actually read.
Put it first — it is the natural scan axis for a live feed.
## 2. Five columns exist in the schema and are invisible
Measured by diffing `PRAGMA table_info(route_decisions)` against what
`decision_row` exposes:
| column | shipped in | why it matters |
|---|---|---|
| `profile` | PR #25 | which named profile served the request |
| `exploration` | proficiency branch | whether epsilon-greedy picked this, not the ranking |
| `pinch_original_tokens` | PR #18 | context pruning input |
| `pinch_final_tokens` | PR #18 | context pruning output |
| `request_id` | proficiency branch | the join key to `/outcome` reports |
| `session_key` | earlier | hashed session fingerprint |
**Do not add six more columns to the table.** It already has eight and the
terminal is not wide. Instead:
- Add **`profile`** to the table proper. It changes which models were even
considered, so a decision cannot be read without it — the same argument that
earned it a column in the admin portal.
- Add an **`E` flag** in the existing flags idiom for `exploration`, alongside
how `flex` is already rendered. An exploratory pick is not a ranking result
and must be visually distinguishable, or the operator reads a deliberate
random sample as the router's judgement.
- Put **`pinch_*`, `request_id`, `session_key`** in the **detail popup**
(`tui_screens.py`, opened with Enter or `e`), which already shows the full
decision JSON. They are per-decision forensics, not scan-axis data.
## 3. The quota panel shows the wrong thing
**This section depends on `plans/quota-balance-and-burn-rate.md` and must not
land before it.** That plan replaces percentage-of-plan with balance and burn
rate, because the current framing is measurably wrong: the warning says "a
quota is a wall, not a bill — requests fail rather than costing more" while
usage sat at 146% of plan and nothing failed, since the provider bills overage
against a credit balance.
Once that lands, the TUI panel (`#quota-panel`, `#quota-progress`,
`#quota-legend`) should lead with **balance and projected runway**
("$12.19 left, ~23h at current burn") and demote the percentage bar to
secondary. A progress bar against a plan figure that is routinely exceeded is
actively misleading — it implies a ceiling that does not exist.
**Sequencing:** if the quota plan has not landed when this one runs, do items
1, 2 and 4 and leave the panel alone. Do NOT reimplement balance/burn
independently in the TUI — `metrics.quota_burn` is the single source and the
TUI reads `/metrics`.
## 4. Surface the warnings that already exist
`coverage.warnings` from `/metrics` already carries catalog staleness, quota
burn, scoring coverage gaps and ceiling warnings. `#warnings-panel` exists.
Confirm every warning class actually reaches it — the vision-ceiling incident
on 2026-09-04 showed a whole warning family that was computed and never
displayed, and the fix there was surfacing, not computing.
This is a verification task as much as a feature: for each warning the
`/metrics` `coverage.warnings` list can emit, assert it renders.
## Non-goals
- No new data. Everything here is already in `route_decisions` or `/metrics`.
- Do not widen the decision table beyond one added column plus one flag.
- Do not reimplement any metric in the TUI. `tui_model.py` is the pure data
layer over `/metrics` and `/events/decisions`; keep the computation in
`metrics.py`.
- Do not import `textual` outside the TUI modules. The dispatch path must stay
free of the UI dependency — that separation is deliberate and tested.
- No colour/theme rework. This is about information, not appearance.
## Success criteria
- The decision table shows a short `HH:MM:SS` time column, first.
- `profile` is a column; `exploration` renders as a flag beside `flex`.
- `pinch_original_tokens`, `pinch_final_tokens`, `request_id` and
`session_key` appear in the detail popup.
- A test diffs `PRAGMA table_info(route_decisions)` against what the TUI
model exposes and fails if a column is added to the schema without a
decision about surfacing it. **This is the test that stops the drift
recurring** — the rest of this plan is a one-time catch-up, this is the part
that keeps it caught up.
- Every warning class `/metrics` can emit renders in `#warnings-panel`.
- Quota panel leads with balance and runway **if** the quota plan has landed;
otherwise untouched and noted.
- `textual` still imported only by TUI modules (existing test stays green).
- Full suite green with `local_energy.enabled` both true and false.
- The user's `config/config.local.yaml` is byte-identical after the run
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). **Corrected
2026-09-05:** this previously named `config/config.yaml`. PR #26 moved
deployment values into the gitignored overlay, so `config/config.yaml` is
clean and tracked, and editing it is an ordinary commit. The file that
cannot be recovered is the overlay — it is not in git history at all.