Files
6krrt/plans/no-progress-lift-a2.md
2026-09-26 22:46:00 -04:00

166 lines
8.7 KiB
Markdown

# Lift A.2: loops and stalling models, surfaced with the lever beside them
Status: planned -- brief ready; worktree created from origin/main 5aaf1e6 (after #103)
Date: 2026-09-26
This brief is the contract. Lift A (PR #102) is the base; do NOT read
`plans/no-progress-detection.md` (reference only).
## Goal
Meet the surfacing half of North Star #4 (`CLAUDE.md`): "Waste is surfaced in
the admin portal, and stopping it is one click." Lift A built the detection,
the alerts and the data. The portal does not yet show the evidence or put the
lever beside it:
- **Loops panel** (`admin/frontend/index.html#loops`, `loadLoops`) shows a
severity badge, the raw key `opencode-loop:<root id>` and `opened_at`.
Nothing an operator can act on.
- **Models page** has no per-model stall rollup, no Block button beside any
evidence, and no Blocked list. `blocked` exists only as a fourth value in
the Model Availability override dropdown.
The data is already there. On the live DB, 18 of 19 `watchdog_verdicts` rows
carry `model_id`, and all carry `calls_since_landed` and
`cost_since_landed_usd`, now that conversation identity arrives (fixed in
#102).
**Out of scope:** detection changes, recovery modes, alert channels other than
desktop, Lift B, and any change to how `blocked` is enforced (it already
excludes routing and pinned requests via `_admin_excluded_models`).
## Facts to build on (checked on origin/main 2026-09-26)
- `watchdog_verdicts`: `tick_id, session_id, session_root, agent, model_id,
provider, flagged, dup, top, top_what, landed, slow, coverage,
calls_since_landed, cost_since_landed_usd, llm_second_opinion, created_at`.
One row per judged session per tick. **`agent` holds the opencode session
TITLE**, not the agent slug. `cost_since_landed_usd` is cumulative, so it
grows tick by tick for the same session.
- `watchdog_alerts`: `dedup_key` (`opencode-loop:<root id>`), `state`,
`severity`, `flagged_ticks`, `opened_at`, `last_fired_at`, `resolved_at`.
- `admin_model_overrides`: `model_id, provider, availability, reason,
updated_at`. `POST /admin/api/models/{model_id:path}/{provider}/availability`
takes `{availability, reason}` (`_AvailabilityBody`, `src/admin.py` ~885);
`DELETE` on the same path clears the override (~2116).
- Existing endpoints: `GET /admin/api/watchdog/status`, `/loops` (~2834),
`/channels` GET/POST, `POST /test-alert`.
- `route_decisions.session_key` is `"c:" + opencode session id` and
`route_decisions.agent` is the agent slug when identity arrives.
- Models page cards: "Model Availability" (`models.html` ~226) and
"Per-model usage" (~310). JS unit harness: `tests/test_admin_js_units.py`
(marker-delimited pure functions, run under node).
## Components
### 1. Loops panel shows the evidence (`index.html#loops`)
- Extend `GET /admin/api/watchdog/loops` to return each OPEN alert joined to
its root's most recent FLAGGED verdict (any session in that root). Fields:
`dedup_key, root_id, session_id, title, agent_slug, model_id, provider,
severity, flagged_ticks, opened_at, top_what, dup, top, coverage,
calls_since_landed, cost_since_landed_usd, llm_second_opinion`.
`agent_slug` comes from the latest `route_decisions.agent` for
`"c:" + session_id` (NULL if none). Also return `resolved` (alerts resolved
in the last 24h, same fields, capped at 20).
- One row per open alert: severity badge, title (truncate ~40ch, full title in
`title=`), model, top repeated target, calls since last landed, $ since last
landed (`$0.00`, `n/a` when NULL), open for (relative time), and a **Block
model** button (component 3 flow) when `model_id` is known.
- Resolved rows below, dimmed. Session ids truncated with a tooltip.
- Empty state stays "No open alerts".
### 2. Per-model stall rollup (`models.html`, new card "Stalls by model")
- `GET /admin/api/watchdog/rollup?hours=N`, N in {1, 24, 168}. Per
`(model_id, provider)` with at least one flagged verdict in the window:
- `stalled_sessions` = `COUNT(DISTINCT session_root)` over flagged rows. Not
row counts: one row is written per tick.
- `stalled_usd` = per flagged `session_id`, take the MAX
`cost_since_landed_usd` in the window (it is cumulative), then sum per
model.
- `traffic_share` = the model's share of `route_decisions` in the same
window, its own column; never divide sessions by requests.
- `effective_availability` (override if any, else catalog).
- Window picker is radio-style (1h / 24h / 7d), default 24h.
- Each row: model, provider, stalled sessions, $ in stalled sessions, traffic
share, availability, and **Block model** (hidden when already blocked).
- Card sits ABOVE "Model Availability". Empty state: "No stalled sessions in
this window".
### 3. Block from the evidence, with a reason
- Block (from 1 or 2) opens an INLINE confirm row, not a browser dialog:
reason input prefilled `stalled <N> sessions in <window>` (from the Loops
panel: `looping session <title>`), editable, then Confirm / Cancel.
- Confirm posts `{availability: "blocked", reason}` to the existing endpoint,
then refreshes the rollup, the Blocked list and the availability table.
- The existing dropdown path keeps working; when it sets `blocked` it sends
reason `set from Models dropdown`.
### 4. Blocked list (`models.html`, card "Blocked models")
- Rows from `admin_model_overrides WHERE availability = 'blocked'`: model,
provider, reason, blocked since (`updated_at`), **Unblock**
(`DELETE .../availability`, inline confirm).
- Empty state: "No blocked models". Endpoint: reuse an existing models/overrides
endpoint if one returns `reason` and `updated_at`; otherwise add
`GET /admin/api/models/blocked`.
### 5. Alert text carries the evidence
- In `src/watchdog.py`, the trigger and escalate `AlertEvent.summary` becomes
`<title>: <calls> calls, $<cost> since last landed change on <model>.
<dashboard_base_url>#loops` (omit the `on <model>` / `$` parts when NULL).
Keep it one line; `notify-send` shows title and summary only.
## Todos (small; one commit each, by explicit path)
1. Loops endpoint join + `resolved` + endpoint tests (seeded tables: two
ticks for one session so the latest flagged verdict wins; a root with two
sessions; NULL model/cost).
2. Loops panel UI + pure formatting helpers (money, relative time, truncation)
in marker-delimited functions with node cases in `test_admin_js_units.py`.
3. Rollup endpoint + tests (distinct roots across ticks; MAX-then-SUM cost;
traffic share; window filter).
4. Stalls-by-model card + window radio + node cases for its helpers.
5. Inline Block confirm (shared by Loops and rollup) + reason on the dropdown
path + tests that the POST carries the reason.
6. Blocked list card + Unblock + tests.
7. Alert summary text in `watchdog.py` + a test asserting the summary.
8. Visual QA (below) and the done report.
## Visual QA (required; HTTP 200 is not QA)
- On 8081 (`scripts/sandbox.sh`) against a sqlite-backup copy of the live DB.
Seed the COPY with one open alert whose verdicts carry a model, calls and
cost, one resolved alert, and one blocked override with a reason. Never
write to the live DB.
- Screenshots of Home (Loops) and Models at 1400px and 800px wide, plus zoomed
crops of each new card. Check truncation, alignment, empty states, and that
Block -> Confirm -> Blocked list -> Unblock round-trips on screen.
- Save screenshots under the scratch/evidence dir, not the repo.
## Ground rules (lessons from Lift A)
- Worktree already exists: `/home/alee/Sources/6krrt-worktrees/no-progress-lift-a2`,
branch `feat/no-progress-lift-a2` from `origin/main` `5aaf1e6` (includes #103).
`/start-work` runs with `--worktree` set to it. Nothing in the main checkout.
- Dispatch to Sisyphus-Junior categories or `general`. NEVER to
`oh-my-claudecode:*` agents: they are pinned to Anthropic models this
opencode cannot reach, and return "completed" with zero work. A subagent
session with 0 tool calls did nothing.
- One commit per todo; no squashed "full implementation" commits.
- No test exemptions or skips for labelled cases. If a test cannot pass, stop
and report why.
- Lint: pinned `uvx ruff@0.16.9`, no NEW findings vs `origin/main` in touched
files (recipe in `plans/no-progress-lift-a.md` Ground rules). Paste the
gate's empty output in the done report.
- Full pytest: record the `origin/main` baseline first; the only known failure
is `test_incumbent_routing::test_debug_log_emits_incumbent_identity`.
- Port 8080 is production: no requests to it, not even GETs.
- Do not push, do not open a PR, do not remove the worktree. The done report
lists commits, tests, screenshots and operator steps (router restart).
- ASCII only; no middle dots. No prose paragraphs in the admin UI; mutually
exclusive choices are radio-style.
- Stop a worker whose context passes ~150k tokens and start a fresh one.