Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
166 lines
8.7 KiB
Markdown
166 lines
8.7 KiB
Markdown
# Lift A.2: loops and stalling models, surfaced with the lever beside them
|
|
|
|
Status: planned -- brief ready; worktree created from origin/main 5aaf1e6 (after #103)
|
|
Date: 2026-09-26
|
|
This brief is the contract. Lift A (PR #102) is the base; do NOT read
|
|
`plans/no-progress-detection.md` (reference only).
|
|
|
|
## Goal
|
|
|
|
Meet the surfacing half of North Star #4 (`CLAUDE.md`): "Waste is surfaced in
|
|
the admin portal, and stopping it is one click." Lift A built the detection,
|
|
the alerts and the data. The portal does not yet show the evidence or put the
|
|
lever beside it:
|
|
|
|
- **Loops panel** (`admin/frontend/index.html#loops`, `loadLoops`) shows a
|
|
severity badge, the raw key `opencode-loop:<root id>` and `opened_at`.
|
|
Nothing an operator can act on.
|
|
- **Models page** has no per-model stall rollup, no Block button beside any
|
|
evidence, and no Blocked list. `blocked` exists only as a fourth value in
|
|
the Model Availability override dropdown.
|
|
|
|
The data is already there. On the live DB, 18 of 19 `watchdog_verdicts` rows
|
|
carry `model_id`, and all carry `calls_since_landed` and
|
|
`cost_since_landed_usd`, now that conversation identity arrives (fixed in
|
|
#102).
|
|
|
|
**Out of scope:** detection changes, recovery modes, alert channels other than
|
|
desktop, Lift B, and any change to how `blocked` is enforced (it already
|
|
excludes routing and pinned requests via `_admin_excluded_models`).
|
|
|
|
## Facts to build on (checked on origin/main 2026-09-26)
|
|
|
|
- `watchdog_verdicts`: `tick_id, session_id, session_root, agent, model_id,
|
|
provider, flagged, dup, top, top_what, landed, slow, coverage,
|
|
calls_since_landed, cost_since_landed_usd, llm_second_opinion, created_at`.
|
|
One row per judged session per tick. **`agent` holds the opencode session
|
|
TITLE**, not the agent slug. `cost_since_landed_usd` is cumulative, so it
|
|
grows tick by tick for the same session.
|
|
- `watchdog_alerts`: `dedup_key` (`opencode-loop:<root id>`), `state`,
|
|
`severity`, `flagged_ticks`, `opened_at`, `last_fired_at`, `resolved_at`.
|
|
- `admin_model_overrides`: `model_id, provider, availability, reason,
|
|
updated_at`. `POST /admin/api/models/{model_id:path}/{provider}/availability`
|
|
takes `{availability, reason}` (`_AvailabilityBody`, `src/admin.py` ~885);
|
|
`DELETE` on the same path clears the override (~2116).
|
|
- Existing endpoints: `GET /admin/api/watchdog/status`, `/loops` (~2834),
|
|
`/channels` GET/POST, `POST /test-alert`.
|
|
- `route_decisions.session_key` is `"c:" + opencode session id` and
|
|
`route_decisions.agent` is the agent slug when identity arrives.
|
|
- Models page cards: "Model Availability" (`models.html` ~226) and
|
|
"Per-model usage" (~310). JS unit harness: `tests/test_admin_js_units.py`
|
|
(marker-delimited pure functions, run under node).
|
|
|
|
## Components
|
|
|
|
### 1. Loops panel shows the evidence (`index.html#loops`)
|
|
|
|
- Extend `GET /admin/api/watchdog/loops` to return each OPEN alert joined to
|
|
its root's most recent FLAGGED verdict (any session in that root). Fields:
|
|
`dedup_key, root_id, session_id, title, agent_slug, model_id, provider,
|
|
severity, flagged_ticks, opened_at, top_what, dup, top, coverage,
|
|
calls_since_landed, cost_since_landed_usd, llm_second_opinion`.
|
|
`agent_slug` comes from the latest `route_decisions.agent` for
|
|
`"c:" + session_id` (NULL if none). Also return `resolved` (alerts resolved
|
|
in the last 24h, same fields, capped at 20).
|
|
- One row per open alert: severity badge, title (truncate ~40ch, full title in
|
|
`title=`), model, top repeated target, calls since last landed, $ since last
|
|
landed (`$0.00`, `n/a` when NULL), open for (relative time), and a **Block
|
|
model** button (component 3 flow) when `model_id` is known.
|
|
- Resolved rows below, dimmed. Session ids truncated with a tooltip.
|
|
- Empty state stays "No open alerts".
|
|
|
|
### 2. Per-model stall rollup (`models.html`, new card "Stalls by model")
|
|
|
|
- `GET /admin/api/watchdog/rollup?hours=N`, N in {1, 24, 168}. Per
|
|
`(model_id, provider)` with at least one flagged verdict in the window:
|
|
- `stalled_sessions` = `COUNT(DISTINCT session_root)` over flagged rows. Not
|
|
row counts: one row is written per tick.
|
|
- `stalled_usd` = per flagged `session_id`, take the MAX
|
|
`cost_since_landed_usd` in the window (it is cumulative), then sum per
|
|
model.
|
|
- `traffic_share` = the model's share of `route_decisions` in the same
|
|
window, its own column; never divide sessions by requests.
|
|
- `effective_availability` (override if any, else catalog).
|
|
- Window picker is radio-style (1h / 24h / 7d), default 24h.
|
|
- Each row: model, provider, stalled sessions, $ in stalled sessions, traffic
|
|
share, availability, and **Block model** (hidden when already blocked).
|
|
- Card sits ABOVE "Model Availability". Empty state: "No stalled sessions in
|
|
this window".
|
|
|
|
### 3. Block from the evidence, with a reason
|
|
|
|
- Block (from 1 or 2) opens an INLINE confirm row, not a browser dialog:
|
|
reason input prefilled `stalled <N> sessions in <window>` (from the Loops
|
|
panel: `looping session <title>`), editable, then Confirm / Cancel.
|
|
- Confirm posts `{availability: "blocked", reason}` to the existing endpoint,
|
|
then refreshes the rollup, the Blocked list and the availability table.
|
|
- The existing dropdown path keeps working; when it sets `blocked` it sends
|
|
reason `set from Models dropdown`.
|
|
|
|
### 4. Blocked list (`models.html`, card "Blocked models")
|
|
|
|
- Rows from `admin_model_overrides WHERE availability = 'blocked'`: model,
|
|
provider, reason, blocked since (`updated_at`), **Unblock**
|
|
(`DELETE .../availability`, inline confirm).
|
|
- Empty state: "No blocked models". Endpoint: reuse an existing models/overrides
|
|
endpoint if one returns `reason` and `updated_at`; otherwise add
|
|
`GET /admin/api/models/blocked`.
|
|
|
|
### 5. Alert text carries the evidence
|
|
|
|
- In `src/watchdog.py`, the trigger and escalate `AlertEvent.summary` becomes
|
|
`<title>: <calls> calls, $<cost> since last landed change on <model>.
|
|
<dashboard_base_url>#loops` (omit the `on <model>` / `$` parts when NULL).
|
|
Keep it one line; `notify-send` shows title and summary only.
|
|
|
|
## Todos (small; one commit each, by explicit path)
|
|
|
|
1. Loops endpoint join + `resolved` + endpoint tests (seeded tables: two
|
|
ticks for one session so the latest flagged verdict wins; a root with two
|
|
sessions; NULL model/cost).
|
|
2. Loops panel UI + pure formatting helpers (money, relative time, truncation)
|
|
in marker-delimited functions with node cases in `test_admin_js_units.py`.
|
|
3. Rollup endpoint + tests (distinct roots across ticks; MAX-then-SUM cost;
|
|
traffic share; window filter).
|
|
4. Stalls-by-model card + window radio + node cases for its helpers.
|
|
5. Inline Block confirm (shared by Loops and rollup) + reason on the dropdown
|
|
path + tests that the POST carries the reason.
|
|
6. Blocked list card + Unblock + tests.
|
|
7. Alert summary text in `watchdog.py` + a test asserting the summary.
|
|
8. Visual QA (below) and the done report.
|
|
|
|
## Visual QA (required; HTTP 200 is not QA)
|
|
|
|
- On 8081 (`scripts/sandbox.sh`) against a sqlite-backup copy of the live DB.
|
|
Seed the COPY with one open alert whose verdicts carry a model, calls and
|
|
cost, one resolved alert, and one blocked override with a reason. Never
|
|
write to the live DB.
|
|
- Screenshots of Home (Loops) and Models at 1400px and 800px wide, plus zoomed
|
|
crops of each new card. Check truncation, alignment, empty states, and that
|
|
Block -> Confirm -> Blocked list -> Unblock round-trips on screen.
|
|
- Save screenshots under the scratch/evidence dir, not the repo.
|
|
|
|
## Ground rules (lessons from Lift A)
|
|
|
|
- Worktree already exists: `/home/alee/Sources/6krrt-worktrees/no-progress-lift-a2`,
|
|
branch `feat/no-progress-lift-a2` from `origin/main` `5aaf1e6` (includes #103).
|
|
`/start-work` runs with `--worktree` set to it. Nothing in the main checkout.
|
|
- Dispatch to Sisyphus-Junior categories or `general`. NEVER to
|
|
`oh-my-claudecode:*` agents: they are pinned to Anthropic models this
|
|
opencode cannot reach, and return "completed" with zero work. A subagent
|
|
session with 0 tool calls did nothing.
|
|
- One commit per todo; no squashed "full implementation" commits.
|
|
- No test exemptions or skips for labelled cases. If a test cannot pass, stop
|
|
and report why.
|
|
- Lint: pinned `uvx ruff@0.16.9`, no NEW findings vs `origin/main` in touched
|
|
files (recipe in `plans/no-progress-lift-a.md` Ground rules). Paste the
|
|
gate's empty output in the done report.
|
|
- Full pytest: record the `origin/main` baseline first; the only known failure
|
|
is `test_incumbent_routing::test_debug_log_emits_incumbent_identity`.
|
|
- Port 8080 is production: no requests to it, not even GETs.
|
|
- Do not push, do not open a PR, do not remove the worktree. The done report
|
|
lists commits, tests, screenshots and operator steps (router restart).
|
|
- ASCII only; no middle dots. No prose paragraphs in the admin UI; mutually
|
|
exclusive choices are radio-style.
|
|
- Stop a worker whose context passes ~150k tokens and start a fresh one.
|