Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
8.7 KiB
Lift A.2: loops and stalling models, surfaced with the lever beside them
Status: planned -- brief ready; worktree created from origin/main 5aaf1e6 (after #103)
Date: 2026-09-26
This brief is the contract. Lift A (PR #102) is the base; do NOT read
plans/no-progress-detection.md (reference only).
Goal
Meet the surfacing half of North Star #4 (CLAUDE.md): "Waste is surfaced in
the admin portal, and stopping it is one click." Lift A built the detection,
the alerts and the data. The portal does not yet show the evidence or put the
lever beside it:
- Loops panel (
admin/frontend/index.html#loops,loadLoops) shows a severity badge, the raw keyopencode-loop:<root id>andopened_at. Nothing an operator can act on. - Models page has no per-model stall rollup, no Block button beside any
evidence, and no Blocked list.
blockedexists only as a fourth value in the Model Availability override dropdown.
The data is already there. On the live DB, 18 of 19 watchdog_verdicts rows
carry model_id, and all carry calls_since_landed and
cost_since_landed_usd, now that conversation identity arrives (fixed in
#102).
Out of scope: detection changes, recovery modes, alert channels other than
desktop, Lift B, and any change to how blocked is enforced (it already
excludes routing and pinned requests via _admin_excluded_models).
Facts to build on (checked on origin/main 2026-09-26)
watchdog_verdicts:tick_id, session_id, session_root, agent, model_id, provider, flagged, dup, top, top_what, landed, slow, coverage, calls_since_landed, cost_since_landed_usd, llm_second_opinion, created_at. One row per judged session per tick.agentholds the opencode session TITLE, not the agent slug.cost_since_landed_usdis cumulative, so it grows tick by tick for the same session.watchdog_alerts:dedup_key(opencode-loop:<root id>),state,severity,flagged_ticks,opened_at,last_fired_at,resolved_at.admin_model_overrides:model_id, provider, availability, reason, updated_at.POST /admin/api/models/{model_id:path}/{provider}/availabilitytakes{availability, reason}(_AvailabilityBody,src/admin.py~885);DELETEon the same path clears the override (~2116).- Existing endpoints:
GET /admin/api/watchdog/status,/loops(~2834),/channelsGET/POST,POST /test-alert. route_decisions.session_keyis"c:" + opencode session idandroute_decisions.agentis the agent slug when identity arrives.- Models page cards: "Model Availability" (
models.html~226) and "Per-model usage" (~310). JS unit harness:tests/test_admin_js_units.py(marker-delimited pure functions, run under node).
Components
1. Loops panel shows the evidence (index.html#loops)
- Extend
GET /admin/api/watchdog/loopsto return each OPEN alert joined to its root's most recent FLAGGED verdict (any session in that root). Fields:dedup_key, root_id, session_id, title, agent_slug, model_id, provider, severity, flagged_ticks, opened_at, top_what, dup, top, coverage, calls_since_landed, cost_since_landed_usd, llm_second_opinion.agent_slugcomes from the latestroute_decisions.agentfor"c:" + session_id(NULL if none). Also returnresolved(alerts resolved in the last 24h, same fields, capped at 20). - One row per open alert: severity badge, title (truncate ~40ch, full title in
title=), model, top repeated target, calls since last landed, $ since last landed ($0.00,n/awhen NULL), open for (relative time), and a Block model button (component 3 flow) whenmodel_idis known. - Resolved rows below, dimmed. Session ids truncated with a tooltip.
- Empty state stays "No open alerts".
2. Per-model stall rollup (models.html, new card "Stalls by model")
GET /admin/api/watchdog/rollup?hours=N, N in {1, 24, 168}. Per(model_id, provider)with at least one flagged verdict in the window:stalled_sessions=COUNT(DISTINCT session_root)over flagged rows. Not row counts: one row is written per tick.stalled_usd= per flaggedsession_id, take the MAXcost_since_landed_usdin the window (it is cumulative), then sum per model.traffic_share= the model's share ofroute_decisionsin the same window, its own column; never divide sessions by requests.effective_availability(override if any, else catalog).
- Window picker is radio-style (1h / 24h / 7d), default 24h.
- Each row: model, provider, stalled sessions, $ in stalled sessions, traffic share, availability, and Block model (hidden when already blocked).
- Card sits ABOVE "Model Availability". Empty state: "No stalled sessions in this window".
3. Block from the evidence, with a reason
- Block (from 1 or 2) opens an INLINE confirm row, not a browser dialog:
reason input prefilled
stalled <N> sessions in <window>(from the Loops panel:looping session <title>), editable, then Confirm / Cancel. - Confirm posts
{availability: "blocked", reason}to the existing endpoint, then refreshes the rollup, the Blocked list and the availability table. - The existing dropdown path keeps working; when it sets
blockedit sends reasonset from Models dropdown.
4. Blocked list (models.html, card "Blocked models")
- Rows from
admin_model_overrides WHERE availability = 'blocked': model, provider, reason, blocked since (updated_at), Unblock (DELETE .../availability, inline confirm). - Empty state: "No blocked models". Endpoint: reuse an existing models/overrides
endpoint if one returns
reasonandupdated_at; otherwise addGET /admin/api/models/blocked.
5. Alert text carries the evidence
- In
src/watchdog.py, the trigger and escalateAlertEvent.summarybecomes<title>: <calls> calls, $<cost> since last landed change on <model>. <dashboard_base_url>#loops(omit theon <model>/$parts when NULL). Keep it one line;notify-sendshows title and summary only.
Todos (small; one commit each, by explicit path)
- Loops endpoint join +
resolved+ endpoint tests (seeded tables: two ticks for one session so the latest flagged verdict wins; a root with two sessions; NULL model/cost). - Loops panel UI + pure formatting helpers (money, relative time, truncation)
in marker-delimited functions with node cases in
test_admin_js_units.py. - Rollup endpoint + tests (distinct roots across ticks; MAX-then-SUM cost; traffic share; window filter).
- Stalls-by-model card + window radio + node cases for its helpers.
- Inline Block confirm (shared by Loops and rollup) + reason on the dropdown path + tests that the POST carries the reason.
- Blocked list card + Unblock + tests.
- Alert summary text in
watchdog.py+ a test asserting the summary. - Visual QA (below) and the done report.
Visual QA (required; HTTP 200 is not QA)
- On 8081 (
scripts/sandbox.sh) against a sqlite-backup copy of the live DB. Seed the COPY with one open alert whose verdicts carry a model, calls and cost, one resolved alert, and one blocked override with a reason. Never write to the live DB. - Screenshots of Home (Loops) and Models at 1400px and 800px wide, plus zoomed crops of each new card. Check truncation, alignment, empty states, and that Block -> Confirm -> Blocked list -> Unblock round-trips on screen.
- Save screenshots under the scratch/evidence dir, not the repo.
Ground rules (lessons from Lift A)
- Worktree already exists:
/home/alee/Sources/6krrt-worktrees/no-progress-lift-a2, branchfeat/no-progress-lift-a2fromorigin/main5aaf1e6(includes #103)./start-workruns with--worktreeset to it. Nothing in the main checkout. - Dispatch to Sisyphus-Junior categories or
general. NEVER tooh-my-claudecode:*agents: they are pinned to Anthropic models this opencode cannot reach, and return "completed" with zero work. A subagent session with 0 tool calls did nothing. - One commit per todo; no squashed "full implementation" commits.
- No test exemptions or skips for labelled cases. If a test cannot pass, stop and report why.
- Lint: pinned
uvx ruff@0.16.9, no NEW findings vsorigin/mainin touched files (recipe inplans/no-progress-lift-a.mdGround rules). Paste the gate's empty output in the done report. - Full pytest: record the
origin/mainbaseline first; the only known failure istest_incumbent_routing::test_debug_log_emits_incumbent_identity. - Port 8080 is production: no requests to it, not even GETs.
- Do not push, do not open a PR, do not remove the worktree. The done report lists commits, tests, screenshots and operator steps (router restart).
- ASCII only; no middle dots. No prose paragraphs in the admin UI; mutually exclusive choices are radio-style.
- Stop a worker whose context passes ~150k tokens and start a fresh one.