Files
6krrt/plans/no-progress-lift-a2.md
2026-09-26 22:46:00 -04:00

8.7 KiB

Lift A.2: loops and stalling models, surfaced with the lever beside them

Status: planned -- brief ready; worktree created from origin/main 5aaf1e6 (after #103) Date: 2026-09-26 This brief is the contract. Lift A (PR #102) is the base; do NOT read plans/no-progress-detection.md (reference only).

Goal

Meet the surfacing half of North Star #4 (CLAUDE.md): "Waste is surfaced in the admin portal, and stopping it is one click." Lift A built the detection, the alerts and the data. The portal does not yet show the evidence or put the lever beside it:

  • Loops panel (admin/frontend/index.html#loops, loadLoops) shows a severity badge, the raw key opencode-loop:<root id> and opened_at. Nothing an operator can act on.
  • Models page has no per-model stall rollup, no Block button beside any evidence, and no Blocked list. blocked exists only as a fourth value in the Model Availability override dropdown.

The data is already there. On the live DB, 18 of 19 watchdog_verdicts rows carry model_id, and all carry calls_since_landed and cost_since_landed_usd, now that conversation identity arrives (fixed in #102).

Out of scope: detection changes, recovery modes, alert channels other than desktop, Lift B, and any change to how blocked is enforced (it already excludes routing and pinned requests via _admin_excluded_models).

Facts to build on (checked on origin/main 2026-09-26)

  • watchdog_verdicts: tick_id, session_id, session_root, agent, model_id, provider, flagged, dup, top, top_what, landed, slow, coverage, calls_since_landed, cost_since_landed_usd, llm_second_opinion, created_at. One row per judged session per tick. agent holds the opencode session TITLE, not the agent slug. cost_since_landed_usd is cumulative, so it grows tick by tick for the same session.
  • watchdog_alerts: dedup_key (opencode-loop:<root id>), state, severity, flagged_ticks, opened_at, last_fired_at, resolved_at.
  • admin_model_overrides: model_id, provider, availability, reason, updated_at. POST /admin/api/models/{model_id:path}/{provider}/availability takes {availability, reason} (_AvailabilityBody, src/admin.py ~885); DELETE on the same path clears the override (~2116).
  • Existing endpoints: GET /admin/api/watchdog/status, /loops (~2834), /channels GET/POST, POST /test-alert.
  • route_decisions.session_key is "c:" + opencode session id and route_decisions.agent is the agent slug when identity arrives.
  • Models page cards: "Model Availability" (models.html ~226) and "Per-model usage" (~310). JS unit harness: tests/test_admin_js_units.py (marker-delimited pure functions, run under node).

Components

1. Loops panel shows the evidence (index.html#loops)

  • Extend GET /admin/api/watchdog/loops to return each OPEN alert joined to its root's most recent FLAGGED verdict (any session in that root). Fields: dedup_key, root_id, session_id, title, agent_slug, model_id, provider, severity, flagged_ticks, opened_at, top_what, dup, top, coverage, calls_since_landed, cost_since_landed_usd, llm_second_opinion. agent_slug comes from the latest route_decisions.agent for "c:" + session_id (NULL if none). Also return resolved (alerts resolved in the last 24h, same fields, capped at 20).
  • One row per open alert: severity badge, title (truncate ~40ch, full title in title=), model, top repeated target, calls since last landed, $ since last landed ($0.00, n/a when NULL), open for (relative time), and a Block model button (component 3 flow) when model_id is known.
  • Resolved rows below, dimmed. Session ids truncated with a tooltip.
  • Empty state stays "No open alerts".

2. Per-model stall rollup (models.html, new card "Stalls by model")

  • GET /admin/api/watchdog/rollup?hours=N, N in {1, 24, 168}. Per (model_id, provider) with at least one flagged verdict in the window:
    • stalled_sessions = COUNT(DISTINCT session_root) over flagged rows. Not row counts: one row is written per tick.
    • stalled_usd = per flagged session_id, take the MAX cost_since_landed_usd in the window (it is cumulative), then sum per model.
    • traffic_share = the model's share of route_decisions in the same window, its own column; never divide sessions by requests.
    • effective_availability (override if any, else catalog).
  • Window picker is radio-style (1h / 24h / 7d), default 24h.
  • Each row: model, provider, stalled sessions, $ in stalled sessions, traffic share, availability, and Block model (hidden when already blocked).
  • Card sits ABOVE "Model Availability". Empty state: "No stalled sessions in this window".

3. Block from the evidence, with a reason

  • Block (from 1 or 2) opens an INLINE confirm row, not a browser dialog: reason input prefilled stalled <N> sessions in <window> (from the Loops panel: looping session <title>), editable, then Confirm / Cancel.
  • Confirm posts {availability: "blocked", reason} to the existing endpoint, then refreshes the rollup, the Blocked list and the availability table.
  • The existing dropdown path keeps working; when it sets blocked it sends reason set from Models dropdown.

4. Blocked list (models.html, card "Blocked models")

  • Rows from admin_model_overrides WHERE availability = 'blocked': model, provider, reason, blocked since (updated_at), Unblock (DELETE .../availability, inline confirm).
  • Empty state: "No blocked models". Endpoint: reuse an existing models/overrides endpoint if one returns reason and updated_at; otherwise add GET /admin/api/models/blocked.

5. Alert text carries the evidence

  • In src/watchdog.py, the trigger and escalate AlertEvent.summary becomes <title>: <calls> calls, $<cost> since last landed change on <model>. <dashboard_base_url>#loops (omit the on <model> / $ parts when NULL). Keep it one line; notify-send shows title and summary only.

Todos (small; one commit each, by explicit path)

  1. Loops endpoint join + resolved + endpoint tests (seeded tables: two ticks for one session so the latest flagged verdict wins; a root with two sessions; NULL model/cost).
  2. Loops panel UI + pure formatting helpers (money, relative time, truncation) in marker-delimited functions with node cases in test_admin_js_units.py.
  3. Rollup endpoint + tests (distinct roots across ticks; MAX-then-SUM cost; traffic share; window filter).
  4. Stalls-by-model card + window radio + node cases for its helpers.
  5. Inline Block confirm (shared by Loops and rollup) + reason on the dropdown path + tests that the POST carries the reason.
  6. Blocked list card + Unblock + tests.
  7. Alert summary text in watchdog.py + a test asserting the summary.
  8. Visual QA (below) and the done report.

Visual QA (required; HTTP 200 is not QA)

  • On 8081 (scripts/sandbox.sh) against a sqlite-backup copy of the live DB. Seed the COPY with one open alert whose verdicts carry a model, calls and cost, one resolved alert, and one blocked override with a reason. Never write to the live DB.
  • Screenshots of Home (Loops) and Models at 1400px and 800px wide, plus zoomed crops of each new card. Check truncation, alignment, empty states, and that Block -> Confirm -> Blocked list -> Unblock round-trips on screen.
  • Save screenshots under the scratch/evidence dir, not the repo.

Ground rules (lessons from Lift A)

  • Worktree already exists: /home/alee/Sources/6krrt-worktrees/no-progress-lift-a2, branch feat/no-progress-lift-a2 from origin/main 5aaf1e6 (includes #103). /start-work runs with --worktree set to it. Nothing in the main checkout.
  • Dispatch to Sisyphus-Junior categories or general. NEVER to oh-my-claudecode:* agents: they are pinned to Anthropic models this opencode cannot reach, and return "completed" with zero work. A subagent session with 0 tool calls did nothing.
  • One commit per todo; no squashed "full implementation" commits.
  • No test exemptions or skips for labelled cases. If a test cannot pass, stop and report why.
  • Lint: pinned uvx ruff@0.16.9, no NEW findings vs origin/main in touched files (recipe in plans/no-progress-lift-a.md Ground rules). Paste the gate's empty output in the done report.
  • Full pytest: record the origin/main baseline first; the only known failure is test_incumbent_routing::test_debug_log_emits_incumbent_identity.
  • Port 8080 is production: no requests to it, not even GETs.
  • Do not push, do not open a PR, do not remove the worktree. The done report lists commits, tests, screenshots and operator steps (router restart).
  • ASCII only; no middle dots. No prose paragraphs in the admin UI; mutually exclusive choices are radio-style.
  • Stop a worker whose context passes ~150k tokens and start a fresh one.