Files
6krrt/AGENTS.md
2026-10-04 01:14:32 -04:00

363 lines
19 KiB
Markdown

# AGENTS.md — working guide for AI agents on this repo
The three-tier docs:
- **`README.md`** — concise front door. Deep-dive module-by-module walkthroughs
are in **`docs/`** (13 reference files: `admin-portal.md`, `api.md`,
`architecture.md`, `clients.md`, `config-local-overlay.md`, `data-model.md`,
`evaluation.md`, `incidents.md`, `local-models.md`, `operations.md`,
`pinch.md`, `routing.md`, `verification.md`).
The README defers to `docs/` for module-level detail.
- **`CLAUDE.md`** — working state + immediate next steps + design rationale
(the one to trust on what is currently true).
- **`design/local-llm-model-router.md`** — architecture and rationale,
including parts still unbuilt.
This file is the quick-reference for an agent starting work: where things
live, what's safe to touch, and the conventions that aren't obvious from the
code.
## Worktrees: where dispatched work happens (read before editing anything)
**`/home/alee/Sources/6krrt` is the main checkout, and it is production.** The
live router restarts from its `config/config.yaml`, and the live `router.db` and
`config/config.local.yaml` live there. Its uncommitted files are the owner's.
Your session opens in that directory by default, even when your task belongs to
a worktree. So a task that names no worktree edits production files. That
happened on 2026-09-28: a worker rewrote the main checkout's live
`config.yaml`.
**If you are a worker:**
- Your task names a worktree under `/home/alee/Sources/6krrt-worktrees/<name>`.
`cd` there first, and read and edit only paths under it.
- Check where you are before your first edit:
`git -C <worktree> rev-parse --show-toplevel` must print the worktree, not
`/home/alee/Sources/6krrt`.
- If your task names no worktree, stop and ask the orchestrator. Do not default
to the main checkout.
- Never edit, commit, stash, reset or check out anything in the main checkout
unless your task explicitly says "main checkout".
**If you dispatch work (Atlas, any orchestrator):**
- Begin every `task()` prompt with this line:
`WORKTREE: <absolute worktree path>. cd there first; never edit under /home/alee/Sources/6krrt/.`
Include it even for a small or read-only task.
- Give every `task()` a `category` or `subagent_type`. A call with neither
creates no worker, yet still reports "completed".
- Never dispatch to `oh-my-claudecode:*` agents. They are pinned to Anthropic
models this opencode cannot reach, and return "completed" with zero work.
- Run one worker at a time per worktree, unless the plan says the files are
disjoint.
- Before ticking a todo, confirm its commit exists:
`git -C <worktree> log --oneline`.
Additionally:
- A worker, after committing, runs
`python3 scripts/verify_commit.py --repo <worktree> HEAD` and pastes its
output before reporting done.
- An orchestrator, before reporting a wave done, runs
`scripts/oc_dispatch_audit.py <session> --worktree <path>` and pastes its
output.
- A guardrail block is a rule, not an error to route around: fix the call,
do not work around the plugin.
- A test of code that calls an external service (Ollama, a provider, the
opencode API) fakes it at the HTTP boundary (`requests.post`, `fetch`),
never by replacing the project's own wrapper. On 2026-09-29 the
local_decision dispatcher bug (category names passed where option letters
were expected) passed tests that faked `classify_choice` and failed every
live call.
The Git hygiene section below says why the main checkout is shared.
## Stack snapshot
| Dimension | Value |
|---|---|
| Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) |
| Framework | FastAPI + uvicorn |
| Database | SQLite (`router.db`) |
| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
| Testing | `pytest`, 1923 tests, all offline (no provider or local-model calls) |
| Dependencies | Pinned. Bump deliberately, never use `>=` |
## Module map and file boundaries
### Dispatcher / service (I/O: DB + network)
| File | Role | Agent notes |
|---|---|---|
| `dispatcher.py` | FastAPI service: routes, calls providers, logs, streams | ~2500 LOC; only edit the specific function you need. **Must not import `metrics`** (circular). Imports `events` for SSE fan-out. |
| `metrics.py` | Read-only aggregations for `/health` and `/metrics` | Takes `(conn, cfg)` args — **must not import `dispatcher`** to avoid circular import. |
| `events.py` | In-memory decision-event broker | Pure stdlib (`queue`, `collections.deque`). Thread-safe. No Textual import. The SSE endpoint lives in `dispatcher.py` and calls `events.subscribe()`/`publish_decision()`. |
| `config.py` | Pydantic models + YAML loader | All config models inherit `StrictModel` (`extra="forbid"`). Unknown keys fail at load. |
| `capabilities.py` | Request-side capability detection (`tools`, `images`, `json_mode`, `reasoning`) | Reads from the OpenAI-format request body, not from the classifier. |
| `context_prune.py` | Relevance-based context pruning (`pinch`) | Ships **enabled** by default (`pinch.enabled: true`; `pinch.relevance.enabled: true`; budget-gated: the embed only fires when the conversation exceeds `budget_tokens`). Imports `PinchConfig` from `config.py` (no circular import: `config.py` doesn't import `context_prune`). |
| `logs.py` | Structured logging (logfmt, journald, ContextVar trace ids) | `logs.bind()` exists for StreamingResponse generators that lose the ContextVar. |
### Pure scoring/routing modules (no I/O)
| File | Role |
|---|---|
| `scoring.py` | `normalize_inverted` + weighted composite (quality-first, cost as tiebreak) |
| `routing.py` | Hard filters (`select_candidates`) + ranking (`rank_candidates`) |
| `tiering.py` | Pure tier resolver: 1=cheap+small, 2=mid, 3=frontier |
| `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite |
| `iteration.py` | Retry budget per tier, matching retry to failure kind |
### Admin portal
| File | Role |
|---|---|
| `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML |
| `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions |
| `admin/frontend/models.html` | Model availability table with override dropdown |
| `admin/frontend/decisions.html` | Full decision log with filter + search |
| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor |
| `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) |
When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card.
### TUI modules (import `textual`; never imported by the dispatch path)
| File | Role |
|---|---|
| `tui.py` | `DashboardApp` — main Textual app with CSS, bindings, compose, render |
| `tui_model.py` | Pure data layer: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal |
| `tui_sse.py` | Background-thread SSE consumer (`DecisionStream`): reconnects on failure, marshals decisions to UI thread via `call_from_thread` |
| `tui_screens.py` | `DecisionDetailScreen` — `ModalScreen` showing full decision JSON (Enter or `e` key) |
### Other entrypoints
| File | Role |
|---|---|
| `poller.py` | Fetches Neuralwatt catalog, upserts `models` table; also refreshes `provider='ollama-local'` rows each poll via `upsert_local_dispatch_models` |
| `seed_energy.py` | Reference workload sweep → `energy_observations` |
| `seed_local_dispatch_energy.py` | Local reference-shape sweep → per-token rates for `ollama-local` rows |
| `eval_proficiency.py` | Self-eval harness → `proficiency` table |
| `feedback.py` | Folds verification failures into `proficiency` |
| `tier.py` | DB tiering pass |
| `router_cli.py` | One-shot `/route` probe |
| `proficiency_store.py` | DB write path for `proficiency` |
## Conventions
### Code style
- The codebase uses `Optional[X]` (not `X | None`) and `dict` return types
throughout — match the existing style in the file you're editing.
- No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent.
- `except Exception` is acceptable at top-level boundaries with
`# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`).
- Config is strict: every knob belongs in `config/config.yaml`, not only in a
Pydantic default. A default the file never mentions is invisible to a tuner.
- Named constants use `Final` in new code (`events.py`); existing code is
inconsistent — don't refactor just for this.
### Import discipline
- `metrics.py` **must not import `dispatcher`** (circular import).
- `events.py` imports nothing from the project (pure stdlib).
- `tui_model.py` imports `requests` but **not `textual`** — it's the testable data layer.
- `tui_sse.py` imports `requests` but **not `textual`** — it's a thread worker.
- `tui.py` and `tui_screens.py` import `textual` — that's fine, they're TUI.
- The service dispatch path (`dispatcher.py` → provider call) **never touches
`textual`**.
### Database
- SQLite, `PRAGMA foreign_keys = ON`.
- `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with
`row_factory = sqlite3.Row`.
- Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
- Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`)
mirrors the schema for live DBs that predate a feature.
### Testing
- All tests are offline. No test calls a provider or local model.
- Pure modules take rows and config as arguments, so they're testable without DB/network.
- Test files that exercise FastAPI use `starlette.testclient.TestClient` with a
temp SQLite DB (`tests/test_metrics_endpoint.py`).
- TUI tests use `App.run_test()` with a stubbed fetcher and a `_no_real_network`
safety-net fixture that guards both `requests.get` and the SSE consumer.
- **SSE tests must not use `TestClient.stream()` on the infinite endpoint** — it
hangs. Test the generator directly (`dispatcher._decision_event_stream()`)
or stub the generator to a bounded one for header/route checks.
## Shared scripts (`scripts/`)
Mechanics every agent on this repo ends up doing by hand, wrapped once so
they are done the same way each time. Agent-agnostic -- plain scripts, plain
text output. See `scripts/README.md` for the reasoning behind each.
| script | what it does |
|---|---|
| `scripts/preflight.sh` | **run this before starting any multi-step work.** Prints the shared stash stack, then refuses a shared-checkout / wrong-branch / dirty / stale start |
| `scripts/sandbox.sh start\|stop\|status` | throwaway router on **8081** for exercising the admin portal against real data; never binds 8080, only kills a pid it started itself |
| `scripts/mkpr.sh <body.md> "<title>"` | cut a gitea PR with a long markdown body (`tea` has no `--description-file`), refusing an unpushed or non-ahead branch |
Prefer these over hand-rolling the equivalent command: each encodes a
constraint that is easy to get wrong once and expensive to get wrong twice
(production port, kill-by-name, WAL-safe DB copies, overlay protection).
## The rule that matters most
**When a tracker, hook, test, or the working tree disagrees with your own
account of what you did, assume YOUR ACCOUNT is wrong first.** Check the
artifact before concluding the tool is broken, and never modify tracking state
to stop a signal you have not disproven.
Both of the expensive failures on 2026-09-10 had one shape:
```
external signal contradicts internal belief
-> conclude the tool is broken
-> modify state to silence it
```
- A clean working tree contradicted "I made edits", so the conclusion was
"subagent edits are not persisting" and subagents were re-fired twice. The
work was in `stash@{0}` the whole time. The tree was telling the truth.
- A continuation hook contradicted "the work is complete", so the conclusion
was "the plan tracker cannot count" and the proposed fix was to clear the
field the hook reads. Three of seven todos were marked done and had never
been started -- two of them were not in the PR's diff at all. The hook was
telling the truth.
The tracker being genuinely buggy does not make the signal wrong. Fix the
mechanism AND correct the state it was misreporting; a fixed parser reading
7/7 against work that is 4/7 done is worse than a broken one reading 0/7,
because it turns a loud wrong signal into a quiet one.
## Git hygiene
**If the working tree looks unexpectedly clean, run `git stash list` BEFORE
concluding anything about your tooling.**
That single line is here because of a real run on 2026-09-10. An executor
stashed its own partial work with a descriptive message, came back to a clean
tree, concluded "subagent edits are not persisting", and re-fired subagents
twice against that theory. The work was sitting in `stash@{0}` the whole time.
An unexpectedly clean tree is far more often a stash than a broken editor.
`scripts/preflight.sh` prints the stack unconditionally for this reason.
**Never `git stash` or `git stash pop` bare.** This repo has ~20 worktrees on
this machine and they all share ONE stash stack. Another session pushing an
entry shifts every index, and `pop` deletes on success -- so a wrong index is
destructive and unrecoverable. Always:
```bash
git stash push -u -m '<unique-tag>' # push with a tag you can find
git stash list --format='%gd %H %gs' # capture the SHA
git stash apply <SHA> # apply by SHA, never pop by index
```
**Work in a worktree, not the shared checkout.** `/home/alee/Sources/6krrt` is
what the operator uses from their own terminal. Mutating it collides with
their `git pull` mid-run, and a leftover branch there sends your commits
somewhere nobody is looking -- in the same 2026-09-10 run, execution happened
on `docs/admin-screenshots`, a merged PR's branch, against a tree missing two
already-merged PRs.
**The plan is the contract.** If you believe an approved plan item is wrong,
STOP and say so. Do not implement the alternative. The same run re-derived a
design that had been explicitly reviewed and rejected, because the rejected
version sounded more reasonable in isolation -- the reasoning against it was
in the plan's own "Must NOT" list.
## How to run things
```bash
# Setup
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own
# model traffic goes through it (see opencode.json's baseURL) and
# `Restart=always` resurrects it ~5s after any kill — so never
# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the
# port". That fights the supervisor, and on this repo it can cut off your
# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the
# process forever" section for the incident this note comes from.
#
# To pick up a dispatcher.py change on the real instance:
systemctl --user restart llm-router.service
# For an ad hoc/manual run — iterating with --reload, a throwaway instance
# for Playwright smoke tests against the admin frontend, anything that
# isn't "use the real router" — bind a different port so it can't collide:
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
# Run the TUI (service must be running)
PYTHONPATH=src python -m tui
# Run the full test suite
python -m pytest
# Quick routing probe (no spend)
PYTHONPATH=src python -m router_cli "Refactor this Django view"
# Populate the catalog
PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
```
## Gitea / `tea` CLI (not `gh`)
This repo is hosted on **self-hosted Gitea** (`git.adlee.work/alee/6krrt`).
The GitHub CLI `gh` is **not installed** and will fail — use `tea` instead.
- **Login:** `tea login` as `alee`. The `alee` login is NOT the default tea
login — always pass `--remote origin` so tea resolves the right context.
- **Remote:** `origin` resolves to `git.adlee.work/alee/6krrt`.
- **List open PRs:** `tea pulls --remote origin`
- **PR metadata (JSON):**
```
tea pulls <PR_INDEX> --remote origin --fields index,state,draft,title,mergeable,base,head,body -o json
```
Use the `base` and `head` from this output for local diff commands.
- **PR creation / update / maintenance:** use the `tea pr` family (e.g. `tea pr create --base main --title ... --description ...`). Confirm against `tea pr --help` for the current flag set.
- **Post a comment to a PR** (uses the `tea api` proxy to Gitea's REST API):
```
tea api --remote origin "/repos/{owner}/{repo}/issues/<PR_INDEX>/comments" -F body=@-
```
- **CRITICAL:** use `-F` (typed field), **not** `-f`. `-f key=@file` writes the literal string `@file`; only `-F` reads stdin or the contents of a file.
- The endpoint **must** include the `/repos/` prefix — omitting it returns a 404.
- `tea api` already prints JSON to stdout — do **not** pass `-o json` on top of that (you would write the body to a file literally named `json`).
- Fallback if stdin is awkward: `-F body=@/tmp/review-body.md`.
- **Getting a PR diff:** prefer local `git diff origin/<base>...origin/<head>`
(base and head from the PR JSON). In tea v0.14.1 the `diff`/`patch` fields
return empty even when requested — do not rely on them.
- **Gitea source links** (not GitHub `blob` links):
`https://git.adlee.work/alee/6krrt/src/commit/<FULL-SHA>/<path>#L<start>-L<end>`
(segment is `src/commit`, not `blob`).
- **Code-review plugin:** this repo ships a project-local fork of the code-review
plugin (`code-review-tea`) that already wires these `tea` calls. Prefer
invoking it (e.g. `/code-review-tea:code-review`) over hand-rolling.
## What's NOT built yet (open items)
1. **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty.
2. **Three models unsettled** on eco stability (attribution noise).
3. **Retry does not reach streaming** — `POST /outcome` is the answer for streamed traffic.
4. **Local energy not on the ledger** — local classifier/verifier electricity is unmeasured.
5. **Session-directory attribution picks the wrong directory** — heuristic resolves to a dependency's source dir instead of the project being edited.
See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
## After a code change
- Run `python -m pytest` — 1923 tests, ~35s.
- If you changed `dispatcher.py`, restart the systemd service:
`systemctl --user restart llm-router.service` (it doesn't auto-reload code).
- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.
- Check `lsp_diagnostics` on changed files.
- Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`.