> How responses are checked. Back to [README](../README.md). ## Verification Pipeline Every completion passes through a two-layer check **before** routing learns from it and **without** adding to client-facing latency for the async layer. ### Structural Verification — `verification.py` Free, exact checks on what a model just returned — no code execution. Extracts fenced code blocks and validates them by **parsing**, not running: | Language | Check | Safeguard | |---|---|---| | Python (`python`, `py`, `python3`) | `ast.parse` | Syntax tree, no evaluation | | JSON (`json`, `jsonc`) | `json.loads` | Strict parse | | YAML (`yaml`, `yml`) | `yaml.safe_load` | No arbitrary objects | | Shell (`bash`, `sh`, `zsh`) | `bash -n` (caller's job) | Parse-only, never runs | | Other | `unverifiable` (not a failure) | Prose answers land here | Plus two universal checks: - **`finish_reason == "length"`** — decisive truncation, even if parsed content looks fine (the most dangerous case: a fragment that parses) - **Unterminated code fences** — the response ran out of tokens mid-block **Key design decisions:** - Never executes model output. The eval harness (`eval_proficiency.py`) does run generated code, but there the prompts are ones this project authored, so what comes back is bounded. Here the code is whatever the user asked for and could do anything. - `unverifiable` is **not** a failure. Most prose answers land there (correctly — there's nothing structural to check), and counting it as wrong would penalize models for the checker's limits. - Measured on real traffic: a wasted cloud completion costs about 52 local checks at 1,500 tokens and 139 at 4,000. A zero-cost check pays trivially. ### Local LLM Verification — `verification.py` (cont.) For answers nothing structural can judge (prose, reasoning, refusals), a local model spot-checks the response **after** it has gone back to the client, so its ~6 s never lands on anyone's latency. The prompt frames the local model as a judge: "Does this answer actually address the user's request?" and returns `{"ok": true|false, "reason": "..."}`. Gating and safeguards: - **Size threshold** — only responses ≥ 600 completion tokens trigger a local check (configurable via `verification.min_completion_tokens`). A check on a 193-token answer costs ~15% of the answer; at 1,500 tokens the break-even failure rate drops to ~1.9%. - **Middle elision** — if the answer is longer than the limit (default 4,000), the **middle** is elided, not the head, so the checker judges the real ending. An elision marker tells the checker not to flag it as a defect. - **Thinking disabled** — the local model is a reasoning model and will otherwise emit unbounded chain-of-thought. Verification uses Ollama's native endpoint with `think=False`, so it answers in ~25 tokens instead of burning through the budget on reasoning. - **Malfunction safety** — an empty or unparseable verdict produces no sample rather than a false failure. An empty response is settled by an if-statement, never asked of a model what code can decide. ### Feedback Loop — `feedback.py` Observation → learning. `feedback.py` folds verification failures into `proficiency` so routing improves on **your traffic**, not just the fixed 43-task benchmark: ```bash PYTHONPATH=src python -m feedback --dry-run # preview what would change PYTHONPATH=src python -m feedback # apply ``` Key behaviors: - **Only failures are folded in from the automatic checks.** A structural 'ok' means the code parsed, not that it was correct — scoring a parse-success as a proficiency 1.0 would flatten every score toward the ceiling. The coding categories already sit at 1.00 for every model; this would spread that flatness everywhere. **[`POST /outcome`](api.md) is the one exception** — a client-reported `ok: true` folds in as a real success, because the client actually ran the result. That's the only signal in the system where a passing result counts, not just a failing one. - **One 0.0 (or 1.0, for a client success) sample per report**, added to the running mean so the effect scales with rate rather than replacing the benchmark score outright. - **Idempotent** — `applied_at` marks consumed rows so re-running cannot penalize (or credit) a model repeatedly for the same response. - **Not the model's fault** — a client that sets `max_tokens=40` and gets a truncated answer caused that itself. Those failures are recorded in the verifications table but excluded from proficiency feedback. ### Retry Budget — `iteration.py` A verification failure can also buy a corrective attempt on the *same* request, before the response ever reaches the client. The budget is a function of tier and latency mode — interactive is capped below its tier's batch budget because every retry doubles time-to-answer, and in interactive use latency is itself a quality loss: | Tier | Batch retries | Interactive retries | |---|---|---| | 1 | 0 | 0 | | 2 | 1 | 1 | | 3 | 2 | 1 | The retry is matched to what actually failed, since the causes differ: - **`truncated`** — raise the token budget on the same model first; a different model would just run out too. If there's no cap left to raise, the model's own output ceiling is the wall, so escalate to a candidate that can emit more instead. - **`malformed`** — more tokens won't make unparseable output parse, so escalate straight to the next-ranked candidate. - **`ok` / `unverifiable`** — buys nothing. Retrying `unverifiable` would burn quota across the majority of ordinary prose traffic for no signal, since there was nothing wrong to correct. This only reaches the non-streaming path — once bytes have gone to a streaming client there's nothing left to take back. `POST /outcome` is the answer for streamed traffic: it arrives after the fact and works identically either way.