Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
120 lines
5.8 KiB
Markdown
120 lines
5.8 KiB
Markdown
> How responses are checked. Back to [README](../README.md).
|
|
|
|
## Verification Pipeline
|
|
|
|
Every completion passes through a two-layer check **before** routing learns from
|
|
it and **without** adding to client-facing latency for the async layer.
|
|
|
|
### Structural Verification — `verification.py`
|
|
|
|
Free, exact checks on what a model just returned — no code execution.
|
|
Extracts fenced code blocks and validates them by **parsing**, not running:
|
|
|
|
| Language | Check | Safeguard |
|
|
|---|---|---|
|
|
| Python (`python`, `py`, `python3`) | `ast.parse` | Syntax tree, no evaluation |
|
|
| JSON (`json`, `jsonc`) | `json.loads` | Strict parse |
|
|
| YAML (`yaml`, `yml`) | `yaml.safe_load` | No arbitrary objects |
|
|
| Shell (`bash`, `sh`, `zsh`) | `bash -n` (caller's job) | Parse-only, never runs |
|
|
| Other | `unverifiable` (not a failure) | Prose answers land here |
|
|
|
|
Plus two universal checks:
|
|
- **`finish_reason == "length"`** — decisive truncation, even if parsed content
|
|
looks fine (the most dangerous case: a fragment that parses)
|
|
- **Unterminated code fences** — the response ran out of tokens mid-block
|
|
|
|
**Key design decisions:**
|
|
- Never executes model output. The eval harness (`eval_proficiency.py`) does
|
|
run generated code, but there the prompts are ones this project authored,
|
|
so what comes back is bounded. Here the code is whatever the user asked for
|
|
and could do anything.
|
|
- `unverifiable` is **not** a failure. Most prose answers land there (correctly
|
|
— there's nothing structural to check), and counting it as wrong would
|
|
penalize models for the checker's limits.
|
|
- Measured on real traffic: a wasted cloud completion costs about 52 local
|
|
checks at 1,500 tokens and 139 at 4,000. A zero-cost check pays trivially.
|
|
|
|
### Local LLM Verification — `verification.py` (cont.)
|
|
|
|
For answers nothing structural can judge (prose, reasoning, refusals), a local
|
|
model spot-checks the response **after** it has gone back to the client, so its
|
|
~6 s never lands on anyone's latency.
|
|
|
|
The prompt frames the local model as a judge: "Does this answer actually
|
|
address the user's request?" and returns `{"ok": true|false, "reason": "..."}`.
|
|
|
|
Gating and safeguards:
|
|
- **Size threshold** — only responses ≥ 600 completion tokens trigger a local
|
|
check (configurable via `verification.min_completion_tokens`). A check on
|
|
a 193-token answer costs ~15% of the answer; at 1,500 tokens the break-even
|
|
failure rate drops to ~1.9%.
|
|
- **Middle elision** — if the answer is longer than the limit (default 4,000),
|
|
the **middle** is elided, not the head, so the checker judges the real ending.
|
|
An elision marker tells the checker not to flag it as a defect.
|
|
- **Thinking disabled** — the local model is a reasoning model and will otherwise
|
|
emit unbounded chain-of-thought. Verification uses Ollama's native endpoint
|
|
with `think=False`, so it answers in ~25 tokens instead of burning through
|
|
the budget on reasoning.
|
|
- **Malfunction safety** — an empty or unparseable verdict produces no sample
|
|
rather than a false failure. An empty response is settled by an if-statement,
|
|
never asked of a model what code can decide.
|
|
|
|
### Feedback Loop — `feedback.py`
|
|
|
|
Observation → learning. `feedback.py` folds verification failures into
|
|
`proficiency` so routing improves on **your traffic**, not just the fixed
|
|
43-task benchmark:
|
|
|
|
```bash
|
|
PYTHONPATH=src python -m feedback --dry-run # preview what would change
|
|
PYTHONPATH=src python -m feedback # apply
|
|
```
|
|
|
|
Key behaviors:
|
|
- **Only failures are folded in from the automatic checks.** A structural 'ok'
|
|
means the code parsed, not that it was correct — scoring a parse-success as
|
|
a proficiency 1.0 would flatten every score toward the ceiling. The coding
|
|
categories already sit at 1.00 for every model; this would spread that
|
|
flatness everywhere. **[`POST /outcome`](api.md) is the one exception** —
|
|
a client-reported `ok: true` folds in as a real success, because the client
|
|
actually ran the result. That's the only signal in the system where a
|
|
passing result counts, not just a failing one.
|
|
- **One 0.0 (or 1.0, for a client success) sample per report**, added to the
|
|
running mean so the effect scales with rate rather than replacing the
|
|
benchmark score outright.
|
|
- **Idempotent** — `applied_at` marks consumed rows so re-running cannot
|
|
penalize (or credit) a model repeatedly for the same response.
|
|
- **Not the model's fault** — a client that sets `max_tokens=40` and gets a
|
|
truncated answer caused that itself. Those failures are recorded in the
|
|
verifications table but excluded from proficiency feedback.
|
|
|
|
### Retry Budget — `iteration.py`
|
|
|
|
A verification failure can also buy a corrective attempt on the *same*
|
|
request, before the response ever reaches the client. The budget is a
|
|
function of tier and latency mode — interactive is capped below its tier's
|
|
batch budget because every retry doubles time-to-answer, and in interactive
|
|
use latency is itself a quality loss:
|
|
|
|
| Tier | Batch retries | Interactive retries |
|
|
|---|---|---|
|
|
| 1 | 0 | 0 |
|
|
| 2 | 1 | 1 |
|
|
| 3 | 2 | 1 |
|
|
|
|
The retry is matched to what actually failed, since the causes differ:
|
|
- **`truncated`** — raise the token budget on the same model first; a
|
|
different model would just run out too. If there's no cap left to raise,
|
|
the model's own output ceiling is the wall, so escalate to a candidate that
|
|
can emit more instead.
|
|
- **`malformed`** — more tokens won't make unparseable output parse, so
|
|
escalate straight to the next-ranked candidate.
|
|
- **`ok` / `unverifiable`** — buys nothing. Retrying `unverifiable` would burn
|
|
quota across the majority of ordinary prose traffic for no signal, since
|
|
there was nothing wrong to correct.
|
|
|
|
This only reaches the non-streaming path — once bytes have gone to a
|
|
streaming client there's nothing left to take back. `POST /outcome` is the
|
|
answer for streamed traffic: it arrives after the fact and works identically
|
|
either way.
|