Files
6krrt/docs/evaluation.md

120 lines
6.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
> Eval harness and classifier notes. Back to [README](../README.md).
## Self-Eval Harness (`eval_proficiency.py`)
```bash
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
PYTHONPATH=src python -m eval_proficiency --categories coding_general
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
```
- Runs every task through the target provider, scores it, writes to `proficiency`.
- Scores **accumulate** (running mean), so repeated runs tighten estimates.
- `propagate_to_variants` copies evaluated scores to equivalent serving variants.
- Safety: code tasks run in a temp directory with a 15s wall-clock timeout —
bounded isolation, not a container.
- Judge tasks skip the model being judged (avoids self-scoring bias).
- Local `ollama-local` identities resolve their endpoint from `cfg.local_dispatch_models`
and write proficiency rows with `provider='ollama-local'`. Judges remain cloud-only.
## Encoder classifier harness (`eval_classifier.py`)
```bash
PYTHONPATH=src python src/eval_classifier.py # run against the real model
PYTHONPATH=src python src/eval_classifier.py --dry-run # plan only, no model load
```
- Measures zero-shot classification accuracy of `local_encoder.classify_zero_shot`
over the 46-task eval set (`evals/tasks.yaml` minus the 11 `tool_use_agentic`
rows) with three noise variants per prompt (clean, short-noise, long-noise).
- Prints top-1 accuracy, a full confusion matrix, per-category precision/recall,
confidence distribution split by correct/incorrect, and the per-cell
`blended_score` gap from the `proficiency` table (when `router.db` is present).
- Standalone: imports `transformers`/`torch` only when actually running against
the model, and lives under `requirements-encoder.txt`, so the router's dispatch
path never touches them.
## Proficiency
The harness is a **prior**, not the routing score. The benchmark produces a
best-guess quality estimate before real traffic arrives. Routing turns that
candidate score into an expected pass rate on real traffic, then updates it
with client outcomes via `POST /outcome` and `feedback.py`. The benchmark is
what you have when no one has reported whether the answer worked yet; it is
not the final truth.
## Two new categories for local dispatch
The shipped `qwen2.5-coder-router:14b` row is eligible for `file_summarization`
and `diff_checking`:
| Category | Kind | Why |
|---|---|---|
| `diff_checking` | `exact` | A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. |
| `file_summarization` | `judge` | A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. |
Both categories run through the same harness as the cloud rows. The only
difference is that an `ollama-local` identity resolves its Ollama endpoint from
`cfg.local_dispatch_models` and its cost model is seeded separately (see below).
## Seeding local-dispatch costs
Cloud rows arrive with catalog prices; local rows start with NULL token rates
and must be measured. `seed_local_dispatch_energy.py` runs a small reference
sweep (`sum_small`, `sum_large`, `diff_small`, `long_answer`), samples GPU draw
with `nvidia-smi`, and derives per-token USD rates via through-origin OLS on
`energy_kwh ~ a*prompt_tokens + b*completion_tokens` at the user's tariff.
An intercept is intentionally excluded. Routing's cost model is purely
per-token, so fitting a fixed overhead would add a term `estimated_cost()`
cannot express. Consistency with the ledger beats theoretical purity, and that
choice is stated rather than hidden. Repeat the sweep with
`--samples 5` or more to tighten the fit; the resulting rates are written to
`models.cost_per_1m_prompt` / `cost_per_1m_completion` for `provider='ollama-local'`
rows and are never overwritten by later poller runs.
## Classifier Reliability Notes
The classifier is the one blocking LLM call on the request path, so it sets
the latency floor for every routed request. `classifier.base_url` /
`api_key_env` / `model` accept any OpenAI-compatible endpoint — a local
Ollama, an Ollama on another machine across a VPN, or a cloud model.
**Pick a non-reasoning model, and judge it by the tail, not the mean.** The
classifier's entire output is ~45 tokens of JSON, and a reasoning model
spends its budget getting there instead. This project has swapped classifiers
twice on exactly that evidence — see
[local-models.md](local-models.md#three-generations-and-the-one-lesson-that-outlived-all-of-them)
for the full three-generation comparison table rather than duplicating it
here; the current default is `qwen2.5-coder-router:14b`.
Four settings keep a slow or misbehaving classifier from cascading into
worse failures — full rationale for each in
[local-models.md](local-models.md#the-four-settings-and-why-each-exists):
| setting | value | one-line reason |
|---|---|---|
| `max_retries` | `0` | the SDK's own retries would triple the timeout silently |
| `max_output_tokens` | `1024` | bounds a reasoning model's chain of thought |
| `max_input_chars` | `8000` | the classifier needs the instruction, not the document |
| `temperature` | `0` | reproducible classification, not a coin flip per request |
Two more settings are graceful degradation rather than prevention:
- **Fallback tier/category** (`fallback_tier: 2` / `fallback_category:
general_chat`) is what a timeout, error, or unparseable output degrades to,
instead of a 502/503. A coding agent would rather have a mid-tier answer
than an error. Escalation deliberately skips fallbacks, so an unavailable
local model does not silently promote every request to the frontier tier.
- **Session classification cache** (`session_cache.enabled` in
`config/config.yaml`, off by default) remembers the last
`task_category`/`task_tier` decision per session for `staleness_seconds`
(default 1200), so a long agent session skips the classifier round-trip on
every turn. It only ever short-circuits the classifier call — capability
flags (tools/images/json) are still read fresh from each request body,
fallback classifications are never cached, and there is no persistence
beyond process memory. A cache-hit decision is observable via
`route_decisions.classification_source = "cached"`.