120 lines
6.5 KiB
Markdown
120 lines
6.5 KiB
Markdown
> Eval harness and classifier notes. Back to [README](../README.md).
|
||
|
||
## Self-Eval Harness (`eval_proficiency.py`)
|
||
|
||
```bash
|
||
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
|
||
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
|
||
PYTHONPATH=src python -m eval_proficiency --categories coding_general
|
||
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
|
||
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
|
||
```
|
||
|
||
- Runs every task through the target provider, scores it, writes to `proficiency`.
|
||
- Scores **accumulate** (running mean), so repeated runs tighten estimates.
|
||
- `propagate_to_variants` copies evaluated scores to equivalent serving variants.
|
||
- Safety: code tasks run in a temp directory with a 15s wall-clock timeout —
|
||
bounded isolation, not a container.
|
||
- Judge tasks skip the model being judged (avoids self-scoring bias).
|
||
- Local `ollama-local` identities resolve their endpoint from `cfg.local_dispatch_models`
|
||
and write proficiency rows with `provider='ollama-local'`. Judges remain cloud-only.
|
||
|
||
## Encoder classifier harness (`eval_classifier.py`)
|
||
|
||
```bash
|
||
PYTHONPATH=src python src/eval_classifier.py # run against the real model
|
||
PYTHONPATH=src python src/eval_classifier.py --dry-run # plan only, no model load
|
||
```
|
||
|
||
- Measures zero-shot classification accuracy of `local_encoder.classify_zero_shot`
|
||
over the 46-task eval set (`evals/tasks.yaml` minus the 11 `tool_use_agentic`
|
||
rows) with three noise variants per prompt (clean, short-noise, long-noise).
|
||
- Prints top-1 accuracy, a full confusion matrix, per-category precision/recall,
|
||
confidence distribution split by correct/incorrect, and the per-cell
|
||
`blended_score` gap from the `proficiency` table (when `router.db` is present).
|
||
- Standalone: imports `transformers`/`torch` only when actually running against
|
||
the model, and lives under `requirements-encoder.txt`, so the router's dispatch
|
||
path never touches them.
|
||
|
||
## Proficiency
|
||
|
||
The harness is a **prior**, not the routing score. The benchmark produces a
|
||
best-guess quality estimate before real traffic arrives. Routing turns that
|
||
candidate score into an expected pass rate on real traffic, then updates it
|
||
with client outcomes via `POST /outcome` and `feedback.py`. The benchmark is
|
||
what you have when no one has reported whether the answer worked yet; it is
|
||
not the final truth.
|
||
|
||
## Two new categories for local dispatch
|
||
|
||
The shipped `qwen2.5-coder-router:14b` row is eligible for `file_summarization`
|
||
and `diff_checking`:
|
||
|
||
| Category | Kind | Why |
|
||
|---|---|---|
|
||
| `diff_checking` | `exact` | A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. |
|
||
| `file_summarization` | `judge` | A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. |
|
||
|
||
Both categories run through the same harness as the cloud rows. The only
|
||
difference is that an `ollama-local` identity resolves its Ollama endpoint from
|
||
`cfg.local_dispatch_models` and its cost model is seeded separately (see below).
|
||
|
||
## Seeding local-dispatch costs
|
||
|
||
Cloud rows arrive with catalog prices; local rows start with NULL token rates
|
||
and must be measured. `seed_local_dispatch_energy.py` runs a small reference
|
||
sweep (`sum_small`, `sum_large`, `diff_small`, `long_answer`), samples GPU draw
|
||
with `nvidia-smi`, and derives per-token USD rates via through-origin OLS on
|
||
`energy_kwh ~ a*prompt_tokens + b*completion_tokens` at the user's tariff.
|
||
|
||
An intercept is intentionally excluded. Routing's cost model is purely
|
||
per-token, so fitting a fixed overhead would add a term `estimated_cost()`
|
||
cannot express. Consistency with the ledger beats theoretical purity, and that
|
||
choice is stated rather than hidden. Repeat the sweep with
|
||
`--samples 5` or more to tighten the fit; the resulting rates are written to
|
||
`models.cost_per_1m_prompt` / `cost_per_1m_completion` for `provider='ollama-local'`
|
||
rows and are never overwritten by later poller runs.
|
||
|
||
## Classifier Reliability Notes
|
||
|
||
The classifier is the one blocking LLM call on the request path, so it sets
|
||
the latency floor for every routed request. `classifier.base_url` /
|
||
`api_key_env` / `model` accept any OpenAI-compatible endpoint — a local
|
||
Ollama, an Ollama on another machine across a VPN, or a cloud model.
|
||
|
||
**Pick a non-reasoning model, and judge it by the tail, not the mean.** The
|
||
classifier's entire output is ~45 tokens of JSON, and a reasoning model
|
||
spends its budget getting there instead. This project has swapped classifiers
|
||
twice on exactly that evidence — see
|
||
[local-models.md](local-models.md#three-generations-and-the-one-lesson-that-outlived-all-of-them)
|
||
for the full three-generation comparison table rather than duplicating it
|
||
here; the current default is `qwen2.5-coder-router:14b`.
|
||
|
||
Four settings keep a slow or misbehaving classifier from cascading into
|
||
worse failures — full rationale for each in
|
||
[local-models.md](local-models.md#the-four-settings-and-why-each-exists):
|
||
|
||
| setting | value | one-line reason |
|
||
|---|---|---|
|
||
| `max_retries` | `0` | the SDK's own retries would triple the timeout silently |
|
||
| `max_output_tokens` | `1024` | bounds a reasoning model's chain of thought |
|
||
| `max_input_chars` | `8000` | the classifier needs the instruction, not the document |
|
||
| `temperature` | `0` | reproducible classification, not a coin flip per request |
|
||
|
||
Two more settings are graceful degradation rather than prevention:
|
||
|
||
- **Fallback tier/category** (`fallback_tier: 2` / `fallback_category:
|
||
general_chat`) is what a timeout, error, or unparseable output degrades to,
|
||
instead of a 502/503. A coding agent would rather have a mid-tier answer
|
||
than an error. Escalation deliberately skips fallbacks, so an unavailable
|
||
local model does not silently promote every request to the frontier tier.
|
||
- **Session classification cache** (`session_cache.enabled` in
|
||
`config/config.yaml`, off by default) remembers the last
|
||
`task_category`/`task_tier` decision per session for `staleness_seconds`
|
||
(default 1200), so a long agent session skips the classifier round-trip on
|
||
every turn. It only ever short-circuits the classifier call — capability
|
||
flags (tools/images/json) are still read fresh from each request body,
|
||
fallback classifications are never cached, and there is no persistence
|
||
beyond process memory. A cache-hit decision is observable via
|
||
`route_decisions.classification_source = "cached"`.
|