> Eval harness and classifier notes. Back to [README](../README.md). ## Self-Eval Harness (`eval_proficiency.py`) ```bash PYTHONPATH=src python -m eval_proficiency # every routable model × every task PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models PYTHONPATH=src python -m eval_proficiency --categories coding_general PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking PYTHONPATH=src python -m eval_proficiency --dry-run # plan only ``` - Runs every task through the target provider, scores it, writes to `proficiency`. - Scores **accumulate** (running mean), so repeated runs tighten estimates. - `propagate_to_variants` copies evaluated scores to equivalent serving variants. - Safety: code tasks run in a temp directory with a 15s wall-clock timeout — bounded isolation, not a container. - Judge tasks skip the model being judged (avoids self-scoring bias). - Local `ollama-local` identities resolve their endpoint from `cfg.local_dispatch_models` and write proficiency rows with `provider='ollama-local'`. Judges remain cloud-only. ## Encoder classifier harness (`eval_classifier.py`) ```bash PYTHONPATH=src python src/eval_classifier.py # run against the real model PYTHONPATH=src python src/eval_classifier.py --dry-run # plan only, no model load ``` - Measures zero-shot classification accuracy of `local_encoder.classify_zero_shot` over the 46-task eval set (`evals/tasks.yaml` minus the 11 `tool_use_agentic` rows) with three noise variants per prompt (clean, short-noise, long-noise). - Prints top-1 accuracy, a full confusion matrix, per-category precision/recall, confidence distribution split by correct/incorrect, and the per-cell `blended_score` gap from the `proficiency` table (when `router.db` is present). - Standalone: imports `transformers`/`torch` only when actually running against the model, and lives under `requirements-encoder.txt`, so the router's dispatch path never touches them. ## Proficiency The harness is a **prior**, not the routing score. The benchmark produces a best-guess quality estimate before real traffic arrives. Routing turns that candidate score into an expected pass rate on real traffic, then updates it with client outcomes via `POST /outcome` and `feedback.py`. The benchmark is what you have when no one has reported whether the answer worked yet; it is not the final truth. ## Two new categories for local dispatch The shipped `qwen2.5-coder-router:14b` row is eligible for `file_summarization` and `diff_checking`: | Category | Kind | Why | |---|---|---| | `diff_checking` | `exact` | A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. | | `file_summarization` | `judge` | A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. | Both categories run through the same harness as the cloud rows. The only difference is that an `ollama-local` identity resolves its Ollama endpoint from `cfg.local_dispatch_models` and its cost model is seeded separately (see below). ## Seeding local-dispatch costs Cloud rows arrive with catalog prices; local rows start with NULL token rates and must be measured. `seed_local_dispatch_energy.py` runs a small reference sweep (`sum_small`, `sum_large`, `diff_small`, `long_answer`), samples GPU draw with `nvidia-smi`, and derives per-token USD rates via through-origin OLS on `energy_kwh ~ a*prompt_tokens + b*completion_tokens` at the user's tariff. An intercept is intentionally excluded. Routing's cost model is purely per-token, so fitting a fixed overhead would add a term `estimated_cost()` cannot express. Consistency with the ledger beats theoretical purity, and that choice is stated rather than hidden. Repeat the sweep with `--samples 5` or more to tighten the fit; the resulting rates are written to `models.cost_per_1m_prompt` / `cost_per_1m_completion` for `provider='ollama-local'` rows and are never overwritten by later poller runs. ## Classifier Reliability Notes The classifier is the one blocking LLM call on the request path, so it sets the latency floor for every routed request. `classifier.base_url` / `api_key_env` / `model` accept any OpenAI-compatible endpoint — a local Ollama, an Ollama on another machine across a VPN, or a cloud model. **Pick a non-reasoning model, and judge it by the tail, not the mean.** The classifier's entire output is ~45 tokens of JSON, and a reasoning model spends its budget getting there instead. This project has swapped classifiers twice on exactly that evidence — see [local-models.md](local-models.md#three-generations-and-the-one-lesson-that-outlived-all-of-them) for the full three-generation comparison table rather than duplicating it here; the current default is `qwen2.5-coder-router:14b`. Four settings keep a slow or misbehaving classifier from cascading into worse failures — full rationale for each in [local-models.md](local-models.md#the-four-settings-and-why-each-exists): | setting | value | one-line reason | |---|---|---| | `max_retries` | `0` | the SDK's own retries would triple the timeout silently | | `max_output_tokens` | `1024` | bounds a reasoning model's chain of thought | | `max_input_chars` | `8000` | the classifier needs the instruction, not the document | | `temperature` | `0` | reproducible classification, not a coin flip per request | Two more settings are graceful degradation rather than prevention: - **Fallback tier/category** (`fallback_tier: 2` / `fallback_category: general_chat`) is what a timeout, error, or unparseable output degrades to, instead of a 502/503. A coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. - **Session classification cache** (`session_cache.enabled` in `config/config.yaml`, off by default) remembers the last `task_category`/`task_tier` decision per session for `staleness_seconds` (default 1200), so a long agent session skips the classifier round-trip on every turn. It only ever short-circuits the classifier call — capability flags (tools/images/json) are still read fresh from each request body, fallback classifications are never cached, and there is no persistence beyond process memory. A cache-hit decision is observable via `route_decisions.classification_source = "cached"`.