6.5 KiB
Eval harness and classifier notes. Back to README.
Self-Eval Harness (eval_proficiency.py)
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
PYTHONPATH=src python -m eval_proficiency --categories coding_general
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
- Runs every task through the target provider, scores it, writes to
proficiency. - Scores accumulate (running mean), so repeated runs tighten estimates.
propagate_to_variantscopies evaluated scores to equivalent serving variants.- Safety: code tasks run in a temp directory with a 15s wall-clock timeout — bounded isolation, not a container.
- Judge tasks skip the model being judged (avoids self-scoring bias).
- Local
ollama-localidentities resolve their endpoint fromcfg.local_dispatch_modelsand write proficiency rows withprovider='ollama-local'. Judges remain cloud-only.
Encoder classifier harness (eval_classifier.py)
PYTHONPATH=src python src/eval_classifier.py # run against the real model
PYTHONPATH=src python src/eval_classifier.py --dry-run # plan only, no model load
- Measures zero-shot classification accuracy of
local_encoder.classify_zero_shotover the 46-task eval set (evals/tasks.yamlminus the 11tool_use_agenticrows) with three noise variants per prompt (clean, short-noise, long-noise). - Prints top-1 accuracy, a full confusion matrix, per-category precision/recall,
confidence distribution split by correct/incorrect, and the per-cell
blended_scoregap from theproficiencytable (whenrouter.dbis present). - Standalone: imports
transformers/torchonly when actually running against the model, and lives underrequirements-encoder.txt, so the router's dispatch path never touches them.
Proficiency
The harness is a prior, not the routing score. The benchmark produces a
best-guess quality estimate before real traffic arrives. Routing turns that
candidate score into an expected pass rate on real traffic, then updates it
with client outcomes via POST /outcome and feedback.py. The benchmark is
what you have when no one has reported whether the answer worked yet; it is
not the final truth.
Two new categories for local dispatch
The shipped qwen2.5-coder-router:14b row is eligible for file_summarization
and diff_checking:
| Category | Kind | Why |
|---|---|---|
diff_checking |
exact |
A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. |
file_summarization |
judge |
A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. |
Both categories run through the same harness as the cloud rows. The only
difference is that an ollama-local identity resolves its Ollama endpoint from
cfg.local_dispatch_models and its cost model is seeded separately (see below).
Seeding local-dispatch costs
Cloud rows arrive with catalog prices; local rows start with NULL token rates
and must be measured. seed_local_dispatch_energy.py runs a small reference
sweep (sum_small, sum_large, diff_small, long_answer), samples GPU draw
with nvidia-smi, and derives per-token USD rates via through-origin OLS on
energy_kwh ~ a*prompt_tokens + b*completion_tokens at the user's tariff.
An intercept is intentionally excluded. Routing's cost model is purely
per-token, so fitting a fixed overhead would add a term estimated_cost()
cannot express. Consistency with the ledger beats theoretical purity, and that
choice is stated rather than hidden. Repeat the sweep with
--samples 5 or more to tighten the fit; the resulting rates are written to
models.cost_per_1m_prompt / cost_per_1m_completion for provider='ollama-local'
rows and are never overwritten by later poller runs.
Classifier Reliability Notes
The classifier is the one blocking LLM call on the request path, so it sets
the latency floor for every routed request. classifier.base_url /
api_key_env / model accept any OpenAI-compatible endpoint — a local
Ollama, an Ollama on another machine across a VPN, or a cloud model.
Pick a non-reasoning model, and judge it by the tail, not the mean. The
classifier's entire output is ~45 tokens of JSON, and a reasoning model
spends its budget getting there instead. This project has swapped classifiers
twice on exactly that evidence — see
local-models.md
for the full three-generation comparison table rather than duplicating it
here; the current default is qwen2.5-coder-router:14b.
Four settings keep a slow or misbehaving classifier from cascading into worse failures — full rationale for each in local-models.md:
| setting | value | one-line reason |
|---|---|---|
max_retries |
0 |
the SDK's own retries would triple the timeout silently |
max_output_tokens |
1024 |
bounds a reasoning model's chain of thought |
max_input_chars |
8000 |
the classifier needs the instruction, not the document |
temperature |
0 |
reproducible classification, not a coin flip per request |
Two more settings are graceful degradation rather than prevention:
- Fallback tier/category (
fallback_tier: 2/fallback_category: general_chat) is what a timeout, error, or unparseable output degrades to, instead of a 502/503. A coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. - Session classification cache (
session_cache.enabledinconfig/config.yaml, off by default) remembers the lasttask_category/task_tierdecision per session forstaleness_seconds(default 1200), so a long agent session skips the classifier round-trip on every turn. It only ever short-circuits the classifier call — capability flags (tools/images/json) are still read fresh from each request body, fallback classifications are never cached, and there is no persistence beyond process memory. A cache-hit decision is observable viaroute_decisions.classification_source = "cached".