Files
6krrt/docs/evaluation.md

6.5 KiB
Raw Permalink Blame History

Eval harness and classifier notes. Back to README.

Self-Eval Harness (eval_proficiency.py)

PYTHONPATH=src python -m eval_proficiency                      # every routable model × every task
PYTHONPATH=src python -m eval_proficiency --models kimi-k3     # subset of models
PYTHONPATH=src python -m eval_proficiency --categories coding_general
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
PYTHONPATH=src python -m eval_proficiency --dry-run            # plan only
  • Runs every task through the target provider, scores it, writes to proficiency.
  • Scores accumulate (running mean), so repeated runs tighten estimates.
  • propagate_to_variants copies evaluated scores to equivalent serving variants.
  • Safety: code tasks run in a temp directory with a 15s wall-clock timeout — bounded isolation, not a container.
  • Judge tasks skip the model being judged (avoids self-scoring bias).
  • Local ollama-local identities resolve their endpoint from cfg.local_dispatch_models and write proficiency rows with provider='ollama-local'. Judges remain cloud-only.

Encoder classifier harness (eval_classifier.py)

PYTHONPATH=src python src/eval_classifier.py            # run against the real model
PYTHONPATH=src python src/eval_classifier.py --dry-run  # plan only, no model load
  • Measures zero-shot classification accuracy of local_encoder.classify_zero_shot over the 46-task eval set (evals/tasks.yaml minus the 11 tool_use_agentic rows) with three noise variants per prompt (clean, short-noise, long-noise).
  • Prints top-1 accuracy, a full confusion matrix, per-category precision/recall, confidence distribution split by correct/incorrect, and the per-cell blended_score gap from the proficiency table (when router.db is present).
  • Standalone: imports transformers/torch only when actually running against the model, and lives under requirements-encoder.txt, so the router's dispatch path never touches them.

Proficiency

The harness is a prior, not the routing score. The benchmark produces a best-guess quality estimate before real traffic arrives. Routing turns that candidate score into an expected pass rate on real traffic, then updates it with client outcomes via POST /outcome and feedback.py. The benchmark is what you have when no one has reported whether the answer worked yet; it is not the final truth.

Two new categories for local dispatch

The shipped qwen2.5-coder-router:14b row is eligible for file_summarization and diff_checking:

Category Kind Why
diff_checking exact A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs.
file_summarization judge A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files.

Both categories run through the same harness as the cloud rows. The only difference is that an ollama-local identity resolves its Ollama endpoint from cfg.local_dispatch_models and its cost model is seeded separately (see below).

Seeding local-dispatch costs

Cloud rows arrive with catalog prices; local rows start with NULL token rates and must be measured. seed_local_dispatch_energy.py runs a small reference sweep (sum_small, sum_large, diff_small, long_answer), samples GPU draw with nvidia-smi, and derives per-token USD rates via through-origin OLS on energy_kwh ~ a*prompt_tokens + b*completion_tokens at the user's tariff.

An intercept is intentionally excluded. Routing's cost model is purely per-token, so fitting a fixed overhead would add a term estimated_cost() cannot express. Consistency with the ledger beats theoretical purity, and that choice is stated rather than hidden. Repeat the sweep with --samples 5 or more to tighten the fit; the resulting rates are written to models.cost_per_1m_prompt / cost_per_1m_completion for provider='ollama-local' rows and are never overwritten by later poller runs.

Classifier Reliability Notes

The classifier is the one blocking LLM call on the request path, so it sets the latency floor for every routed request. classifier.base_url / api_key_env / model accept any OpenAI-compatible endpoint — a local Ollama, an Ollama on another machine across a VPN, or a cloud model.

Pick a non-reasoning model, and judge it by the tail, not the mean. The classifier's entire output is ~45 tokens of JSON, and a reasoning model spends its budget getting there instead. This project has swapped classifiers twice on exactly that evidence — see local-models.md for the full three-generation comparison table rather than duplicating it here; the current default is qwen2.5-coder-router:14b.

Four settings keep a slow or misbehaving classifier from cascading into worse failures — full rationale for each in local-models.md:

setting value one-line reason
max_retries 0 the SDK's own retries would triple the timeout silently
max_output_tokens 1024 bounds a reasoning model's chain of thought
max_input_chars 8000 the classifier needs the instruction, not the document
temperature 0 reproducible classification, not a coin flip per request

Two more settings are graceful degradation rather than prevention:

  • Fallback tier/category (fallback_tier: 2 / fallback_category: general_chat) is what a timeout, error, or unparseable output degrades to, instead of a 502/503. A coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier.
  • Session classification cache (session_cache.enabled in config/config.yaml, off by default) remembers the last task_category/task_tier decision per session for staleness_seconds (default 1200), so a long agent session skips the classifier round-trip on every turn. It only ever short-circuits the classifier call — capability flags (tools/images/json) are still read fresh from each request body, fallback classifications are never cached, and there is no persistence beyond process memory. A cache-hit decision is observable via route_decisions.classification_source = "cached".