Change the unit of session_cache.staleness from minutes to seconds so it can express finer-grained (sub-minute) staleness windows. This is a straight rename, not an additive/compat knob — no deprecated alias, per the project's convention of updating every consumer in the same change. New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was 120 minutes). Default: 1200 seconds (was 20 minutes). The validator's reasoning is unit-independent and carries over: the floor is deliberately > 0 because 0 would make session_cache.get() miss every turn while put() still writes and the classifier-failure cascade's stale_read ignores staleness; the ceiling reasoning (unbounded window = never-expiring cache, 7200s still >> 840s real max run) also carries over in seconds. Every consumer updated in the same commit: - src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX = 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated - src/dispatcher.py: drop the * 60 conversion (field is native seconds) - src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST, _CONFIG_GET_ORDER, _runtime_state, error message template - admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS - config/config.yaml: staleness_seconds: 1200 - tests: test_admin_runtime/config/frontend/knob_coverage, plus stale comment in test_chat_completions - docs: admin-portal.md, evaluation.md, README.md config.local.yaml is gitignored and will be migrated separately.
5.6 KiB
Eval harness and classifier notes. Back to README.
Self-Eval Harness (eval_proficiency.py)
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
PYTHONPATH=src python -m eval_proficiency --categories coding_general
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
- Runs every task through the target provider, scores it, writes to
proficiency. - Scores accumulate (running mean), so repeated runs tighten estimates.
propagate_to_variantscopies evaluated scores to equivalent serving variants.- Safety: code tasks run in a temp directory with a 15s wall-clock timeout — bounded isolation, not a container.
- Judge tasks skip the model being judged (avoids self-scoring bias).
- Local
ollama-localidentities resolve their endpoint fromcfg.local_dispatch_modelsand write proficiency rows withprovider='ollama-local'. Judges remain cloud-only.
Proficiency
The harness is a prior, not the routing score. The benchmark produces a
best-guess quality estimate before real traffic arrives. Routing turns that
candidate score into an expected pass rate on real traffic, then updates it
with client outcomes via POST /outcome and feedback.py. The benchmark is
what you have when no one has reported whether the answer worked yet; it is
not the final truth.
Two new categories for local dispatch
The shipped qwen2.5-coder-router:14b row is eligible for file_summarization
and diff_checking:
| Category | Kind | Why |
|---|---|---|
diff_checking |
exact |
A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. |
file_summarization |
judge |
A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. |
Both categories run through the same harness as the cloud rows. The only
difference is that an ollama-local identity resolves its Ollama endpoint from
cfg.local_dispatch_models and its cost model is seeded separately (see below).
Seeding local-dispatch costs
Cloud rows arrive with catalog prices; local rows start with NULL token rates
and must be measured. seed_local_dispatch_energy.py runs a small reference
sweep (sum_small, sum_large, diff_small, long_answer), samples GPU draw
with nvidia-smi, and derives per-token USD rates via through-origin OLS on
energy_kwh ~ a*prompt_tokens + b*completion_tokens at the user's tariff.
An intercept is intentionally excluded. Routing's cost model is purely
per-token, so fitting a fixed overhead would add a term estimated_cost()
cannot express. Consistency with the ledger beats theoretical purity, and that
choice is stated rather than hidden. Repeat the sweep with
--samples 5 or more to tighten the fit; the resulting rates are written to
models.cost_per_1m_prompt / cost_per_1m_completion for provider='ollama-local'
rows and are never overwritten by later poller runs.
Classifier Reliability Notes
The classifier is the one blocking LLM call on the request path, so it sets
the latency floor for every routed request. classifier.base_url /
api_key_env / model accept any OpenAI-compatible endpoint — a local
Ollama, an Ollama on another machine across a VPN, or a cloud model.
Pick a non-reasoning model, and judge it by the tail, not the mean. The
classifier's entire output is ~45 tokens of JSON, and a reasoning model
spends its budget getting there instead. This project has swapped classifiers
twice on exactly that evidence — see
local-models.md
for the full three-generation comparison table rather than duplicating it
here; the current default is qwen2.5-coder-router:14b.
Four settings keep a slow or misbehaving classifier from cascading into worse failures — full rationale for each in local-models.md:
| setting | value | one-line reason |
|---|---|---|
max_retries |
0 |
the SDK's own retries would triple the timeout silently |
max_output_tokens |
1024 |
bounds a reasoning model's chain of thought |
max_input_chars |
8000 |
the classifier needs the instruction, not the document |
temperature |
0 |
reproducible classification, not a coin flip per request |
Two more settings are graceful degradation rather than prevention:
- Fallback tier/category (
fallback_tier: 2/fallback_category: general_chat) is what a timeout, error, or unparseable output degrades to, instead of a 502/503. A coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. - Session classification cache (
session_cache.enabledinconfig/config.yaml, off by default) remembers the lasttask_category/task_tierdecision per session forstaleness_seconds(default 1200), so a long agent session skips the classifier round-trip on every turn. It only ever short-circuits the classifier call — capability flags (tools/images/json) are still read fresh from each request body, fallback classifications are never cached, and there is no persistence beyond process memory. A cache-hit decision is observable viaroute_decisions.classification_source = "cached".