Change the unit of session_cache.staleness from minutes to seconds so it can express finer-grained (sub-minute) staleness windows. This is a straight rename, not an additive/compat knob — no deprecated alias, per the project's convention of updating every consumer in the same change. New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was 120 minutes). Default: 1200 seconds (was 20 minutes). The validator's reasoning is unit-independent and carries over: the floor is deliberately > 0 because 0 would make session_cache.get() miss every turn while put() still writes and the classifier-failure cascade's stale_read ignores staleness; the ceiling reasoning (unbounded window = never-expiring cache, 7200s still >> 840s real max run) also carries over in seconds. Every consumer updated in the same commit: - src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX = 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated - src/dispatcher.py: drop the * 60 conversion (field is native seconds) - src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST, _CONFIG_GET_ORDER, _runtime_state, error message template - admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS - config/config.yaml: staleness_seconds: 1200 - tests: test_admin_runtime/config/frontend/knob_coverage, plus stale comment in test_chat_completions - docs: admin-portal.md, evaluation.md, README.md config.local.yaml is gitignored and will be migrated separately.
103 lines
5.6 KiB
Markdown
103 lines
5.6 KiB
Markdown
> Eval harness and classifier notes. Back to [README](../README.md).
|
||
|
||
## Self-Eval Harness (`eval_proficiency.py`)
|
||
|
||
```bash
|
||
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
|
||
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
|
||
PYTHONPATH=src python -m eval_proficiency --categories coding_general
|
||
PYTHONPATH=src python -m eval_proficiency --models qwen2.5-coder-router:14b --categories file_summarization,diff_checking
|
||
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
|
||
```
|
||
|
||
- Runs every task through the target provider, scores it, writes to `proficiency`.
|
||
- Scores **accumulate** (running mean), so repeated runs tighten estimates.
|
||
- `propagate_to_variants` copies evaluated scores to equivalent serving variants.
|
||
- Safety: code tasks run in a temp directory with a 15s wall-clock timeout —
|
||
bounded isolation, not a container.
|
||
- Judge tasks skip the model being judged (avoids self-scoring bias).
|
||
- Local `ollama-local` identities resolve their endpoint from `cfg.local_dispatch_models`
|
||
and write proficiency rows with `provider='ollama-local'`. Judges remain cloud-only.
|
||
|
||
## Proficiency
|
||
|
||
The harness is a **prior**, not the routing score. The benchmark produces a
|
||
best-guess quality estimate before real traffic arrives. Routing turns that
|
||
candidate score into an expected pass rate on real traffic, then updates it
|
||
with client outcomes via `POST /outcome` and `feedback.py`. The benchmark is
|
||
what you have when no one has reported whether the answer worked yet; it is
|
||
not the final truth.
|
||
|
||
## Two new categories for local dispatch
|
||
|
||
The shipped `qwen2.5-coder-router:14b` row is eligible for `file_summarization`
|
||
and `diff_checking`:
|
||
|
||
| Category | Kind | Why |
|
||
|---|---|---|
|
||
| `diff_checking` | `exact` | A diff either introduces a regression or it doesn't, so a YES/NO check is the right signal. The 8 tasks cover four safe vs buggy refactor pairs. |
|
||
| `file_summarization` | `judge` | A summary is prose; a rubric checks whether it states the non-obvious gotcha rather than repeating the happy path. The 6 tasks exercise real-shaped files. |
|
||
|
||
Both categories run through the same harness as the cloud rows. The only
|
||
difference is that an `ollama-local` identity resolves its Ollama endpoint from
|
||
`cfg.local_dispatch_models` and its cost model is seeded separately (see below).
|
||
|
||
## Seeding local-dispatch costs
|
||
|
||
Cloud rows arrive with catalog prices; local rows start with NULL token rates
|
||
and must be measured. `seed_local_dispatch_energy.py` runs a small reference
|
||
sweep (`sum_small`, `sum_large`, `diff_small`, `long_answer`), samples GPU draw
|
||
with `nvidia-smi`, and derives per-token USD rates via through-origin OLS on
|
||
`energy_kwh ~ a*prompt_tokens + b*completion_tokens` at the user's tariff.
|
||
|
||
An intercept is intentionally excluded. Routing's cost model is purely
|
||
per-token, so fitting a fixed overhead would add a term `estimated_cost()`
|
||
cannot express. Consistency with the ledger beats theoretical purity, and that
|
||
choice is stated rather than hidden. Repeat the sweep with
|
||
`--samples 5` or more to tighten the fit; the resulting rates are written to
|
||
`models.cost_per_1m_prompt` / `cost_per_1m_completion` for `provider='ollama-local'`
|
||
rows and are never overwritten by later poller runs.
|
||
|
||
## Classifier Reliability Notes
|
||
|
||
The classifier is the one blocking LLM call on the request path, so it sets
|
||
the latency floor for every routed request. `classifier.base_url` /
|
||
`api_key_env` / `model` accept any OpenAI-compatible endpoint — a local
|
||
Ollama, an Ollama on another machine across a VPN, or a cloud model.
|
||
|
||
**Pick a non-reasoning model, and judge it by the tail, not the mean.** The
|
||
classifier's entire output is ~45 tokens of JSON, and a reasoning model
|
||
spends its budget getting there instead. This project has swapped classifiers
|
||
twice on exactly that evidence — see
|
||
[local-models.md](local-models.md#three-generations-and-the-one-lesson-that-outlived-all-of-them)
|
||
for the full three-generation comparison table rather than duplicating it
|
||
here; the current default is `qwen2.5-coder-router:14b`.
|
||
|
||
Four settings keep a slow or misbehaving classifier from cascading into
|
||
worse failures — full rationale for each in
|
||
[local-models.md](local-models.md#the-four-settings-and-why-each-exists):
|
||
|
||
| setting | value | one-line reason |
|
||
|---|---|---|
|
||
| `max_retries` | `0` | the SDK's own retries would triple the timeout silently |
|
||
| `max_output_tokens` | `1024` | bounds a reasoning model's chain of thought |
|
||
| `max_input_chars` | `8000` | the classifier needs the instruction, not the document |
|
||
| `temperature` | `0` | reproducible classification, not a coin flip per request |
|
||
|
||
Two more settings are graceful degradation rather than prevention:
|
||
|
||
- **Fallback tier/category** (`fallback_tier: 2` / `fallback_category:
|
||
general_chat`) is what a timeout, error, or unparseable output degrades to,
|
||
instead of a 502/503. A coding agent would rather have a mid-tier answer
|
||
than an error. Escalation deliberately skips fallbacks, so an unavailable
|
||
local model does not silently promote every request to the frontier tier.
|
||
- **Session classification cache** (`session_cache.enabled` in
|
||
`config/config.yaml`, off by default) remembers the last
|
||
`task_category`/`task_tier` decision per session for `staleness_seconds`
|
||
(default 1200), so a long agent session skips the classifier round-trip on
|
||
every turn. It only ever short-circuits the classifier call — capability
|
||
flags (tools/images/json) are still read fresh from each request body,
|
||
fallback classifications are never cached, and there is no persistence
|
||
beyond process memory. A cache-hit decision is observable via
|
||
`route_decisions.classification_source = "cached"`.
|