From d4d69a41fd3ca82dfb3d50d51a8b7af79f87a2a1 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Sun, 6 Sep 2026 22:10:11 -0400 Subject: [PATCH 1/2] docs: refresh CLAUDE.md and companion docs for PRs #40-#47 Covers per-provider quota measurement + credit-aware attenuation (#40), the :path routing fix (#42), quota UI billing-shape clarity (#43), the pinch protected-window size cap (#44, docs/pinch.md verified against what actually shipped), the OpenRouter opt-in allowlist (#45, with the firm x-ai/openai/anthropic exclusions stated as settled, not pending), and the models-page deprecated/stale toggle (#46). Also extends the classifier.mode section with the #48-#50 incident chain (confidence_threshold validation, the gated-model crash, and the natural-language-labels accuracy fix) since it's directly continuous with the local_encoder-is-live-but-not-yet-default note already there. --- CLAUDE.md | 96 +++++++++++++++++++++++++++++++++++++++++--- docs/admin-portal.md | 42 +++++++++++++++---- docs/api.md | 2 +- docs/architecture.md | 2 +- docs/local-models.md | 29 +++++++++++++ docs/pinch.md | 31 +++++++++++--- 6 files changed, 180 insertions(+), 22 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 51cd028..2aa5d66 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -20,9 +20,12 @@ the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it. -Neuralwatt is the only provider. OpenRouter was removed — the `provider` -column and the `(model_id, provider)` primary key stay so a second provider -can be added later without a migration. +Neuralwatt remains the primary provider. OpenRouter is back as an opt-in +provider gated by the `provider_model_allowlist` table — a provider with +`require_allowlist: true` is ignored unless its model_id is explicitly +allowlisted, and any previously-upserted row that drops off the list is +marked `deprecated` on the next poll. The `provider` column and the +`(model_id, provider)` primary key are unchanged, so this adds no migration. ## Stack @@ -207,7 +210,7 @@ rather than from months of history. ## What's built and working - `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md). -- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). +- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below. - `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md). - `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md). - `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md). @@ -216,9 +219,9 @@ rather than from months of history. - `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates. - `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below). - `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md). +- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded. - `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md). - `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md). -- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md). - `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md). - `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable. - `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md). @@ -227,10 +230,50 @@ rather than from months of history. - `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted. - `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted. - `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md). -- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). +- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor. - `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md). +- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below. +- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted. - `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md). +## #45 — OpenRouter is an opt-in allowlist provider + +OpenRouter used to be ingested whole, then removed, and is now back — but only +as an opt-in provider. The base config sets `openrouter.require_allowlist: true` +and ships a short seed allowlist. Models on that seed list are upserted and +kept active; anything else in the OpenRouter catalog is filtered out before +upsert and any previously-active OpenRouter row that is not on the list is +marked `deprecated` on the next poll. + +This is deliberately different from the old ingest-everything behavior. The +previous approach once pulled in a non-chat model (`lyria/...`) that returned +HTTP 404 on dispatch because the endpoint expected chat completions. The router +had paid for the classification, selected the model, and then failed on the +provider call. Allowlist-gating prevents that class of failure by default: if a +model id has not been reviewed and explicitly added, the router acts as if it +does not exist. + +The seed allowlist is short and has firm exclusions. It does NOT include: + +- `x-ai/*` (Grok) +- `openai/*` +- `anthropic/*` + +Those exclusions are non-negotiable. They are not "currently excluded" or +planned for future inclusion; they are deliberately absent from the seed list. +Adding one requires editing both the seed allowlist and this file. + +The admin portal exposes the allowlist under `/admin`: the providers page shows +which providers require one, and each provider row links to an allowlist editor +where entries can be added or removed. The underlying four API endpoints are +`GET /admin/api/providers/{name}/allowlist`, +`POST /admin/api/providers/{name}/allowlist`, +`DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers +page itself surfaces `require_allowlist` with a link to the allowlist editor. + +Neuralwatt remains the primary, ungated provider. The allowlist behavior only +fires for providers with `require_allowlist: true`. + ## Proficiency: category now changes routing `proficiency_score` is the ONLY category-dependent term in the ranking, so @@ -695,6 +738,42 @@ weeks. ### Which implementation is PRIMARY is now a config choice +`classifier.mode` in the live deployment is currently `local_encoder`, set in +`config.local.yaml` with `device: cuda` on the classifier host. That is the +mode answering real traffic right now. The explicit caveat is that a +CPU-vs-CUDA latency and confidence comparison on the live classifier host is +still pending; until that measurement exists, flipping the project default to +`local_encoder` is not decided. + +**Turning that mode on for real found three real bugs in one afternoon +(2026-09-06), each a fresh instance of this project's own recurring lesson — +verify against the live system, not the plan.** First, the config-load +validator for `confidence_threshold` didn't exist yet: the admin UI saved a +raw `80` (meant as 80%) straight into the overlay with no conversion, which +would have made every real confidence score read as below-threshold on the +next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability +exceeds 1.0). Caught before the restart, not after — see `confidence_threshold` +in [local-models](docs/local-models.md) for the full incident and the fix. +Second, once that was corrected and the service actually restarted, it +crash-looped twice more before coming up clean: the shipped default model +(`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace +sometime after this project picked it (401 on an unauthenticated GET of its +own model page), and separately `HF_HOME`'s default cache path falls outside +this service's `ProtectHome=read-only` sandbox exception — both are now fixed +(switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo). +Third, and most consequential: the very first real classifications measured +only 5 of 9 test categories correct, because `classify_zero_shot` was passing +raw config identifiers like `tool_use_agentic` and `diff_checking` directly as +zero-shot candidate labels — HF's pipeline scores a label against a hypothesis +template ("This example is {}."), and an underscored code token is not a +sentence the model's NLI training ever saw. Mapping each category to a natural- +language description before scoring, plus `multi_label=True` (the pipeline's +single-label default forces every candidate to compete for the same +probability mass), brought that to 8 of 9 correct with confidence scores +0.77-0.999 on the hits — the one remaining miss scored below the configured +threshold and correctly fell through to the safe fallback rather than +mis-routing. + `classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks what answers a classification request — a peer concept to the cascade above, **not a replacement for it**. Whichever mode is primary, a failure still @@ -732,6 +811,11 @@ through the admin portal's Classifier card, which reports the **live** resolved primary for `cloud_primary_auto` rather than echoing the config value (see [admin-portal](docs/admin-portal.md)). +The default in `config.yaml` is still `local_llm`. `local_encoder` is +intentionally an opt-in per-deployment choice rather than the repository +default until the pending CPU-vs-CUDA comparison on the live classifier host is +available. + ## Local dispatch model A second local model can now be dispatched directly for specific categories. diff --git a/docs/admin-portal.md b/docs/admin-portal.md index 1ac56f1..def0557 100644 --- a/docs/admin-portal.md +++ b/docs/admin-portal.md @@ -9,13 +9,20 @@ endpoints it is reachable only from `127.0.0.1`. The portal is a **six-page, glass dark-mode** dashboard built on Tabler/Bootstrap with a few lighter inline SVG icons (see `plans/admin-design-standards.md` for the visual language and `plans/admin-work-framework.md` for the method -used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for -its future tables and `admin/frontend/` for the browser UI. +used to evolve it). The six pages are Dashboard, Models, Profiles, Decisions, +Controls, and Providers. It is implemented in `admin.py`, with +`config/admin_schema.sql` for its future tables and `admin/frontend/` for the +browser UI. **Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip (burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict mix and category breakdown, five history mini-charts, and a recent-decisions -table. +table. Clicking the chip opens a detail modal with per-provider balances and +runway; each line is labeled by its billing shape, so `"telemetry"` balances +read as "overage allowance" and `"polled"` balances read as "credits". Tiny +values keep their sign: a negative balance smaller than half a cent renders as +`~$0.00 (slight overage)`, a positive tiny balance renders as `<$0.01`, and an +exact zero balance renders as `$0.00`. **Pinch savings** — the dashboard card shows pinch's effect: share pruned (as a percent), tokens saved, median saved, and 30d dollars saved. When pinch is @@ -30,7 +37,9 @@ for the full model. **Models** — `GET /admin/models` is the model-availability table. Each row shows the serving class, tier, status, and an override dropdown (`active` / -`deprecated` / `stale`) that writes through to the routing hard filters. +`deprecated` / `stale`) that writes through to the routing hard filters. By +default the table hides deprecated and stale rows; enable the **Show deprecated +/ stale** toggle to include them.
Admin models page @@ -127,6 +136,21 @@ can't safely represent: Admin controls page
+**Providers** — `GET /admin/providers` is the dispatch-provider management page. +The left card adds a new provider (`name`, `base_url`, `api_key_env`, +`has_energy_telemetry`, `enabled`); the right card lists configured providers +with their source badge (base config vs overlay) and edit/delete actions. Base +providers are read-only; overlay providers can be edited or deleted. Creating, +editing, or deleting a provider writes to `config/config.local.yaml`, and a +service reload is required before dispatch sees the change. + +Below the provider list is the **Provider allowlists** card. It lists every +provider and whether `require_allowlist` is enabled. For providers with an +allowlist, click **Manage** to see the current allowed model ids, remove +entries, and browse a live upstream catalog preview to add models with an +optional note. The catalog preview calls `poller.fetch_openrouter` directly, so +it is only available for OpenRouter providers that require an allowlist. + **Decisions** — `GET /admin/decisions` is the full decision log with kind/category/tier filters and free-text search. @@ -179,7 +203,7 @@ the attributable set, so `POST /outcome` keeps training proficiency normally. **Read-only dashboards** — `GET /admin/api/snapshot` exposes data for: -- quota: energy burn against `objective.plan_kwh_per_period`, plus per-provider balance, burn rate, and runway under `by_provider`, and a `total_balance_usd` sum (old flat balance/burn/runway keys removed) +- quota: energy burn against `objective.plan_kwh_per_period`, plus per-provider balance, burn rate, and runway under `by_provider`, and a `total_balance_usd` sum (old flat balance/burn/runway keys removed). Each provider entry carries `balance_source`: `"telemetry"` for balances served by the provider's own SSE energy comments (NeuralWatt's overage allowance) or `"polled"` for balances fetched from a provider credits endpoint (OpenRouter's prepaid balance) - per-model usage from `energy_observations` - live routing decisions from `route_decisions` - verdict mix and scoring coverage @@ -226,6 +250,8 @@ intermediate invalid state could land on disk. The GET response includes **Model availability overrides** -`POST /admin/api/models/{model_id}/{provider}/availability` marks a model as -`active`, `deprecated`, or `stale`. `DELETE` on the same path removes the override. -Deprecation feeds into the routing hard filters. +`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model +as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the +override. The `{model_id:path}` converter accepts model ids that contain +slashes, so providers like OpenRouter with slash-bearing ids are handled the +same way as plain ids. Deprecation feeds into the routing hard filters. diff --git a/docs/api.md b/docs/api.md index 9212d12..b9d1a09 100644 --- a/docs/api.md +++ b/docs/api.md @@ -20,7 +20,7 @@ allowance. `/metrics` returns a single JSON object with these top-level keys: -- `quota` — kWh metered in the last 30 days against `objective.plan_kwh_per_period` +- `quota` — kWh metered in the last 30 days against `objective.plan_kwh_per_period`; `by_provider` gives each provider its own balance, burn rate, and runway, with a `balance_source` of either `telemetry` or `polled`; `total_balance_usd` is the sum across providers. Old flat top-level `balance`/`burn`/`runway` keys were removed. - `coverage` — routable-model counts with energy/proficiency data plus warnings - `recent_decisions` — last 50 rows from `route_decisions` - `per_model` — per-model aggregates over the last 30 days of `energy_observations` diff --git a/docs/architecture.md b/docs/architecture.md index 35c7124..51c7087 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -31,7 +31,7 @@ | **`baseline_report.py`** (repo root, not `src/`) | Read-only retrospective: replays `route_decisions` against two trivial baselines (always-cheapest, always-highest-proficiency) to check whether scoring earns its complexity | Yes (DB) | | **`metrics.py`** | Read-only aggregations for `/health` and `GET /metrics`: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency. Never imports `dispatcher` | Yes (DB) | | **`events.py`** | In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure | -| **`admin.py`** | `/admin` management portal: serves five frontend pages (dashboard, models, profiles, decisions, controls) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers) | Yes (DB, network, file) | +| **`admin.py`** | `/admin` management portal: serves seven frontend pages (dashboard, models, profiles, proficiency, decisions, controls, providers) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers, provider-model allowlist) | Yes (DB, network, file) | | **`tui.py`** | Textual terminal dashboard over `GET /metrics` and `GET /events/decisions`; live routing feed, detail popup, category breakdown. Foreground tool, not a service | Yes (network) | | **`tui_model.py`** | Pure data layer for the TUI: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | Pure | | **`tui_sse.py`** | Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) | diff --git a/docs/local-models.md b/docs/local-models.md index d05d425..bf81d98 100644 --- a/docs/local-models.md +++ b/docs/local-models.md @@ -229,6 +229,35 @@ opt-in capture feature (scoped in `CLAUDE.md`'s "What's NOT built yet", not built). A below-`confidence_threshold` result is treated as a failure and walks the same cascade a local-LLM parse failure would, unchanged. +`confidence_threshold` is a **probability in [0.0, 1.0]**, not a percentage. +`classify_zero_shot` returns a confidence score between 0 and 1, so a value of +`0.5` means 50% confidence. The shipped default is `0.5`. The validator in +`src/config.py` rejects anything outside [0.0, 1.0] with an error that says +exactly this: it is a probability, not a percent. + +The admin UI lets you type 0-100 and scales to the probability before saving, +so typing `80` in the UI writes `0.8` to the overlay. The config file and any +other direct writer must supply the raw probability. That boundary conversion +matters: a raw `80` in the YAML is two orders of magnitude above the maximum +and would make every real score read as below threshold, since no probability +exceeds 1.0. + +That validator exists because of a **production near-miss** (issue #47). The +admin UI used to save the threshold as written; an `80` was entered, the UI +wrote `80.0` straight into the overlay, and it was only caught because the +service had not yet restarted and loaded the bad value. Once restarted, +classification would have failed on *every* request. The UI now converts +0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line +of defense for every other caller. + +`local_encoder` runs on **cpu by default** in the shipped example; setting it +to `cuda` requires a `config.local.yaml` overlay and a host that actually has +the hardware. The project does not ship a settled default device. Match the +device to the machine the router is running on, otherwise `torch` will fail +fast with a clear error at first classification. CPU inference is the safe, +portable starting point; CUDA is host-dependent and should be chosen only +where the router process has a card available. + `transformers`/`torch` are optional (`requirements-encoder.txt`, not the main `requirements.txt` — pinned-dependency discipline applies here too) and imported lazily inside `local_encoder.py`, the same rule `tui.py` follows for diff --git a/docs/pinch.md b/docs/pinch.md index 4b810ab..19bb800 100644 --- a/docs/pinch.md +++ b/docs/pinch.md @@ -60,12 +60,15 @@ below): prune → persist → `_check_pinned_capabilities`. Pinch's tailoring is deliberate and safe: -- **Only `tool` results** older than the protected window are trimmed or - summarized. Tool results are where a long agent session's tokens actually - live, they are the least likely to still be needed in full by the time a - later turn is answered, and replacing one with a short placeholder is - reversible at the semantic level — a wrong guess costs context but never - breaks the request. +- **Only `tool` results** are trimmed or summarized. Older results outside the + protected window are the main target, but any **outsized result inside the + window** that is longer than `pinch.protected_max_chars` is also capped. That + threshold is deliberately a much higher bar than `max_summarize_chars`; recent + results are more likely to still matter, so it only catches true outliers. + Tool results are where a long agent session's tokens actually live, they are + the least likely to still be needed in full by the time a later turn is + answered, and replacing one with a short placeholder is reversible at the + semantic level — a wrong guess costs context but never breaks the request. - **Never user / assistant / system messages.** Trimming one of those can change what the model is being asked, so they are always kept verbatim. - **Never the classifier input.** The classification turn uses @@ -82,6 +85,22 @@ valid conversation. Only runs (and only mutates anything) when the estimate actually exceeds the budget; otherwise the original list is returned untouched. +### Capping outsized results in the protected window + +The protected window, controlled by `keep_last_turns`, is intentionally generous: +anything in the last few user turns stays verbatim because it is most likely to +still be needed. But that created a gap. An autonomous tool-call loop (code +execution, file reads, test runs) can produce one enormous tool result without a +new user message, and that whole result used to count as part of the protected +window. Pinch only fired when the total conversation crossed the budget, but +once it did, those outsized protected results shipped verbatim regardless of +size. In one measured case a 324k-token conversation shrank only ~8% because +nearly all of it sat inside the protected window. `protected_max_chars` closes +that gap: tool results inside the window are still preserved unless they cross +this outlier threshold, in which case they receive the same head/tail elision +applied to older candidates. The default is high enough to avoid touching normal +output; it only intervenes when a single result threatens the budget. + ## Config block All pinch configuration lives under `pinch:` in `config/config.yaml`, validated -- 2.49.1 From 5ef411496bdd7d6659fc59c1f8c3a07441f34fe7 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Sun, 6 Sep 2026 22:15:05 -0400 Subject: [PATCH 2/2] docs: dedupe confidence_threshold section, extend with the #50 accuracy fix The merge from origin/main brought in #48's own confidence_threshold paragraph, which duplicated this branch's own independent write-up of the same incident. Kept the more detailed version, removed the redundant one, and added the #50 natural-language-labels accuracy fix (5/9 -> 8/9 correct, confidence 0.15-0.90 -> 0.39-0.999) since it's the direct continuation of the same incident thread. --- docs/local-models.md | 26 +++++++++++++++++--------- 1 file changed, 17 insertions(+), 9 deletions(-) diff --git a/docs/local-models.md b/docs/local-models.md index 3650642..656f5be 100644 --- a/docs/local-models.md +++ b/docs/local-models.md @@ -223,15 +223,6 @@ its own model page sometime after this project picked it, discovered live refused to boot rather than fail opaquely on the first request, but the service still crash-looped until the default was corrected.) -`classifier.encoder.confidence_threshold` is a `[0.0, 1.0]` probability -(`classify_zero_shot`'s own output), not a percent — the admin UI takes 0-100 -for a human to type and converts at the save boundary, but a config file -edit or any other caller must use the raw probability. Config load now -validates the range; a value like `80` used to be silently accepted and -would make every real confidence score read as below-threshold, since none -can exceed `1.0` (also caught live 2026-09-06, before a restart made it -active). - The trade is real, not free. It only produces `task_category` — no `task_tier` signal exists in a zero-shot label score, so tier falls back to `classifier.fallback_tier` for this mode, a documented limitation rather than @@ -264,6 +255,23 @@ classification would have failed on *every* request. The UI now converts 0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line of defense for every other caller. +**Candidate labels are natural-language descriptions, never the raw config +identifier.** HF's zero-shot pipeline scores a candidate label against the +input via a hypothesis template (`"This example is {}."`), so the label +itself has to read as a sentence for entailment scoring to work — feeding it +`tool_use_agentic` or `diff_checking` verbatim asks the model to judge +`"This example is tool_use_agentic."`, not a sentence its NLI training ever +saw. `local_encoder.py`'s `_CATEGORY_DESCRIPTIONS` maps each category to a +real description before scoring, and `multi_label=True` scores each +candidate independently rather than normalizing every score to sum to 1 (the +pipeline's single-label default, which drags a genuinely good match down +whenever another category is also plausible). Measured live 2026-09-06 on the +same 9-prompt set: raw identifiers scored 5/9 correct at confidence +0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A +category missing from the map falls back to its own raw string rather than +raising, so a newly-added config category degrades gracefully instead of +crashing. + `local_encoder` runs on **cpu by default** in the shipped example; setting it to `cuda` requires a `config.local.yaml` overlay and a host that actually has the hardware. The project does not ship a settled default device. Match the -- 2.49.1