docs: refresh CLAUDE.md and companion docs for PRs #40-#50 #51
96
CLAUDE.md
96
CLAUDE.md
@@ -20,9 +20,12 @@ the conclusion, but treat them as observations with a date on them, not as
|
||||
constants — the catalog, prices, grid intensity and pool load all move. When
|
||||
a number here decides something, re-run the measurement before trusting it.
|
||||
|
||||
Neuralwatt is the only provider. OpenRouter was removed — the `provider`
|
||||
column and the `(model_id, provider)` primary key stay so a second provider
|
||||
can be added later without a migration.
|
||||
Neuralwatt remains the primary provider. OpenRouter is back as an opt-in
|
||||
provider gated by the `provider_model_allowlist` table — a provider with
|
||||
`require_allowlist: true` is ignored unless its model_id is explicitly
|
||||
allowlisted, and any previously-upserted row that drops off the list is
|
||||
marked `deprecated` on the next poll. The `provider` column and the
|
||||
`(model_id, provider)` primary key are unchanged, so this adds no migration.
|
||||
|
||||
## Stack
|
||||
|
||||
@@ -207,7 +210,7 @@ rather than from months of history.
|
||||
## What's built and working
|
||||
|
||||
- `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md).
|
||||
- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md).
|
||||
- `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below.
|
||||
- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md).
|
||||
- `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md).
|
||||
- `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md).
|
||||
@@ -216,9 +219,9 @@ rather than from months of history.
|
||||
- `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates.
|
||||
- `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below).
|
||||
- `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md).
|
||||
- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded.
|
||||
- `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md).
|
||||
- `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md).
|
||||
- `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md).
|
||||
- `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md).
|
||||
- `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable.
|
||||
- `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md).
|
||||
@@ -227,10 +230,50 @@ rather than from months of history.
|
||||
- `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted.
|
||||
- `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.
|
||||
- `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md).
|
||||
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md).
|
||||
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor.
|
||||
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
|
||||
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
|
||||
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
|
||||
- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md).
|
||||
|
||||
## #45 — OpenRouter is an opt-in allowlist provider
|
||||
|
||||
OpenRouter used to be ingested whole, then removed, and is now back — but only
|
||||
as an opt-in provider. The base config sets `openrouter.require_allowlist: true`
|
||||
and ships a short seed allowlist. Models on that seed list are upserted and
|
||||
kept active; anything else in the OpenRouter catalog is filtered out before
|
||||
upsert and any previously-active OpenRouter row that is not on the list is
|
||||
marked `deprecated` on the next poll.
|
||||
|
||||
This is deliberately different from the old ingest-everything behavior. The
|
||||
previous approach once pulled in a non-chat model (`lyria/...`) that returned
|
||||
HTTP 404 on dispatch because the endpoint expected chat completions. The router
|
||||
had paid for the classification, selected the model, and then failed on the
|
||||
provider call. Allowlist-gating prevents that class of failure by default: if a
|
||||
model id has not been reviewed and explicitly added, the router acts as if it
|
||||
does not exist.
|
||||
|
||||
The seed allowlist is short and has firm exclusions. It does NOT include:
|
||||
|
||||
- `x-ai/*` (Grok)
|
||||
- `openai/*`
|
||||
- `anthropic/*`
|
||||
|
||||
Those exclusions are non-negotiable. They are not "currently excluded" or
|
||||
planned for future inclusion; they are deliberately absent from the seed list.
|
||||
Adding one requires editing both the seed allowlist and this file.
|
||||
|
||||
The admin portal exposes the allowlist under `/admin`: the providers page shows
|
||||
which providers require one, and each provider row links to an allowlist editor
|
||||
where entries can be added or removed. The underlying four API endpoints are
|
||||
`GET /admin/api/providers/{name}/allowlist`,
|
||||
`POST /admin/api/providers/{name}/allowlist`,
|
||||
`DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers
|
||||
page itself surfaces `require_allowlist` with a link to the allowlist editor.
|
||||
|
||||
Neuralwatt remains the primary, ungated provider. The allowlist behavior only
|
||||
fires for providers with `require_allowlist: true`.
|
||||
|
||||
## Proficiency: category now changes routing
|
||||
|
||||
`proficiency_score` is the ONLY category-dependent term in the ranking, so
|
||||
@@ -695,6 +738,42 @@ weeks.
|
||||
|
||||
### Which implementation is PRIMARY is now a config choice
|
||||
|
||||
`classifier.mode` in the live deployment is currently `local_encoder`, set in
|
||||
`config.local.yaml` with `device: cuda` on the classifier host. That is the
|
||||
mode answering real traffic right now. The explicit caveat is that a
|
||||
CPU-vs-CUDA latency and confidence comparison on the live classifier host is
|
||||
still pending; until that measurement exists, flipping the project default to
|
||||
`local_encoder` is not decided.
|
||||
|
||||
**Turning that mode on for real found three real bugs in one afternoon
|
||||
(2026-09-06), each a fresh instance of this project's own recurring lesson —
|
||||
verify against the live system, not the plan.** First, the config-load
|
||||
validator for `confidence_threshold` didn't exist yet: the admin UI saved a
|
||||
raw `80` (meant as 80%) straight into the overlay with no conversion, which
|
||||
would have made every real confidence score read as below-threshold on the
|
||||
next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability
|
||||
exceeds 1.0). Caught before the restart, not after — see `confidence_threshold`
|
||||
in [local-models](docs/local-models.md) for the full incident and the fix.
|
||||
Second, once that was corrected and the service actually restarted, it
|
||||
crash-looped twice more before coming up clean: the shipped default model
|
||||
(`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace
|
||||
sometime after this project picked it (401 on an unauthenticated GET of its
|
||||
own model page), and separately `HF_HOME`'s default cache path falls outside
|
||||
this service's `ProtectHome=read-only` sandbox exception — both are now fixed
|
||||
(switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo).
|
||||
Third, and most consequential: the very first real classifications measured
|
||||
only 5 of 9 test categories correct, because `classify_zero_shot` was passing
|
||||
raw config identifiers like `tool_use_agentic` and `diff_checking` directly as
|
||||
zero-shot candidate labels — HF's pipeline scores a label against a hypothesis
|
||||
template ("This example is {}."), and an underscored code token is not a
|
||||
sentence the model's NLI training ever saw. Mapping each category to a natural-
|
||||
language description before scoring, plus `multi_label=True` (the pipeline's
|
||||
single-label default forces every candidate to compete for the same
|
||||
probability mass), brought that to 8 of 9 correct with confidence scores
|
||||
0.77-0.999 on the hits — the one remaining miss scored below the configured
|
||||
threshold and correctly fell through to the safe fallback rather than
|
||||
mis-routing.
|
||||
|
||||
`classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`) picks
|
||||
what answers a classification request — a peer concept to the cascade above,
|
||||
**not a replacement for it**. Whichever mode is primary, a failure still
|
||||
@@ -732,6 +811,11 @@ through the admin portal's Classifier card, which reports the **live**
|
||||
resolved primary for `cloud_primary_auto` rather than echoing the config
|
||||
value (see [admin-portal](docs/admin-portal.md)).
|
||||
|
||||
The default in `config.yaml` is still `local_llm`. `local_encoder` is
|
||||
intentionally an opt-in per-deployment choice rather than the repository
|
||||
default until the pending CPU-vs-CUDA comparison on the live classifier host is
|
||||
available.
|
||||
|
||||
## Local dispatch model
|
||||
|
||||
A second local model can now be dispatched directly for specific categories.
|
||||
|
||||
@@ -9,13 +9,20 @@ endpoints it is reachable only from `127.0.0.1`.
|
||||
The portal is a **six-page, glass dark-mode** dashboard built on Tabler/Bootstrap
|
||||
with a few lighter inline SVG icons (see `plans/admin-design-standards.md`
|
||||
for the visual language and `plans/admin-work-framework.md` for the method
|
||||
used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for
|
||||
its future tables and `admin/frontend/` for the browser UI.
|
||||
used to evolve it). The six pages are Dashboard, Models, Profiles, Decisions,
|
||||
Controls, and Providers. It is implemented in `admin.py`, with
|
||||
`config/admin_schema.sql` for its future tables and `admin/frontend/` for the
|
||||
browser UI.
|
||||
|
||||
**Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip
|
||||
(burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict
|
||||
mix and category breakdown, five history mini-charts, and a recent-decisions
|
||||
table.
|
||||
table. Clicking the chip opens a detail modal with per-provider balances and
|
||||
runway; each line is labeled by its billing shape, so `"telemetry"` balances
|
||||
read as "overage allowance" and `"polled"` balances read as "credits". Tiny
|
||||
values keep their sign: a negative balance smaller than half a cent renders as
|
||||
`~$0.00 (slight overage)`, a positive tiny balance renders as `<$0.01`, and an
|
||||
exact zero balance renders as `$0.00`.
|
||||
|
||||
**Pinch savings** — the dashboard card shows pinch's effect: share pruned (as a
|
||||
percent), tokens saved, median saved, and 30d dollars saved. When pinch is
|
||||
@@ -30,7 +37,9 @@ for the full model.
|
||||
|
||||
**Models** — `GET /admin/models` is the model-availability table. Each row shows
|
||||
the serving class, tier, status, and an override dropdown (`active` /
|
||||
`deprecated` / `stale`) that writes through to the routing hard filters.
|
||||
`deprecated` / `stale`) that writes through to the routing hard filters. By
|
||||
default the table hides deprecated and stale rows; enable the **Show deprecated
|
||||
/ stale** toggle to include them.
|
||||
|
||||
<div align="center">
|
||||
<img src="../assets/screenshots/models.png" alt="Admin models page" width="820">
|
||||
@@ -127,6 +136,21 @@ can't safely represent:
|
||||
<img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820">
|
||||
</div>
|
||||
|
||||
**Providers** — `GET /admin/providers` is the dispatch-provider management page.
|
||||
The left card adds a new provider (`name`, `base_url`, `api_key_env`,
|
||||
`has_energy_telemetry`, `enabled`); the right card lists configured providers
|
||||
with their source badge (base config vs overlay) and edit/delete actions. Base
|
||||
providers are read-only; overlay providers can be edited or deleted. Creating,
|
||||
editing, or deleting a provider writes to `config/config.local.yaml`, and a
|
||||
service reload is required before dispatch sees the change.
|
||||
|
||||
Below the provider list is the **Provider allowlists** card. It lists every
|
||||
provider and whether `require_allowlist` is enabled. For providers with an
|
||||
allowlist, click **Manage** to see the current allowed model ids, remove
|
||||
entries, and browse a live upstream catalog preview to add models with an
|
||||
optional note. The catalog preview calls `poller.fetch_openrouter` directly, so
|
||||
it is only available for OpenRouter providers that require an allowlist.
|
||||
|
||||
**Decisions** — `GET /admin/decisions` is the full decision log with
|
||||
kind/category/tier filters and free-text search.
|
||||
|
||||
@@ -179,7 +203,7 @@ the attributable set, so `POST /outcome` keeps training proficiency normally.
|
||||
|
||||
**Read-only dashboards** — `GET /admin/api/snapshot` exposes data for:
|
||||
|
||||
- quota: energy burn against `objective.plan_kwh_per_period`, plus per-provider balance, burn rate, and runway under `by_provider`, and a `total_balance_usd` sum (old flat balance/burn/runway keys removed)
|
||||
- quota: energy burn against `objective.plan_kwh_per_period`, plus per-provider balance, burn rate, and runway under `by_provider`, and a `total_balance_usd` sum (old flat balance/burn/runway keys removed). Each provider entry carries `balance_source`: `"telemetry"` for balances served by the provider's own SSE energy comments (NeuralWatt's overage allowance) or `"polled"` for balances fetched from a provider credits endpoint (OpenRouter's prepaid balance)
|
||||
- per-model usage from `energy_observations`
|
||||
- live routing decisions from `route_decisions`
|
||||
- verdict mix and scoring coverage
|
||||
@@ -226,6 +250,8 @@ intermediate invalid state could land on disk. The GET response includes
|
||||
|
||||
**Model availability overrides**
|
||||
|
||||
`POST /admin/api/models/{model_id}/{provider}/availability` marks a model as
|
||||
`active`, `deprecated`, or `stale`. `DELETE` on the same path removes the override.
|
||||
Deprecation feeds into the routing hard filters.
|
||||
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
|
||||
as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the
|
||||
override. The `{model_id:path}` converter accepts model ids that contain
|
||||
slashes, so providers like OpenRouter with slash-bearing ids are handled the
|
||||
same way as plain ids. Deprecation feeds into the routing hard filters.
|
||||
|
||||
@@ -20,7 +20,7 @@ allowance.
|
||||
|
||||
`/metrics` returns a single JSON object with these top-level keys:
|
||||
|
||||
- `quota` — kWh metered in the last 30 days against `objective.plan_kwh_per_period`
|
||||
- `quota` — kWh metered in the last 30 days against `objective.plan_kwh_per_period`; `by_provider` gives each provider its own balance, burn rate, and runway, with a `balance_source` of either `telemetry` or `polled`; `total_balance_usd` is the sum across providers. Old flat top-level `balance`/`burn`/`runway` keys were removed.
|
||||
- `coverage` — routable-model counts with energy/proficiency data plus warnings
|
||||
- `recent_decisions` — last 50 rows from `route_decisions`
|
||||
- `per_model` — per-model aggregates over the last 30 days of `energy_observations`
|
||||
|
||||
@@ -31,7 +31,7 @@
|
||||
| **`baseline_report.py`** (repo root, not `src/`) | Read-only retrospective: replays `route_decisions` against two trivial baselines (always-cheapest, always-highest-proficiency) to check whether scoring earns its complexity | Yes (DB) |
|
||||
| **`metrics.py`** | Read-only aggregations for `/health` and `GET /metrics`: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency. Never imports `dispatcher` | Yes (DB) |
|
||||
| **`events.py`** | In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure |
|
||||
| **`admin.py`** | `/admin` management portal: serves five frontend pages (dashboard, models, profiles, decisions, controls) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers) | Yes (DB, network, file) |
|
||||
| **`admin.py`** | `/admin` management portal: serves seven frontend pages (dashboard, models, profiles, proficiency, decisions, controls, providers) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers, provider-model allowlist) | Yes (DB, network, file) |
|
||||
| **`tui.py`** | Textual terminal dashboard over `GET /metrics` and `GET /events/decisions`; live routing feed, detail popup, category breakdown. Foreground tool, not a service | Yes (network) |
|
||||
| **`tui_model.py`** | Pure data layer for the TUI: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | Pure |
|
||||
| **`tui_sse.py`** | Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) |
|
||||
|
||||
@@ -223,15 +223,6 @@ its own model page sometime after this project picked it, discovered live
|
||||
refused to boot rather than fail opaquely on the first request, but the
|
||||
service still crash-looped until the default was corrected.)
|
||||
|
||||
`classifier.encoder.confidence_threshold` is a `[0.0, 1.0]` probability
|
||||
(`classify_zero_shot`'s own output), not a percent — the admin UI takes 0-100
|
||||
for a human to type and converts at the save boundary, but a config file
|
||||
edit or any other caller must use the raw probability. Config load now
|
||||
validates the range; a value like `80` used to be silently accepted and
|
||||
would make every real confidence score read as below-threshold, since none
|
||||
can exceed `1.0` (also caught live 2026-09-06, before a restart made it
|
||||
active).
|
||||
|
||||
The trade is real, not free. It only produces `task_category` — no `task_tier`
|
||||
signal exists in a zero-shot label score, so tier falls back to
|
||||
`classifier.fallback_tier` for this mode, a documented limitation rather than
|
||||
@@ -243,6 +234,52 @@ opt-in capture feature (scoped in `CLAUDE.md`'s "What's NOT built yet", not
|
||||
built). A below-`confidence_threshold` result is treated as a failure and
|
||||
walks the same cascade a local-LLM parse failure would, unchanged.
|
||||
|
||||
`confidence_threshold` is a **probability in [0.0, 1.0]**, not a percentage.
|
||||
`classify_zero_shot` returns a confidence score between 0 and 1, so a value of
|
||||
`0.5` means 50% confidence. The shipped default is `0.5`. The validator in
|
||||
`src/config.py` rejects anything outside [0.0, 1.0] with an error that says
|
||||
exactly this: it is a probability, not a percent.
|
||||
|
||||
The admin UI lets you type 0-100 and scales to the probability before saving,
|
||||
so typing `80` in the UI writes `0.8` to the overlay. The config file and any
|
||||
other direct writer must supply the raw probability. That boundary conversion
|
||||
matters: a raw `80` in the YAML is two orders of magnitude above the maximum
|
||||
and would make every real score read as below threshold, since no probability
|
||||
exceeds 1.0.
|
||||
|
||||
That validator exists because of a **production near-miss** (issue #47). The
|
||||
admin UI used to save the threshold as written; an `80` was entered, the UI
|
||||
wrote `80.0` straight into the overlay, and it was only caught because the
|
||||
service had not yet restarted and loaded the bad value. Once restarted,
|
||||
classification would have failed on *every* request. The UI now converts
|
||||
0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line
|
||||
of defense for every other caller.
|
||||
|
||||
**Candidate labels are natural-language descriptions, never the raw config
|
||||
identifier.** HF's zero-shot pipeline scores a candidate label against the
|
||||
input via a hypothesis template (`"This example is {}."`), so the label
|
||||
itself has to read as a sentence for entailment scoring to work — feeding it
|
||||
`tool_use_agentic` or `diff_checking` verbatim asks the model to judge
|
||||
`"This example is tool_use_agentic."`, not a sentence its NLI training ever
|
||||
saw. `local_encoder.py`'s `_CATEGORY_DESCRIPTIONS` maps each category to a
|
||||
real description before scoring, and `multi_label=True` scores each
|
||||
candidate independently rather than normalizing every score to sum to 1 (the
|
||||
pipeline's single-label default, which drags a genuinely good match down
|
||||
whenever another category is also plausible). Measured live 2026-09-06 on the
|
||||
same 9-prompt set: raw identifiers scored 5/9 correct at confidence
|
||||
0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A
|
||||
category missing from the map falls back to its own raw string rather than
|
||||
raising, so a newly-added config category degrades gracefully instead of
|
||||
crashing.
|
||||
|
||||
`local_encoder` runs on **cpu by default** in the shipped example; setting it
|
||||
to `cuda` requires a `config.local.yaml` overlay and a host that actually has
|
||||
the hardware. The project does not ship a settled default device. Match the
|
||||
device to the machine the router is running on, otherwise `torch` will fail
|
||||
fast with a clear error at first classification. CPU inference is the safe,
|
||||
portable starting point; CUDA is host-dependent and should be chosen only
|
||||
where the router process has a card available.
|
||||
|
||||
`transformers`/`torch` are optional (`requirements-encoder.txt`, not the main
|
||||
`requirements.txt` — pinned-dependency discipline applies here too) and
|
||||
imported lazily inside `local_encoder.py`, the same rule `tui.py` follows for
|
||||
|
||||
@@ -60,12 +60,15 @@ below): prune → persist → `_check_pinned_capabilities`.
|
||||
|
||||
Pinch's tailoring is deliberate and safe:
|
||||
|
||||
- **Only `tool` results** older than the protected window are trimmed or
|
||||
summarized. Tool results are where a long agent session's tokens actually
|
||||
live, they are the least likely to still be needed in full by the time a
|
||||
later turn is answered, and replacing one with a short placeholder is
|
||||
reversible at the semantic level — a wrong guess costs context but never
|
||||
breaks the request.
|
||||
- **Only `tool` results** are trimmed or summarized. Older results outside the
|
||||
protected window are the main target, but any **outsized result inside the
|
||||
window** that is longer than `pinch.protected_max_chars` is also capped. That
|
||||
threshold is deliberately a much higher bar than `max_summarize_chars`; recent
|
||||
results are more likely to still matter, so it only catches true outliers.
|
||||
Tool results are where a long agent session's tokens actually live, they are
|
||||
the least likely to still be needed in full by the time a later turn is
|
||||
answered, and replacing one with a short placeholder is reversible at the
|
||||
semantic level — a wrong guess costs context but never breaks the request.
|
||||
- **Never user / assistant / system messages.** Trimming one of those can
|
||||
change what the model is being asked, so they are always kept verbatim.
|
||||
- **Never the classifier input.** The classification turn uses
|
||||
@@ -82,6 +85,22 @@ valid conversation.
|
||||
Only runs (and only mutates anything) when the estimate actually exceeds the
|
||||
budget; otherwise the original list is returned untouched.
|
||||
|
||||
### Capping outsized results in the protected window
|
||||
|
||||
The protected window, controlled by `keep_last_turns`, is intentionally generous:
|
||||
anything in the last few user turns stays verbatim because it is most likely to
|
||||
still be needed. But that created a gap. An autonomous tool-call loop (code
|
||||
execution, file reads, test runs) can produce one enormous tool result without a
|
||||
new user message, and that whole result used to count as part of the protected
|
||||
window. Pinch only fired when the total conversation crossed the budget, but
|
||||
once it did, those outsized protected results shipped verbatim regardless of
|
||||
size. In one measured case a 324k-token conversation shrank only ~8% because
|
||||
nearly all of it sat inside the protected window. `protected_max_chars` closes
|
||||
that gap: tool results inside the window are still preserved unless they cross
|
||||
this outlier threshold, in which case they receive the same head/tail elision
|
||||
applied to older candidates. The default is high enough to avoid touching normal
|
||||
output; it only intervenes when a single result threatens the budget.
|
||||
|
||||
## Config block
|
||||
|
||||
All pinch configuration lives under `pinch:` in `config/config.yaml`, validated
|
||||
|
||||
Reference in New Issue
Block a user