plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
139 lines
7.0 KiB
Markdown
139 lines
7.0 KiB
Markdown
# Classifier failure should degrade gradually, not to a fixed guess
|
|
|
|
Status: done -- the four-step cascade in classify()
|
|
|
|
**Status: FINAL — decision-complete.** Written 2026-09-04.
|
|
|
|
## The gap
|
|
|
|
`classify()` (`src/dispatcher.py:~555-590`) has exactly one endpoint
|
|
(`_classifier_client()`, `cfg.classifier.base_url`) and one failure behaviour:
|
|
return `cfg.classifier.fallback_category` / `fallback_tier` — currently
|
|
`general_chat` / tier 2 — flagged `source="fallback"`.
|
|
|
|
That is a fixed guess, and `general_chat` is the worst possible one:
|
|
`proficiency_score` is the ONLY category-dependent term in ranking, so a
|
|
request mislabelled `general_chat` loses the single signal that makes routing
|
|
category-aware. Every request in the outage window routes as if it were small
|
|
talk.
|
|
|
|
**The session cache cannot help, because of where it sits.**
|
|
`session_cache.get()` is called at `dispatcher.py:3077`, *before* classifying.
|
|
`classify()` resolves its own failure internally and never sees the session
|
|
key. So a session with 200 prior turns, every one classified
|
|
`coding_refactor`, still degrades to `general_chat` the moment the local model
|
|
is unavailable.
|
|
|
|
## Why this matters now, and why it hasn't yet
|
|
|
|
`source='fallback'` appears **zero times in 16,744 recorded decisions**
|
|
(`cached` 11,089, `classifier` 5,631, `override` 24). The path is real but
|
|
cold — which is precisely why it should be hardened rather than trusted:
|
|
untested code that only runs during an outage is the worst kind.
|
|
|
|
And it is about to get warmer. The operator stops Ollama to play games on the
|
|
same GPU. That is a deliberate, recurring, whole-session outage of the local
|
|
classifier — exactly the condition this path exists for.
|
|
|
|
## The fix: a cascade, cheapest first
|
|
|
|
Replace the single static fallback with an ordered cascade. Each step is tried
|
|
only if the previous failed, and each records a **distinct**
|
|
`classification_source` so the fallback mix is visible in `route_decisions`
|
|
rather than collapsed into one bucket.
|
|
|
|
1. **Local classifier** — unchanged. `source='classifier'`.
|
|
2. **Stale session reuse** — `source='session_stale'`. On classifier failure,
|
|
reuse this session's most recent classification **ignoring the staleness
|
|
bound**. A stale but real classification of the same conversation beats a
|
|
fixed guess, and this is free: no network, no tokens, works when everything
|
|
external is down.
|
|
3. **Durable session history** — `source='session_history'`. If the in-memory
|
|
cache has nothing (it is module-level and resets on restart, so a router
|
|
restart during an Ollama outage empties it), read the most recent
|
|
`task_category` / `task_tier` for this `session_key` from
|
|
`route_decisions`. One indexed query on a table that is already written on
|
|
every decision.
|
|
4. **Cloud classifier** — `source='classifier_cloud'`. Optional, off unless
|
|
configured. See below.
|
|
5. **Static fallback** — unchanged, `source='fallback'`. Last resort only.
|
|
|
|
Steps 2 and 3 are the valuable ones: free, offline, and they cover the actual
|
|
scenario (a running session whose local model went away mid-conversation).
|
|
Step 4 covers a *new* session started during an outage, which steps 2-3 cannot.
|
|
|
|
## The cloud classifier step
|
|
|
|
`cfg.classifier` already accepts any OpenAI-compatible endpoint, so this is a
|
|
**second** endpoint plus failover, not new transport.
|
|
|
|
It is worth having, and it is cheap. Measured previously and recorded in
|
|
`CLAUDE.md`: `deepseek-v4-flash` classifying the same prompts returned
|
|
**1.02s mean, 5/5 categories correct, 0 hard failures**, at **$0.093 per 1,000
|
|
calls** — 0.19% of the plan allowance. Against a local `qwen3.5` that averaged
|
|
11.58s with one hard failure.
|
|
|
|
Requirements:
|
|
|
|
- New optional block (e.g. `classifier.cloud_fallback`) with its own
|
|
`base_url`, `model`, `api_key_env`, `timeout_seconds`. **Default disabled** —
|
|
it spends the user's credits, so it must be opt-in.
|
|
- **Do not reuse the primary `classifier` block's fields.** The whole point is
|
|
a different endpoint.
|
|
- `max_retries=0`, same as the primary. The primary's timeout already becomes a
|
|
3x wall-clock bound with SDK retries on; do not repeat that mistake.
|
|
- Its timeout must be **short**. This runs only after the local attempt already
|
|
failed, so the request has already spent that budget. A slow cloud fallback
|
|
turns one bad request into two.
|
|
- **Never fall back to cloud when the account is out of credits.** Plan 5 added
|
|
`_account_level_refusal`; a cloud classifier call during a credit outage
|
|
wastes latency to reach a foregone 4xx. Skip step 4 if a recent
|
|
account-level refusal is known.
|
|
|
|
## Observability
|
|
|
|
The distinct `classification_source` values are the point. Today every
|
|
degradation is one bucket, so an operator cannot tell "Ollama is down but
|
|
sessions are being reused correctly" from "everything failed, we are guessing".
|
|
|
|
- `metrics.scoring_coverage` warnings: warn when the non-`classifier`,
|
|
non-`cached` share of recent decisions exceeds a threshold. That is the
|
|
signal that the classifier is down, and nothing currently reports it.
|
|
- Surface the source mix on the admin dashboard. `decisions.html` already
|
|
displays `source`; the values just need to be recognised and styled.
|
|
|
|
## Non-goals
|
|
|
|
- Do not change `fallback_category` / `fallback_tier` semantics or defaults.
|
|
They remain the last resort.
|
|
- Do not make the cloud classifier the primary. Local-first is the project's
|
|
premise, and the local model currently wins on latency (1.07s vs 1.02s is
|
|
parity, and local costs no credits).
|
|
- Do not classify with a *dispatch* model as a side effect. Step 4 is an
|
|
explicit configured endpoint, not an opportunistic reuse of a routed call.
|
|
- Do not extend the session-cache staleness bound for the normal path. Step 2
|
|
ignores staleness ONLY on the failure path; normal operation keeps its
|
|
existing freshness guarantee.
|
|
|
|
## Success criteria
|
|
|
|
- A test with the classifier raising and a populated session cache yields the
|
|
session's prior category with `source='session_stale'`, not `general_chat`.
|
|
- A test with the classifier raising and an EMPTY in-memory cache but existing
|
|
`route_decisions` rows for that `session_key` yields
|
|
`source='session_history'`.
|
|
- A test with the classifier raising, no session data, and cloud fallback
|
|
disabled yields `source='fallback'` — the current behaviour, preserved.
|
|
- A test with cloud fallback configured and the local classifier raising yields
|
|
`source='classifier_cloud'`.
|
|
- Cloud fallback is skipped when a recent account-level refusal is known.
|
|
- A warning fires when the degraded share of recent decisions crosses the
|
|
threshold.
|
|
- Full suite green with `local_energy.enabled` both true and false.
|
|
- The user's `config/config.local.yaml` is byte-identical after the run
|
|
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). **Corrected
|
|
2026-09-05:** this previously named `config/config.yaml`. PR #26 moved
|
|
deployment values into the gitignored overlay, so `config/config.yaml` is
|
|
clean and tracked, and editing it is an ordinary commit. The file that
|
|
cannot be recovered is the overlay — it is not in git history at all.
|