feat(classify): cascading fallback instead of a dumb static guess #34

Merged
alee merged 5 commits from feat/classifier-fallback-cascade into main 2026-09-05 21:25:29 +00:00
Owner

When the local classifier failed, the router jumped straight to a fixed
tier/category. Cheap and predictable — but it discards two signals the router
already holds, and stamps unrelated work with general_chat.

The cascade

local -> stale session cache -> session history -> optional cloud -> static

Steps 2 and 3 are free and local. Step 4 costs money and is disabled by
default
, so this is zero-cost until someone configures cloud_fallback.

Three things bound step 4's cost:

  • a global cooldown (30s) after any classifier failure — a sustained
    outage buys one cloud attempt per window across all requests, not one per
    request;
  • an account-refusal gate wired at both _account_level_refusal sites: an
    out-of-credit account makes a cloud classification a guaranteed wasted
    request, so the step is skipped;
  • the session_cache.put() gate now admits classifier_cloud, bounding a
    session to one cloud call per staleness window rather than one per turn.
    session_stale/session_history stay out — borrowed answers must not
    renew a staleness clock they never earned.

_session_history_lookup filters degraded sources, so one outage's guess
cannot propagate through every later turn of a session.

The part the review bought: attribution

A fallback-classified request is recorded as general_chat — which is a
fully scored category (13 proficiency rows live) and the configured
fallback_category. Without a guard, POST /outcome on mislabelled outage
traffic folds into proficiency(model, general_chat) and drags real scores
toward whatever was flowing while the classifier was down.

report_outcome now resolves classification_source and passes
model_attributable. feedback.py is untouched — its existing
AND model_attributable = 1 already implements keep-the-record,
don't-steer-routing.

Attributable: classifier, cached, override, classifier_cloud. Not:
fallback, session_stale, session_history. The asymmetry is deliberate —
a degraded source is good enough to route one visibly-flagged request, but a
proficiency score is consulted by every future request.

Fails open on an unknown decision row; an over-applied exclusion already
starved feedback once (client_capped).

Two deviations from the plan, both forced by the code

session_key travels in a ContextVar, not threaded through
dispatch -> route -> classify. The kwarg broke 78 tests that stub
route with a lambda — that is the signal a signature change here is a
public API change. logs.py already carries the trace id this way.

session_cache.classify_one returns a plain dict. The plan wanted it in
session_cache and that module import-free; returning a dispatcher type
breaks the second.

Verification

  • 1234 passed, local_energy.enabled both true and false.
  • Attribution tests: 9 fail on unmodified main, 6 pass (pre-existing).
  • Cascade tests assert each step with the earlier ones disabled, so a pass
    proves that step ran.
  • Root config/config.local.yaml byte-identical (cc3d4c2a…).
  • Evidence + notepads under .omo/.

⚠️ Merge-order note

PR #33 adds a tripwire that fails on any unregistered warning class. This
PR adds one (classifier_degradation_warning). Whichever merges second turns
the other red — that is the tripwire working. Fix is one line:

"classifier-degraded": r"came from a degraded source",

Do not disable the test.

When the local classifier failed, the router jumped straight to a fixed tier/category. Cheap and predictable — but it discards two signals the router already holds, and stamps unrelated work with `general_chat`. ## The cascade `local -> stale session cache -> session history -> optional cloud -> static` Steps 2 and 3 are free and local. Step 4 costs money and is **disabled by default**, so this is zero-cost until someone configures `cloud_fallback`. Three things bound step 4's cost: - a **global** cooldown (30s) after any classifier failure — a sustained outage buys one cloud attempt per window across *all* requests, not one per request; - an **account-refusal gate** wired at both `_account_level_refusal` sites: an out-of-credit account makes a cloud classification a guaranteed wasted request, so the step is skipped; - the `session_cache.put()` gate now admits `classifier_cloud`, bounding a session to one cloud call per staleness window rather than one per turn. `session_stale`/`session_history` stay **out** — borrowed answers must not renew a staleness clock they never earned. `_session_history_lookup` filters degraded sources, so one outage's guess cannot propagate through every later turn of a session. ## The part the review bought: attribution A fallback-classified request is recorded as `general_chat` — which is a **fully scored category** (13 proficiency rows live) *and* the configured `fallback_category`. Without a guard, `POST /outcome` on mislabelled outage traffic folds into `proficiency(model, general_chat)` and drags real scores toward whatever was flowing while the classifier was down. `report_outcome` now resolves `classification_source` and passes `model_attributable`. **`feedback.py` is untouched** — its existing `AND model_attributable = 1` already implements keep-the-record, don't-steer-routing. Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not: `fallback`, `session_stale`, `session_history`. The asymmetry is deliberate — a degraded source is good enough to route one visibly-flagged request, but a proficiency score is consulted by every future request. **Fails open** on an unknown decision row; an over-applied exclusion already starved feedback once (`client_capped`). ## Two deviations from the plan, both forced by the code **`session_key` travels in a ContextVar**, not threaded through `dispatch -> route -> classify`. The kwarg broke **78 tests** that stub `route` with a lambda — that is the signal a signature change here is a public API change. `logs.py` already carries the trace id this way. **`session_cache.classify_one` returns a plain dict.** The plan wanted it in `session_cache` *and* that module import-free; returning a dispatcher type breaks the second. ## Verification - **1234 passed**, `local_energy.enabled` both true and false. - Attribution tests: **9 fail on unmodified main**, 6 pass (pre-existing). - Cascade tests assert each step with the earlier ones disabled, so a pass proves that step ran. - Root `config/config.local.yaml` byte-identical (`cc3d4c2a…`). - Evidence + notepads under `.omo/`. ## ⚠️ Merge-order note PR #33 adds a tripwire that fails on any **unregistered** warning class. This PR adds one (`classifier_degradation_warning`). Whichever merges second turns the other red — that is the tripwire working. Fix is one line: ```python "classifier-degraded": r"came from a degraded source", ``` Do not disable the test.
alee added 3 commits 2026-09-05 06:49:28 +00:00
When the local classifier failed, the router jumped straight to a fixed
tier/category. Cheap and predictable, but it discards two signals the
router already holds, and stamps unrelated work with general_chat.

The cascade is local -> stale session cache -> session history -> optional
cloud -> static fallback. Steps 2 and 3 are free and local. Step 4 costs
money and is DISABLED by default, so the feature is zero-cost until
someone configures cloud_fallback.

Three things bound the cost of step 4:
- a global cooldown (default 30s) after any classifier failure, so a
  sustained outage buys at most one cloud attempt per window across ALL
  requests rather than one per request;
- an account-refusal gate wired at both _account_level_refusal sites: if
  the account is out of credit, a cloud classification is a guaranteed
  waste of a request, so the step is skipped entirely;
- the session_cache.put() gate now admits classifier_cloud, so a session
  pays for at most one cloud call per staleness window instead of one per
  turn. session_stale/session_history stay OUT of the cache: they are
  borrowed answers and must not renew a staleness clock they never earned.

_session_history_lookup filters degraded sources, so a previous outage's
guess cannot propagate through every later turn and masquerade as a real
classification.

Two deviations from the plan, both forced by the code:

session_key travels in a ContextVar rather than threaded through
dispatch -> route -> classify. Adding the kwarg to route() broke 78 tests
that legitimately stub it with a lambda -- that is the signal that a
signature change here is a public API change, and this is not one.
logs.py already carries the trace id exactly this way.

session_cache.classify_one returns a plain dict rather than a
Classification. The plan asked for it in session_cache AND for that module
to stay import-free; returning a dispatcher type would violate the second.
The caller builds the typed object.

Also makes _schema_minus_route_decisions strip every index on the table
rather than one by name -- the hand-kept name list drifted the moment this
change added an index, failing with an opaque "no such table".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
A fallback-classified request is recorded with task_category=general_chat.
general_chat is a FULLY SCORED category -- 13 proficiency rows on the live
DB, kimi-k2.7-code-fast at 0.95 -- and it is also the configured
fallback_category. So without a guard, POST /outcome reports on mislabelled
outage traffic fold into proficiency(model, general_chat) through
feedback.py and drag real scores toward whatever traffic happened to be
flowing while the classifier was down. That pollutes the one table the
router treats as ground truth.

report_outcome now resolves the decision's classification_source and passes
model_attributable accordingly. feedback.py is NOT touched: its existing
`AND model_attributable = 1` filter already implements keep-the-record,
don't-steer-routing -- the same treatment the tool-call false-failures got.

Attributable: classifier, cached, override, classifier_cloud -- the category
was measured on this request, or declared by a client who knows. Not
attributable: fallback (a guess), session_stale and session_history (a real
classification of a DIFFERENT request, already past the staleness bound the
design itself set).

The asymmetry is deliberate. A degraded source is good enough to route one
visibly-flagged request; a proficiency score is consulted by every future
request, so attribution must not trust more than routing does.

FAILS OPEN on an unknown decision row. Withholding ground truth requires
positive knowledge the category was a guess, and an over-applied exclusion
already starved feedback once (client_capped). The test pinning that passes
on main by design -- it guards a future over-correction, not a fixed defect.

Also adds a drift test: the exclusion set must equal the Classification
source Literal minus the trusted four, so a new degraded source cannot be
added without deciding its attribution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
The cascade makes a local-classifier outage survivable, which is the point
-- but survivable failures are exactly the ones that go unnoticed. Routing
keeps working on borrowed categories while nothing says the classifier has
been down since Tuesday. classifier_degradation_warning reports the share
of the last 24h that came from a degraded source.

Silent below degraded_warn_min decisions: a share computed over a handful
of requests is noise, and a warning that flaps on low traffic is one nobody
reads.

Also documents cooldown_seconds, degraded_warn_min, degraded_warn_threshold
and a commented cloud_fallback block in config/config.yaml. CLAUDE.md is
explicit that a knob living only as a Pydantic default is invisible to
anyone tuning it -- and cloud_fallback in particular needs to be visible
precisely because it is absent by default.

Adds tests/test_classifier_cascade.py: each cascade step asserted by making
the earlier ones fail, so a passing test proves that step ran rather than
that an earlier one produced the same answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
alee added 2 commits 2026-09-05 15:41:23 +00:00
PR #33's tripwire fails on any warning class scoring_coverage emits that
nothing displays. This branch adds one, so the tripwire fired on merge --
working exactly as designed.

The fix was three parts, not the one line predicted:

1. the matcher, "came from a degraded source";
2. the CFG stub gains a `classifier` attribute. It never had one, so the
   failure was an AttributeError rather than an unregistered-warning
   report. The real RouterConfig always has .classifier -- the minimal
   stub was simply unfaithful;
3. the chernobyl fixture seeds 20 fallback-classified decisions so the
   class actually fires.

Step 3 needed a correction of its own. Seeding those rows with a SMALL
context silenced the escalation-hazard class, because 20 low rows drag the
tier-1 p95 below the tier-2 ceiling. They now carry the same context as the
demand row, leaving the p95 unmoved.

Worth recording: the chernobyl fixture is a coupled system. A new seed can
silence an existing class, and only the "every registered class fires"
direction catches it -- which is the argument for keeping both directions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
alee merged commit 12615d9e9c into main 2026-09-05 21:25:29 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#34