feat(classify): cascading fallback instead of a dumb static guess #34

Merged
alee merged 5 commits from feat/classifier-fallback-cascade into main 2026-09-05 21:25:29 +00:00

5 Commits

Author SHA1 Message Date
adlee-was-taken
9b0712f12e test(tui): register the classifier-degradation warning class
PR #33's tripwire fails on any warning class scoring_coverage emits that
nothing displays. This branch adds one, so the tripwire fired on merge --
working exactly as designed.

The fix was three parts, not the one line predicted:

1. the matcher, "came from a degraded source";
2. the CFG stub gains a `classifier` attribute. It never had one, so the
   failure was an AttributeError rather than an unregistered-warning
   report. The real RouterConfig always has .classifier -- the minimal
   stub was simply unfaithful;
3. the chernobyl fixture seeds 20 fallback-classified decisions so the
   class actually fires.

Step 3 needed a correction of its own. Seeding those rows with a SMALL
context silenced the escalation-hazard class, because 20 low rows drag the
tier-1 p95 below the tier-2 ceiling. They now carry the same context as the
demand row, leaving the p95 unmoved.

Worth recording: the chernobyl fixture is a coupled system. A new seed can
silence an existing class, and only the "every registered class fires"
direction catches it -- which is the argument for keeping both directions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 11:41:19 -04:00
adlee-was-taken
ebbd8d3193 Merge remote-tracking branch 'origin/main' into feat/classifier-fallback-cascade 2026-09-05 11:36:32 -04:00
adlee-was-taken
9e4ca52602 feat(metrics): classifier degradation warning, and document the knobs
The cascade makes a local-classifier outage survivable, which is the point
-- but survivable failures are exactly the ones that go unnoticed. Routing
keeps working on borrowed categories while nothing says the classifier has
been down since Tuesday. classifier_degradation_warning reports the share
of the last 24h that came from a degraded source.

Silent below degraded_warn_min decisions: a share computed over a handful
of requests is noise, and a warning that flaps on low traffic is one nobody
reads.

Also documents cooldown_seconds, degraded_warn_min, degraded_warn_threshold
and a commented cloud_fallback block in config/config.yaml. CLAUDE.md is
explicit that a knob living only as a Pydantic default is invisible to
anyone tuning it -- and cloud_fallback in particular needs to be visible
precisely because it is absent by default.

Adds tests/test_classifier_cascade.py: each cascade step asserted by making
the earlier ones fail, so a passing test proves that step ran rather than
that an earlier one produced the same answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 02:48:39 -04:00
adlee-was-taken
2afdb40e2b fix(outcome): degraded classifications must not train proficiency
A fallback-classified request is recorded with task_category=general_chat.
general_chat is a FULLY SCORED category -- 13 proficiency rows on the live
DB, kimi-k2.7-code-fast at 0.95 -- and it is also the configured
fallback_category. So without a guard, POST /outcome reports on mislabelled
outage traffic fold into proficiency(model, general_chat) through
feedback.py and drag real scores toward whatever traffic happened to be
flowing while the classifier was down. That pollutes the one table the
router treats as ground truth.

report_outcome now resolves the decision's classification_source and passes
model_attributable accordingly. feedback.py is NOT touched: its existing
`AND model_attributable = 1` filter already implements keep-the-record,
don't-steer-routing -- the same treatment the tool-call false-failures got.

Attributable: classifier, cached, override, classifier_cloud -- the category
was measured on this request, or declared by a client who knows. Not
attributable: fallback (a guess), session_stale and session_history (a real
classification of a DIFFERENT request, already past the staleness bound the
design itself set).

The asymmetry is deliberate. A degraded source is good enough to route one
visibly-flagged request; a proficiency score is consulted by every future
request, so attribution must not trust more than routing does.

FAILS OPEN on an unknown decision row. Withholding ground truth requires
positive knowledge the category was a guess, and an over-applied exclusion
already starved feedback once (client_capped). The test pinning that passes
on main by design -- it guards a future over-correction, not a fixed defect.

Also adds a drift test: the exclusion set must equal the Classification
source Literal minus the trusted four, so a new degraded source cannot be
added without deciding its attribution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 02:43:39 -04:00
adlee-was-taken
2b91f114dc feat(classify): cascading fallback instead of a dumb static guess
When the local classifier failed, the router jumped straight to a fixed
tier/category. Cheap and predictable, but it discards two signals the
router already holds, and stamps unrelated work with general_chat.

The cascade is local -> stale session cache -> session history -> optional
cloud -> static fallback. Steps 2 and 3 are free and local. Step 4 costs
money and is DISABLED by default, so the feature is zero-cost until
someone configures cloud_fallback.

Three things bound the cost of step 4:
- a global cooldown (default 30s) after any classifier failure, so a
  sustained outage buys at most one cloud attempt per window across ALL
  requests rather than one per request;
- an account-refusal gate wired at both _account_level_refusal sites: if
  the account is out of credit, a cloud classification is a guaranteed
  waste of a request, so the step is skipped entirely;
- the session_cache.put() gate now admits classifier_cloud, so a session
  pays for at most one cloud call per staleness window instead of one per
  turn. session_stale/session_history stay OUT of the cache: they are
  borrowed answers and must not renew a staleness clock they never earned.

_session_history_lookup filters degraded sources, so a previous outage's
guess cannot propagate through every later turn and masquerade as a real
classification.

Two deviations from the plan, both forced by the code:

session_key travels in a ContextVar rather than threaded through
dispatch -> route -> classify. Adding the kwarg to route() broke 78 tests
that legitimately stub it with a lambda -- that is the signal that a
signature change here is a public API change, and this is not one.
logs.py already carries the trace id exactly this way.

session_cache.classify_one returns a plain dict rather than a
Classification. The plan asked for it in session_cache AND for that module
to stay import-free; returning a dispatcher type would violate the second.
The caller builds the typed object.

Also makes _schema_minus_route_decisions strip every index on the table
rather than one by name -- the hand-kept name list drifted the moment this
change added an index, failing with an opaque "no such table".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 02:40:07 -04:00