docs: refresh CLAUDE.md for PRs #26-#34, and spec the admin portal uplift #35

Merged
alee merged 3 commits from docs/refresh-claude-md into main 2026-09-06 02:32:00 +00:00

3 Commits

Author SHA1 Message Date
adlee-was-taken
399934fd69 docs(plans): gaming mode requires a cloud classifier, not just a warning
User decision 2026-09-05. The earlier draft only warned when
classifier.cloud_fallback was absent, which would let gaming mode silently
turn the router into a worse one -- every request dropping to
stale-session, session-history, or a static guess.

Also corrects an easy assumption the draft invited. The cloud classifier
machinery IS built (PR #34) but two things stop it working today: the
config.yaml block is commented out, so cfg.classifier.cloud_fallback is
None and the cascade returns early; and it is step 3 of a cascade, reached
only after local fails AND both free session steps miss. Forcing cloud is a
promotion from fallback to primary, not a wiring job.

Keeps cascade steps 1 and 2 ahead of the cloud call and says why: a stale
session classification is free and was a real classification of that same
session, so paying to re-derive it is spending money for nothing. "Force
cloud" replaces the local model, not free correct answers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 19:44:21 -04:00
adlee-was-taken
d6d3f949ea docs(plans): admin portal uplift -- proficiency, profiles, gaming mode
Three gaps found by auditing the portal against the config surface: 5 pages
and 9 editable keys against 21 config sections, with seven sections having
no presence at all.

Proficiency is invisible despite being, per CLAUDE.md, the ONLY
category-dependent term in the ranking -- 148 rows over 17 models and 11
categories, surfaced as a single top-N list. The page must distinguish
measured scores from inherited ones: 71 of the 148 are outcome_prior and 39
self_eval_thin, and treating those as equal to outcome_blended is the
mistake this project already made with -flex rows. Read-only by design; a
hand-edited proficiency score is a fabricated measurement.

The profiles page is a dead end -- all five profiles are built-ins, which
render no action buttons, so nothing on the page is editable and it never
says why. Duplicate-to-edit fixes it with no backend change.

Its zero-admit alarm also cries wolf: locality reports "admits 0 models"
while an active ollama-local row exists, because _profile_probe passes
task_category=None and that model is gated to file_summarization and
diff_checking. A false alarm on the badge that exists to catch incident #3
is worse than no badge.

Gaming mode carries a real defect. _last_classifier_failure is written and
never read, so the circuit breaker does not break the circuit and nothing
stops the router retrying a dead local classifier -- up to 120s per request
against a hung or black-holed Ollama. Records that the toggle should be one
flag the code READS rather than a macro writing five keys, and that it must
skip the local call rather than fail and recover.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 19:42:16 -04:00
adlee-was-taken
3737bd4898 docs: bring CLAUDE.md up to date with PRs #26-#34
CLAUDE.md declares itself the file to trust on what is currently true, and
no commit touched it during the nine merges. Four shipped features were
absent entirely.

Adds: the classifier fallback cascade and, more importantly, the trap it
had to avoid -- a degraded classification lands on general_chat, which is a
fully scored category AND the fallback_category, so unguarded outcome
folding would drag real proficiency toward outage traffic. Records the
attributable/not-attributable split and the fail-open rule.

Adds the capability sub-ceiling and reactive rejection detectors to the
incident-3 section, since that incident recurred on 2026-09-04 through a
dimension the original detector did not model. Keeps the two details that
are easy to undo by accident: grouping on the structured task_tier column
rather than a normalized string, and the novelty-OR-rate signal, which
exists because routine rejections run ~3/hr while the incident was n=2 --
below the noise floor, so no single threshold works.

Notes the two drift tripwires and that the warning fixture is a coupled
system where a new seed can silence an existing class.

Marks the plan queue empty and records the two known follow-ups, including
that the NULL demand crash is latent only -- zero such rows on the live DB,
checked rather than assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 17:32:36 -04:00