feat: local-encoder accuracy rebuild — eval harness, CLS pooling, trainable head #100

Merged
alee merged 9 commits from feat/local-encoder-accuracy-rebuild into main 2026-09-23 20:23:58 +00:00
Owner

What this PR does

Rebuilds the local-encoder (classifier.mode: local_encoder) classification
path for measurable accuracy. It ships the five spec items (B/A/C/D/E) from
plans/local-encoder-accuracy-rebuild.md, plus the docs and evidence:

B — Eval harness (src/eval_classifier.py)

A standalone harness (modeled on eval_proficiency.py, separate from the
pytest suite; needs requirements-encoder.txt) that measures zero-shot
accuracy on the 46-task eval set (evals/tasks.yaml minus the 11
tool_use_agentic rows the encoder cannot emit), with three noise variants
per prompt (clean, short-noise, long-noise). It prints top-1 accuracy, a full
confusion matrix, per-category precision/recall, confidence distribution split
by correct/incorrect, and the per-cell blended_score gap from the proficiency
table. --dry-run plans the run without loading the model.

A — CLS pooling + per-model query prefix

Replaces the body of _mean_pool (never-optimised mean pooling) with per-model
dispatch that reads 1_Pooling/config.json from the model snapshot for the
pooling strategy (CLS vs mean) and tokenizer_config.json "prompts" for the
per-model query prefix. This is the pooling both shipped backbones
(bge-large-en-v1.5, GTE) specify, moving real-prompt accuracy from 0.304 →
~0.391
. Falls back to mean-pool when the snapshot file is absent. Function
signature and call sites unchanged.

C — Confidence knob rename (confidence_threshold → confidence_min)

The knob was documented as a probability but reads a softmax-amplified cosine
similarity with no probabilistic meaning — so confidence_threshold: 0.5
sent every real request below threshold. Renamed to confidence_min (a minimum
similarity score in [0.0, 1.0]) in LocalEncoderConfig, with a logged
deprecation alias for the old key.

D — Trainable head (_TrainableHead)

A convex LogisticRegression(solver="lbfgs") over frozen embeddings, trained
by scripts/train_encoder_head.py on a committed synthetic
corpus (evals/synthetic/encoder-training.jsonl), with per-class Platt
(sigmoid) calibration. When evals/synthetic/encoder-head-coefficients.json
is present (and the backbone/candidate set match), it replaces nearest-centroid
at config time, giving the confidence gate a real P(correct) meaning and
fixing the debugging sink / file_summarization 0/6 problem. Absent or
mismatched artifact degrades to the unchanged centroid path with a logged
warning.

E — Tier-from-features sketch

_TierFeatureClassifier model-class stub plus classifier.encoder.tier_from_features
and classifier.encoder.tier_feature_fields config keys — a concrete sketch,
no training loop.

Evidence baseline

All evidence captured to .omo/evidence/local-encoder-accuracy-rebuild/ (gitignored):

artifact result
final-suite.txt 2254 passed, 1 failed — the single failure is the pre-existing test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity (not touched by this PR)
eval-dry-run.txt 46 tasks × 3 noise levels = 138 classifications planned; no encoder load, no transformers/torch
git-log.txt 6 implementation commits + docs commit

Docs

  • docs/local-models.md — confidence_min rename, eval_classifier.py harness,
    trainable-head section.
  • docs/evaluation.md — brief eval_classifier.py harness section.

Notes

  • The dispatch path never imports transformers/torch; eval_classifier.py
    and the trainer are standalone under requirements-encoder.txt.
  • No binding of the encoder to production decisions, no restart of port 8080.
## What this PR does Rebuilds the local-encoder (`classifier.mode: local_encoder`) classification path for measurable accuracy. It ships the five spec items (B/A/C/D/E) from `plans/local-encoder-accuracy-rebuild.md`, plus the docs and evidence: ### B — Eval harness (`src/eval_classifier.py`) A standalone harness (modeled on `eval_proficiency.py`, separate from the pytest suite; needs `requirements-encoder.txt`) that measures zero-shot accuracy on the **46-task eval set** (`evals/tasks.yaml` minus the 11 `tool_use_agentic` rows the encoder cannot emit), with three noise variants per prompt (clean, short-noise, long-noise). It prints top-1 accuracy, a full confusion matrix, per-category precision/recall, confidence distribution split by correct/incorrect, and the per-cell `blended_score` gap from the proficiency table. `--dry-run` plans the run without loading the model. ### A — CLS pooling + per-model query prefix Replaces the body of `_mean_pool` (never-optimised mean pooling) with per-model dispatch that reads `1_Pooling/config.json` from the model snapshot for the pooling strategy (CLS vs mean) and `tokenizer_config.json` `"prompts"` for the per-model query prefix. This is the pooling both shipped backbones (`bge-large-en-v1.5`, GTE) specify, moving real-prompt accuracy from **0.304 → ~0.391**. Falls back to mean-pool when the snapshot file is absent. Function signature and call sites unchanged. ### C — Confidence knob rename (`confidence_threshold` → `confidence_min`) The knob was documented as a probability but reads a softmax-amplified cosine similarity with no probabilistic meaning — so `confidence_threshold: 0.5` sent every real request below threshold. Renamed to `confidence_min` (a minimum similarity score in [0.0, 1.0]) in `LocalEncoderConfig`, with a logged deprecation alias for the old key. ### D — Trainable head (`_TrainableHead`) A convex `LogisticRegression(solver="lbfgs")` over frozen embeddings, trained by `scripts/train_encoder_head.py` on a committed synthetic corpus (`evals/synthetic/encoder-training.jsonl`), with per-class Platt (sigmoid) calibration. When `evals/synthetic/encoder-head-coefficients.json` is present (and the backbone/candidate set match), it replaces nearest-centroid at config time, giving the confidence gate a real `P(correct)` meaning and fixing the `debugging` sink / `file_summarization` 0/6 problem. Absent or mismatched artifact degrades to the unchanged centroid path with a logged warning. ### E — Tier-from-features sketch `_TierFeatureClassifier` model-class stub plus `classifier.encoder.tier_from_features` and `classifier.encoder.tier_feature_fields` config keys — a concrete sketch, no training loop. ## Evidence baseline All evidence captured to `.omo/evidence/local-encoder-accuracy-rebuild/` (gitignored): | artifact | result | |---|---| | `final-suite.txt` | `2254 passed, 1 failed` — the single failure is the **pre-existing** `test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity` (not touched by this PR) | | `eval-dry-run.txt` | 46 tasks × 3 noise levels = 138 classifications planned; no encoder load, no transformers/torch | | `git-log.txt` | 6 implementation commits + docs commit | ## Docs - `docs/local-models.md` — `confidence_min` rename, `eval_classifier.py` harness, trainable-head section. - `docs/evaluation.md` — brief `eval_classifier.py` harness section. ## Notes - The dispatch path never imports `transformers`/`torch`; `eval_classifier.py` and the trainer are standalone under `requirements-encoder.txt`. - No binding of the encoder to production decisions, no restart of port 8080.
alee added 7 commits 2026-09-23 05:26:49 +00:00
alee added 2 commits 2026-09-23 20:20:26 +00:00
_mean_pool(token_embeddings, attention_mask) reads
list(_model_pooling_strategies.values())[-1] — the last-inserted
value — instead of the strategy belonging to the model whose
embeddings are being pooled. This works today because production
loads one encoder per process and every test populates at most one
dict entry, but breaks silently the moment two backbones share a
process.

Fix: thread model_id through _mean_pool and its call sites
(_build_description_embeddings and _embed_task), resolving the
strategy via _model_pooling_strategies.get(model_id, 'mean')
instead of reading the global tail.

Add test_mean_pool_dispatches_by_model_id_not_load_order which
proves the failure mode is closed: two models with different
strategies (cls, mean) are each pooled correctly regardless of
which was inserted first.
Both fields are read nowhere outside config.py; item E is a design
sketch only. An admin toggle would control nothing real until that
feature is implemented. The existing gap in
test_admin_knob_coverage.py (the whole classifier.* section sits
outside the generic registries) is pre-existing and not addressed
here.
alee merged commit f554fc17d0 into main 2026-09-23 20:23:58 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#100