feat: local-encoder accuracy rebuild — eval harness, CLS pooling, trainable head #100
Reference in New Issue
Block a user
Delete Branch "feat/local-encoder-accuracy-rebuild"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What this PR does
Rebuilds the local-encoder (
classifier.mode: local_encoder) classificationpath for measurable accuracy. It ships the five spec items (B/A/C/D/E) from
plans/local-encoder-accuracy-rebuild.md, plus the docs and evidence:B — Eval harness (
src/eval_classifier.py)A standalone harness (modeled on
eval_proficiency.py, separate from thepytest suite; needs
requirements-encoder.txt) that measures zero-shotaccuracy on the 46-task eval set (
evals/tasks.yamlminus the 11tool_use_agenticrows the encoder cannot emit), with three noise variantsper prompt (clean, short-noise, long-noise). It prints top-1 accuracy, a full
confusion matrix, per-category precision/recall, confidence distribution split
by correct/incorrect, and the per-cell
blended_scoregap from the proficiencytable.
--dry-runplans the run without loading the model.A — CLS pooling + per-model query prefix
Replaces the body of
_mean_pool(never-optimised mean pooling) with per-modeldispatch that reads
1_Pooling/config.jsonfrom the model snapshot for thepooling strategy (CLS vs mean) and
tokenizer_config.json"prompts"for theper-model query prefix. This is the pooling both shipped backbones
(
bge-large-en-v1.5, GTE) specify, moving real-prompt accuracy from 0.304 →~0.391. Falls back to mean-pool when the snapshot file is absent. Function
signature and call sites unchanged.
C — Confidence knob rename (
confidence_threshold→confidence_min)The knob was documented as a probability but reads a softmax-amplified cosine
similarity with no probabilistic meaning — so
confidence_threshold: 0.5sent every real request below threshold. Renamed to
confidence_min(a minimumsimilarity score in [0.0, 1.0]) in
LocalEncoderConfig, with a loggeddeprecation alias for the old key.
D — Trainable head (
_TrainableHead)A convex
LogisticRegression(solver="lbfgs")over frozen embeddings, trainedby
scripts/train_encoder_head.pyon a committed syntheticcorpus (
evals/synthetic/encoder-training.jsonl), with per-class Platt(sigmoid) calibration. When
evals/synthetic/encoder-head-coefficients.jsonis present (and the backbone/candidate set match), it replaces nearest-centroid
at config time, giving the confidence gate a real
P(correct)meaning andfixing the
debuggingsink /file_summarization0/6 problem. Absent ormismatched artifact degrades to the unchanged centroid path with a logged
warning.
E — Tier-from-features sketch
_TierFeatureClassifiermodel-class stub plusclassifier.encoder.tier_from_featuresand
classifier.encoder.tier_feature_fieldsconfig keys — a concrete sketch,no training loop.
Evidence baseline
All evidence captured to
.omo/evidence/local-encoder-accuracy-rebuild/(gitignored):final-suite.txt2254 passed, 1 failed— the single failure is the pre-existingtest_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity(not touched by this PR)eval-dry-run.txtgit-log.txtDocs
docs/local-models.md—confidence_minrename,eval_classifier.pyharness,trainable-head section.
docs/evaluation.md— briefeval_classifier.pyharness section.Notes
transformers/torch;eval_classifier.pyand the trainer are standalone under
requirements-encoder.txt.