"""Pure proficiency blending for the local LLM model router. Like ``scoring.py``, ``tiering.py`` and ``routing.py``, this module is free of I/O: scores and thresholds come in as arguments. ``eval_proficiency.py`` owns the DB writes and the model calls. Two independent sources feed one number per (model, category): - **leaderboard** โ€” a curated prior from published benchmarks. Its job is cold start: NeuralWatt adds models, and a newly listed one has no self-eval history at all. Without a prior it scores the neutral 0.5 and is indistinguishable from a model that was measured and found average. - **self-eval** โ€” this router's own task set, run against the real endpoint. More predictive of actual routing quality, but it accumulates slowly. Design doc ยง3.3 gives the blend as ``0.3 x leaderboard + 0.7 x self_eval`` once self-eval crosses ``self_eval_min_samples``, falling back to the leaderboard alone before that so thin, noisy self-eval data cannot dominate early. That rule assumes a leaderboard entry exists. It often will not โ€” the curated file is hand-maintained and NeuralWatt ships models faster than public benchmarks cover them. Taken literally, a model with no prior and 9 samples would score nothing at all, which is strictly worse than the 9 samples it actually has. So the fallback ladder is: leaderboard + enough self-eval -> weighted blend 'blended' enough self-eval, no prior -> self-eval alone 'self_eval' prior, not enough self-eval -> leaderboard alone 'leaderboard' thin self-eval, no prior -> self-eval alone 'self_eval_thin' neither -> None (neutral 0.5 downstream) ``self_eval_thin`` is deliberately distinguishable: it is real measurement, but from too few samples to trust as much as the label 'self_eval' implies, and a caller wanting to exclude it can. """ from __future__ import annotations from typing import Literal, Optional Source = Literal["blended", "self_eval", "leaderboard", "self_eval_thin"] def blend( leaderboard_score: Optional[float], self_eval_score: Optional[float], self_eval_samples: int, *, leaderboard_weight: float, self_eval_weight: float, min_samples: int, ) -> tuple[Optional[float], Optional[Source]]: """Combine the two sources into one score, with the source that produced it. Returns ``(None, None)`` when neither source has anything, which leaves the candidate on the neutral 0.5 in ``scoring.proficiency_score`` rather than penalizing it for being unmeasured. """ has_leaderboard = leaderboard_score is not None has_self_eval = self_eval_score is not None and self_eval_samples > 0 enough_samples = has_self_eval and self_eval_samples >= min_samples if has_leaderboard and enough_samples: return ( leaderboard_weight * leaderboard_score + self_eval_weight * self_eval_score, "blended", ) if enough_samples: return self_eval_score, "self_eval" if has_leaderboard: return leaderboard_score, "leaderboard" if has_self_eval: return self_eval_score, "self_eval_thin" return None, None def accumulate( previous_score: Optional[float], previous_samples: int, new_scores: list[float], ) -> tuple[Optional[float], int]: """Fold a fresh eval run into a running mean. Keeps a running average rather than replacing, so ``self_eval_samples`` means what the blending rule assumes it means: how much evidence stands behind the score. Re-running the harness therefore tightens an estimate instead of discarding everything learned before it. Returns ``(previous_score, previous_samples)`` unchanged when handed no new scores, so a run where every task errored cannot quietly reset a model's history to zero. """ if not new_scores: return previous_score, previous_samples total_samples = previous_samples + len(new_scores) if previous_score is None or previous_samples <= 0: return sum(new_scores) / len(new_scores), len(new_scores) weighted = previous_score * previous_samples + sum(new_scores) return weighted / total_samples, total_samples