Review follow-up to #62: docstring indentation and invisible U+202F spaces #63
@@ -72,10 +72,10 @@ def recompute_category(
|
||||
For each row:
|
||||
1. Re-compute benchmark via ``blend()`` on stored leaderboard / self_eval
|
||||
components. If neither component exists the row skips the conversion.
|
||||
2. Collect *all* rows with ``outcome_samples > 0`` and compute the
|
||||
category-level *peer_rate* (sample-weighted pooled rate:
|
||||
Σoutcome_samples × outcome_score / Σoutcome_samples) and
|
||||
*peer_benchmark* (mean benchmark_score over the same set).
|
||||
2. Collect *all* rows with ``outcome_samples > 0`` and compute the
|
||||
category-level *peer_rate* (sample-weighted pooled rate:
|
||||
Σoutcome_samples × outcome_score / Σoutcome_samples) and
|
||||
*peer_benchmark* (mean benchmark_score over the same set).
|
||||
3. Call ``expected_success_rate`` per row with the benchmark, outcome, and
|
||||
peer statistics. Write the final ``blended_score`` and ``source``.
|
||||
"""
|
||||
|
||||
Reference in New Issue
Block a user