Files
6krrt/evals
adlee-was-taken 35a593942d fix(evals,seed): unanswerable diff_checking question + two seed-script defects
Three defects, all found by actually running the local-dispatch runbook for
the first time. None could have been caught offline.

1. diff_checking was logically unanswerable (evals/tasks.yaml)

   All four safe/buggy pairs are exact MIRROR IMAGES: BEFORE_safe ==
   AFTER_buggy and AFTER_safe == BEFORE_buggy. The question asked "Does the
   AFTER version change behaviour for any valid input?" -- which is SYMMETRIC:
   if A->B changes behaviour, so does B->A. But the pairs carry OPPOSITE
   labels, so four of the eight tasks were wrong no matter what any model
   answered.

   Measured before the fix: the maximum achievable score was 0.50, and a model
   that blindly answered "no" also scored 0.50. deepseek-v4-flash, which scores
   1.00 on all three coding categories, got 0.25 -- punished for engaging with
   the question. After the fix it scores 0.625 and nemotron-mini:4b's true
   profile is visible (0/4 bugs detected).

   The question is now antisymmetric ("does the AFTER version introduce a bug
   that the BEFORE version does not have?"), which is what opposite labels
   require. tests/test_task_set.py grows a regression test that pins the
   invariant; it was verified to FAIL against the old phrasing, not merely to
   pass against the new one.

   This is the fifth harness bug in this project that scored the rig rather
   than the model.

2. seed_local_dispatch_energy closed its DB connection mid-run

   main() closed conn right after reading the catalog, then used it four more
   times. Every run died on the first sample with "Cannot operate on a closed
   database". The step sits behind the user-tariff gate, so it had never been
   executed and the defect shipped unseen.

3. seed_local_dispatch_energy read its measurement one line too early

   ctx.avg_power_watts was read INSIDE the `with measure(...)` block, but
   measure finalizes on __exit__ (that is where the sampler thread is joined
   and the average computed). It was therefore always None, and the script
   wrote cost_per_1m_prompt = cost_per_1m_completion = $0.0000 -- pricing local
   compute as FREE, the exact failure this feature exists to prevent.

   With both fixed, the measured rates on a Quadro RTX 6000 at $0.159/kWh are
   $0.0054/1M prompt and $0.2286/1M completion (r^2 = 0.9985).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-03 19:33:14 -04:00
..