Three defects, all found by actually running the local-dispatch runbook for
the first time. None could have been caught offline.
1. diff_checking was logically unanswerable (evals/tasks.yaml)
All four safe/buggy pairs are exact MIRROR IMAGES: BEFORE_safe ==
AFTER_buggy and AFTER_safe == BEFORE_buggy. The question asked "Does the
AFTER version change behaviour for any valid input?" -- which is SYMMETRIC:
if A->B changes behaviour, so does B->A. But the pairs carry OPPOSITE
labels, so four of the eight tasks were wrong no matter what any model
answered.
Measured before the fix: the maximum achievable score was 0.50, and a model
that blindly answered "no" also scored 0.50. deepseek-v4-flash, which scores
1.00 on all three coding categories, got 0.25 -- punished for engaging with
the question. After the fix it scores 0.625 and nemotron-mini:4b's true
profile is visible (0/4 bugs detected).
The question is now antisymmetric ("does the AFTER version introduce a bug
that the BEFORE version does not have?"), which is what opposite labels
require. tests/test_task_set.py grows a regression test that pins the
invariant; it was verified to FAIL against the old phrasing, not merely to
pass against the new one.
This is the fifth harness bug in this project that scored the rig rather
than the model.
2. seed_local_dispatch_energy closed its DB connection mid-run
main() closed conn right after reading the catalog, then used it four more
times. Every run died on the first sample with "Cannot operate on a closed
database". The step sits behind the user-tariff gate, so it had never been
executed and the defect shipped unseen.
3. seed_local_dispatch_energy read its measurement one line too early
ctx.avg_power_watts was read INSIDE the `with measure(...)` block, but
measure finalizes on __exit__ (that is where the sampler thread is joined
and the average computed). It was therefore always None, and the script
wrote cost_per_1m_prompt = cost_per_1m_completion = $0.0000 -- pricing local
compute as FREE, the exact failure this feature exists to prevent.
With both fixed, the measured rates on a Quadro RTX 6000 at $0.159/kWh are
$0.0054/1M prompt and $0.2286/1M completion (r^2 = 0.9985).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U