Files
6krrt/tests/test_multi_provider_dispatch.py
adlee-was-taken af18009d59 feat(telemetry): a NULL cached_tokens now says which kind of nothing it was
Wave 1 item 1.2 of plans/token-waste-waves.md.

Three provider shapes collapsed into two stored outcomes, so "the provider
said zero" was indistinguishable from "the provider said nothing":

  details present, cached_tokens: 0   -> stored 0
  details present, no cached count    -> stored NULL
  no details block at all             -> stored NULL

That distinction is the entire question the item exists to settle. If the
field is omitted on a full cache miss, every cache rate computed from the
rows that carry it is conditioned on a hit having occurred and is biased
upward -- which would make the 0.924 figure on pruned turns meaningless.

`cached_tokens_source TEXT` records which shape produced the row:
'reported' | 'details_no_count' | 'no_details'. Additive, idempotent ALTER
alongside the existing one, NULL on existing rows -- which is honest, since
a row written before the column cannot say, and inferring one would
fabricate exactly the evidence being gathered. Invariant:
cached_prompt_tokens IS NOT NULL exactly when source = 'reported'. A JSON
null under an existing cached_tokens key groups with 'details_no_count',
not with a zero; no value is invented anywhere.

Two bugs fell out of writing it:

- `prompt_tokens_details: null` called .get on None and raised
  AttributeError inside the streamed finally block.
- cached tokens were read ONLY from the explicit `usage` argument, so every
  buffered call site (non-streaming dispatch, eval_proficiency, seed_energy)
  dropped the count even when the provider reported it. It now resolves its
  source the same way cost already did.

Measured read-only on the live DB before writing this: the 12.8% NeuralWatt
coverage in the plan is not sparse reporting. Of 425 post-restart NULL rows,
420 are seed_reference sweeps; real dispatch traffic carries a count on
282 of 287 rows, and 24 of those are an explicit 0. So the provider does
report zeros, and the upward-bias worry is much smaller than the plan
assumed -- but the column is what makes that checkable rather than argued.

7 new tests: the three shapes at parse level, JSON null, a non-dict details
block, the buffered payload path, and the three shapes asserted distinct as
STORED rows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 11:04:02 -04:00

25 KiB