Files
6krrt/seed_energy.py
adlee-was-taken 7965394b70 docs: comments that outlived their measurements
This project's comments carry the reasoning, which makes a stale one worse
than none: it is a confident account of a decision that was reversed. CLAUDE.md
already records the reversals correctly, so these were the copies that
disagreed with it, and a reader has no way to tell which one is current.

schema.sql said "That product is what scoring ranks on" about
avg_power_watts * duration_seconds. Scoring reads the ATTRIBUTED figures, and
ranking on the pre-attribution product was tried and rejected -- it discards a
750x between-model signal to suppress a 1.8x within-model one. seed_energy.py's
docstring carried the same claim, and also still described the sweep as feeding
cost; it feeds eco and the energy ceiling now.

load_candidates' docstring described `cost` as the sweep's median.
rank_candidates overwrites it with routing.estimated_cost, priced from catalog
prices scaled to the request -- the measured value survives only as a fallback
for a row the catalog has no price for. The /health warning about missing sweep
data said the same thing and is corrected the same way.

routing.py and scoring.py both still stated the retired weighted composite as
the scoring model. scoring.py additionally now says which of its functions the
router actually calls: eco_score and composite_score are arithmetic with tests
and no callers.

config.yaml's "no flex discount knob" note justified itself with cost scoring
that no longer works that way; the real reason is that latency_tolerance is a
hard filter, and flex and standard carry the same catalog price anyway.

Test counts in README.md and CLAUDE.md were 264/14 files; 320 across 18 now.

CLAUDE.md's "until several sweeps have accumulated" section gets the correction
it needs most: none have. The sweep has been dying on its first billed call
since 6e729ad, so that accumulation starts from the next run, not from months
of history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WSkcSD2Jzkxo1Kw27ecfXJ
2026-08-22 21:00:12 -04:00

263 lines
11 KiB
Python

#!/usr/bin/env python3
"""Seed `energy_observations` by running a fixed reference task per model.
Why this exists: `eco_score` returns the neutral 0.5 for every candidate
until a model has observations, and observations only accrue from real
traffic. That is a cold start the router cannot route its way out of — every
model looks identical on eco, so eco contributes nothing to the very
decisions that would generate the data. This sweep breaks the cycle.
Method: one fixed prompt, one fixed `max_tokens`, `temperature=0`, run N
times per model. The prompt is held constant so the resulting figures are
comparable *across* models rather than reflecting who happened to be asked a
harder question.
N > 1 for a specific reason. The provider's billed `energy_kwh` is
`avg_power_watts * duration_seconds * attribution_ratio`, and that last term
is the request's share of a shared multi-tenant GPU pool — eight identical
calls to one model inside one minute spanned 20x, correlating +0.997 with
the attribution ratio while power and duration held steady. Two earlier
sweeps of these same 13 models disagreed by up to 36x on that basis.
Scoring reads the ATTRIBUTED figures, not `power * duration`. Ranking on the
pre-attribution product was tried and is wrong: the median attribution ratio
spans 750x between models against a 1.8x spread within one, so stripping it
discards the larger real signal to suppress the smaller noisy one. The summary
still prints both spreads side by side, because the gap between them is what
that reasoning rests on.
What this sweep feeds today is `eco` and the per-request energy ceiling
(objective.max_energy_per_request). It no longer feeds `cost`: cost is priced
per request from catalog prices scaled to the request's shape, because a fixed
400-token workload ranks models backwards for real traffic.
Coverage across TIME is the point of running it repeatedly. Attribution tracks
pool load and pool load tracks time of day -- between two sweeps hours apart,
deepseek-v4-flash moved ~50x and qwen3.6-35b ~7x the other way, enough to
invert their ranking. More samples inside one sweep measures one moment more
precisely; `load_candidates` takes the median over ALL seed_reference rows, so
repeated sweeps accumulate into a median across time for free.
What lands in the table are real observations of real calls, identical in
kind to what the dispatcher logs; they are simply generated deliberately
rather than incidentally. They carry `task_category='seed_reference'` so
they can be identified or purged later.
Usage:
python seed_energy.py # 5 samples of every routable model
python seed_energy.py --samples 3
python seed_energy.py --models kimi-k3,gemma-4-31b
python seed_energy.py --dry-run # show the plan and cost, call nothing
"""
from __future__ import annotations
import argparse
import os
import sqlite3
import statistics
import sys
import time
import requests
from config import load_config
from dispatcher import extract_telemetry, gross_energy_kwh, log_observation
# Held constant across models so the energy numbers are comparable. Long
# enough that most models run to the token cap rather than stopping early
# (a short answer burns less energy and would look misleadingly efficient),
# and generic enough that no model is advantaged by domain fit.
REFERENCE_PROMPT = (
"Explain how a B-tree works, including its structure, how lookups "
"descend the tree, and how splits keep it balanced on insert."
)
REFERENCE_MAX_TOKENS = 400
SEED_CATEGORY = "seed_reference"
def routable_models(conn: sqlite3.Connection, allowed_levels: list[str]) -> list[dict]:
placeholders = ",".join("?" * len(allowed_levels))
rows = conn.execute(
f"""
SELECT model_id, provider, latency_class, tier, cost_per_1m_completion
FROM models
WHERE access_level IN ({placeholders})
AND availability = 'active'
ORDER BY model_id
""",
tuple(allowed_levels),
).fetchall()
return [dict(r) for r in rows]
def sample_once(
base_url: str,
api_key: str,
model_id: str,
max_tokens: int = REFERENCE_MAX_TOKENS,
timeout: int = 300,
) -> dict:
"""One reference call. ``max_tokens`` is a parameter because --max-tokens is.
It used to be hardcoded while the banner printed whatever --max-tokens had
been passed, so the flag moved the report and not the request -- and the
resulting rows landed in the same seed_reference median as the 400-token
ones, quietly mixing two workload shapes in the one axis that exists to
hold the workload constant across models.
"""
resp = requests.post(
f"{base_url}/chat/completions",
headers={"authorization": f"Bearer {api_key}"},
json={
"model": model_id,
"messages": [{"role": "user", "content": REFERENCE_PROMPT}],
"max_tokens": max_tokens,
"temperature": 0,
},
timeout=timeout,
)
resp.raise_for_status()
return resp.json()
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--samples", type=int, default=5, help="samples per model (default 5)")
ap.add_argument("--models", help="comma-separated model_ids; default is all routable")
ap.add_argument("--max-tokens", type=int, default=REFERENCE_MAX_TOKENS)
ap.add_argument("--dry-run", action="store_true", help="print the plan, call nothing")
args = ap.parse_args()
cfg = load_config("config.yaml")
conn = sqlite3.connect(cfg.database.path)
conn.row_factory = sqlite3.Row
models = routable_models(conn, cfg.routing.allowed_access_levels)
conn.close()
if args.models:
wanted = {m.strip() for m in args.models.split(",")}
models = [m for m in models if m["model_id"] in wanted]
missing = wanted - {m["model_id"] for m in models}
if missing:
print(f"not routable / unknown: {', '.join(sorted(missing))}", file=sys.stderr)
if not models:
print("no models to sweep", file=sys.stderr)
return 1
total_calls = len(models) * args.samples
print(f"{len(models)} models x {args.samples} samples = {total_calls} calls")
print(f"prompt: {REFERENCE_PROMPT[:60]}... (max_tokens={args.max_tokens}, temperature=0)")
if args.dry_run:
for m in models:
print(f" {m['model_id']:26s} tier {m['tier']} {m['latency_class']}")
return 0
settings = cfg.dispatch_providers["neuralwatt"]
api_key = os.environ.get(settings.api_key_env)
if not api_key:
print(f"{settings.api_key_env} is not set", file=sys.stderr)
return 1
results: dict[str, list[dict]] = {}
allowance_start = allowance_end = None
for m in models:
model_id = m["model_id"]
results[model_id] = []
for i in range(args.samples):
try:
payload = sample_once(
settings.base_url, api_key, model_id, args.max_tokens
)
except requests.RequestException as e:
print(f" {model_id:26s} sample {i + 1}: FAILED {type(e).__name__}: {e}")
continue
usage = payload.get("usage") or {}
telemetry = extract_telemetry(payload)
# Keyword arguments deliberately. log_observation grew request_id,
# session_key and session_dir in the middle of its signature and
# made the trailing three keyword-only; this call still passed six
# positionals, so every sweep died on TypeError after its first
# BILLED call -- and TypeError is not a RequestException, so the
# `except` below never caught it.
log_observation(
model_id,
m["provider"],
SEED_CATEGORY,
payload.get("id"),
prompt_tokens=usage.get("prompt_tokens"),
completion_tokens=usage.get("completion_tokens"),
telemetry=telemetry,
)
if telemetry.allowance_remaining_usd is not None:
if allowance_start is None:
allowance_start = telemetry.allowance_remaining_usd
allowance_end = telemetry.allowance_remaining_usd
gross = None
if telemetry.avg_power_watts and telemetry.duration_seconds:
gross = gross_energy_kwh(
telemetry.avg_power_watts, telemetry.duration_seconds
)
results[model_id].append(
{
"completion_tokens": usage.get("completion_tokens"),
"energy_kwh": telemetry.energy_kwh,
"gross_kwh": gross,
"carbon": telemetry.carbon_g_co2eq,
"cost": telemetry.cost_usd,
}
)
# Be a considerate neighbour on a shared endpoint.
time.sleep(0.3)
done = [d for d in results[model_id] if d["gross_kwh"]]
if done:
med = statistics.median(d["gross_kwh"] for d in done)
print(
f" {model_id:26s} {len(done)}/{args.samples} ok "
f"median gross {med:.3e} kWh"
)
# --- summary ---------------------------------------------------------
# Two spread columns, because the whole point of this sweep is that they
# differ: the billed figure carries a multi-tenancy attribution term that
# the pre-attribution product does not.
print()
print(
f"{'model':26s}{'n':>3}{'gross kWh':>12}{'gross spr':>11}"
f"{'billed kWh':>12}{'billed spr':>12}{'gCO2eq':>11}"
)
for model_id, rows in results.items():
rows = [r for r in rows if r["gross_kwh"]]
if not rows:
print(f"{model_id:26s} 0 (no successful samples)")
continue
gross = [r["gross_kwh"] for r in rows]
billed = [r["energy_kwh"] for r in rows if r["energy_kwh"]]
carbons = [r["carbon"] for r in rows if r["carbon"] is not None]
def spread(xs):
return max(xs) / min(xs) if xs and min(xs) else float("nan")
print(
f"{model_id:26s}{len(rows):>3}{statistics.median(gross):>12.3e}"
f"{spread(gross):>10.1f}x{statistics.median(billed) if billed else 0:>12.3e}"
f"{spread(billed):>11.1f}x"
f"{statistics.median(carbons) if carbons else 0:>11.2e}"
)
if allowance_start is not None and allowance_end is not None:
print()
print(
f"allowance: {allowance_start:.6f} -> {allowance_end:.6f} USD "
f"(spent {allowance_start - allowance_end:.6f})"
)
return 0
if __name__ == "__main__":
raise SystemExit(main())