Two measurement surfaces on /metrics. Neither is read by routing, and a test asserts that against the module source: a live incumbent-cache-pricing experiment is running, and a series that reached the ranker would confound it. cost_calibration -- routing.estimated_cost against the provider's own bill, per (provider, model), joined on request_id. The scale error is not the finding: a uniform overestimate reorders nothing, because the ranking is a comparison and every candidate moves together. The SPREAD does reorder, so that is the headline figure. Live over 168h, 4,152 joined requests: est/billed runs 1.62x (qwen/qwen3.6-35b-a3b, openrouter) to 12.86x (qwen3.6-35b-fast, neuralwatt) -- a 7.9x spread, wider than the 5x the brief was written against. The factors are REPORTED, not applied; whether they are stable enough to trust is the question this exists to answer, and this project has already mistaken one moment of a moving per-model quantity for a constant. The join needs the model as well as the request id. On the live database 5 of 4,143 rows pair a decision that selected qwen3.6-35b (neuralwatt) with a completion billed by qwen/qwen3.6-35b-a3b (openrouter) -- a cross-provider failover, and one model's estimate against another's bill. latency -- p50/p95 of router_wall_seconds and router_ttft_seconds, which landed recently and nothing read. These are the router's own clock, not the provider's duration_seconds, which is why OpenRouter is visible here at all: it reports no duration. z-ai/glm-5.3-flash, a model with "flash" in its name, measures p50 9.60s / p95 40.08s to first token against 1.58s / 3.09s for deepseek-v4-flash on NeuralWatt. Reported, never scored. The wall and TTFT sample counts are kept independent because TTFT is streaming-only by nature. No new warning class, deliberately. Every group is out of band on both series today, so a divergence warning would fire on all of them from the first run -- bare presence, which is the anti-pattern rejection_warnings exists to avoid -- and a warning is a form of trust these factors have not yet earned. Four knobs under objective, all report-only, all recorded in DELIBERATELY_NOT_IN_ADMIN with the reason: they change what /metrics shows, and not even what it warns about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
16 KiB
16 KiB