12 of 30 active OpenRouter rows had an effective context window of 0 and
had never been selected once across 23,000+ decisions -- including
moonshotai/kimi-k3 and z-ai/glm-5.3 at 1M advertised context.
The two providers do not mean the same thing by max_output_tokens.
NeuralWatt reports a genuine per-request output cap: at most 0.16 of
context across all 19 rows. OpenRouter reports
top_provider.max_completion_tokens, which is a ceiling on what you may
ASK for -- 0.8-0.9 of context on a dozen rows.
effective_context_window subtracted it whole:
int(1048576 * 0.75) - 943718 == -157286
and max(_, 0) turned that into 0, which routing.py's context filter
reads as "fits nothing". Nothing warned, because a row that is never
eligible never rejects anything either -- neither the capability-ceiling
detector nor the rejection detector can see a candidate that silently
fails a hard filter.
context.max_output_reserve_fraction (0.5) caps the catalog-derived
reserve at a fraction of the usable window. A per_model_overrides
reserve is a measurement rather than a parsed field, so it is still used
as written.
Simulated over the live catalog: exactly those 12 rows recover, plus a
7% nudge on deepseek/deepseek-v3.2 (57344 -> 61440). No NeuralWatt row
moves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
9.1 KiB
9.1 KiB