# Degraded output as a routing signal Status: planned -- not built; prerequisite landed in `8518114`, second trigger measured 2026-09-08 Two measured instances of one thing: **a provider returning garbage the router can detect for free, with no way to act on it.** Mojibake was the first and turned out to be ours. Empty-response-on-200 is the second and is not -- see "The empty-200 incident" below, which is the stronger case for building this and changes what the trigger has to watch. ## Why this exists, and the thing it must not become The trigger was mojibake in agent output. Investigating it found the cause in the router, not the models: the SSE proxy decoded upstream bytes as latin-1 (requests falls back to ISO-8859-1 for any `text/*` without a charset, and OpenRouter's `text/event-stream` omits one) and re-encoded them as UTF-8, double-encoding every non-ASCII character in every streamed completion. `8518114` fixes that. **Read that again before building anything here.** Had a per-model "mangled output" score existed during that period, it would have recorded a failure against whichever model happened to be routed, and because the bug was triggered by one provider's Content-Type header it would have systematically scored every OpenRouter row worse than every NeuralWatt row, for a reason with nothing to do with the models. CLAUDE.md already lists four cases of "scored the rig rather than the model"; this would have been the fifth, and the first to bias a whole provider. So the governing constraint is not detection quality. It is attribution: **a mangled-output signal is only worth having if it can tell the model's fault from the pipeline's.** Everything below is shaped by that. ## What it is not - **Not a proficiency term.** Three reasons, each sufficient: - *Wrong grain.* `proficiency` is per `(model, category)`. Encoding breakage is not category-dependent, so the signal would be diluted across 11 buckets and need roughly 11x the samples to become visible. - *It reverses a decision already made.* Structural verdicts were deliberately demoted to diagnostics because "a structural failure is not a verified client outcome". An automatic checker feeding the score puts that straight back. - *Wrong instrument.* `proficiency` is a slow mean behind a 20-pseudo-observation prior. Encoding breakage is a step change -- a deploy, a proxy, a quantisation swap, or a router bug. A slow mean is late to notice it and slow to forgive it. - **Not a new objective or weight.** Ranking stays quality-first with cost as the tiebreak inside `quality_tolerance`. There is no knob to tune here. - **Not a reason to store completions.** This router stores no raw task text, and that does not change. Detection is inline; only a verdict is persisted. ## What it is The same shape as `circuit_breaker.py`: a **passive availability skip**, not a score. Mangled output is a health signal, usually transient and usually provider-side, so the right response is to route away now and re-probe soon -- exactly what the breaker already does for 5xx, cooldown and backoff included, clearing on the next clean response. ### 1. Detection (free, structural, inline) A new `mangled` verdict in `verification.py`, checked only when the response is otherwise checkable. Three signatures, in descending confidence: | signature | example | confidence | |---|---|---| | double-encoded UTF-8 | `c3 a2 c2 80 c2 94` (em dash); any `c3 83`/`c3 82` lead pair | high | | replacement character | `U+FFFD` | high | | non-ASCII absent from the input | smart quotes injected into ASCII source | low -- diagnostic only | Only the first two may trip anything. The third is recorded and surfaced but must never gate routing on its own: a model correctly answering in Japanese, or correctly quoting a non-ASCII identifier, produces it. ### 2. Attribution (the load-bearing part) `verify_response(text, finish_reason, has_tool_calls)` has no view of the prompt today, so this needs one new argument -- the input's charset profile, computed once per request. The rule: - prompt was pure ASCII and the completion carries `U+FFFD` or a double-encoded sequence -> **attributable** - prompt already carried non-ASCII -> **not attributable**, because a model echoing what it was given is behaving correctly This is the same lesson as `has_tool_calls` and `client_capped`: the checker is only safe once it knows what is not the model's fault. Both of those were learned by over-applying a verdict, once in each direction. Non-attributable rows are still written, with `model_attributable = 0` -- the treatment degraded-classification outcomes already get. The record survives; it just does not steer. ### 3. Storage (no migration) `verifications` already carries `kind`, `verdict`, `model_attributable` and `applied_at`. A `mangled` value in the existing `verdict` column and a `kind='structural'` row is the whole change. `Verdict` in `verification.py` gains one literal. `feedback.py` is untouched and must stay untouched: its existing `AND model_attributable = 1` plus the client-outcome filter already keep this out of `proficiency`. ### 4. Routing response Trip the breaker. Two decisions inside that: - **Scope: provider, not model.** The incident that motivated this was uniform across every OpenRouter model because the fault was one layer up. A per-model breaker would have tripped them one at a time, taking eleven separate incidents to describe one. Record the `model_id` on the row for forensics; key the trip on `provider`. `circuit_breaker._store` is keyed `(model_id, provider)` and lives in memory, so this needs a provider-level entry or a reserved key -- a small addition, not a redesign. **Scope is per-signature, not global to this feature.** The empty-200 incident below went the other way: one model at 35% malformed while three others across both providers sat at 0%, so tripping the provider would have taken out three healthy models. Encoding faults are provider-shaped because they come from transport; content faults are model-shaped because they come from the model. Whichever signature fires decides the key. - **Threshold: novelty or rate, never mere presence.** One mangled response is noise. The reactive rejection detector in `metrics.py` already had to solve this exact problem and its answer applies unchanged: a signature not seen in the 24h baseline warns at n>=2; a familiar one warns at the configured count. A single fixed count cannot both catch a new incident and stay quiet on a known one. Recovery is automatic on the next clean response, as with 5xx. ### 5. Surfacing - `/metrics`: a `mangled_output_warnings` series alongside the existing detectors, per provider, over the same 24h window. - Admin: the verdict shows up in the proficiency page's client-outcomes table without any new work, since it renders `verdict` already. ## Measured, on the live providers Both probed directly; the difference is the whole story. | provider | SSE `Content-Type` | `r.encoding` | corrupted? | |---|---|---|---| | neuralwatt | `text/event-stream; charset=utf-8` | `utf-8` | no | | openrouter | `text/event-stream` | `ISO-8859-1` | **yes** | So this was **OpenRouter-only**, and never touched NeuralWatt. End to end through the router against a free OpenRouter model, same prompt, same model, only the dispatcher differing: ``` before 'cafeI\x81 naA-ve a\x80\x94 a\x80\x9cquoteda\x80\x9d a\x86\x92 done' after 'cafe(acute) naive (em dash) (open quote)quoted(close quote) (arrow) done' ``` (the second line is the exact prompt string round-tripping byte-perfect; rendered here as descriptions so this file stays ASCII.) Three consequences worth carrying into the build: 1. **The provider-bias argument is measured, not hypothetical.** A per-model score during this window would have marked down every OpenRouter row and left every NeuralWatt row untouched, entirely because of a header. That settles the scope question: trip on the provider. 2. **The corrupted output is VALID UTF-8.** Both runs decoded strictly without error -- the mangled one is well-formed, just wrong. A decode-failure check would catch none of this. Detection has to be the double-encoding signature (`c3 a2`, `c3 83`, `c3 82` lead pairs), not "does it parse". 3. **It correlates with when OpenRouter came back.** #45 reintroduced OpenRouter as an allowlist provider and traffic shifted onto it, which is when the mojibake became visible. Any pre-fix OpenRouter streamed traffic is unusable as a baseline; NeuralWatt traffic over the same period is clean. ## The empty-200 incident, 2026-09-08 Measured on the live router while `onlycheaps` was the default profile and opencode was hammering a free OpenRouter model. Structural verdicts over a 45-minute window: | model | malformed | unverifiable | total | % malformed | |---|---|---|---|---| | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` | **38** | 70 | 108 | **35.2%** | | `qwen3.6-35b` | 0 | 18 | 18 | 0% | | `deepseek-v4-flash` | 0 | 8 | 8 | 0% | | `z-ai/glm-5.3-flash` | 0 | 2 | 2 | 0% | All 38 were `malformed: empty response`. Every one arrived as **HTTP 200**. Nothing was logged as an upstream failure, because from the router's point of view nothing failed. Three things follow, and they are the reason this plan is worth building rather than filing. **1. The circuit breaker cannot see this at all.** It trips on `status_code >= 400` (both paths -- CLAUDE.md's "on a 5xx" is stale). A provider that degrades by returning an empty completion with a success status is invisible to it, so the router kept dispatching: 108 calls in 25 minutes, 38 of them empty, with no backoff and nothing to re-probe. This is a strictly worse failure mode than the 429 case the breaker does handle, because there is no eventual recovery -- only the operator noticing. **2. The signal already exists and has nowhere to go.** `verify_response` caught every one of them, correctly, for free. Structural verdicts are diagnostics only (deliberately -- see CLAUDE.md), so the router detected the degradation 38 times and could do nothing with it. That gap is exactly this plan's subject; the mojibake case just happened to arrive first. **3. Attribution is far easier here than for encoding.** The encoding signal needed the input's charset profile to tell a model echoing non-ASCII from a model mangling it. This one needs no comparison at all: - an empty completion on a **200**, on a turn with **no tool calls**, is unambiguous -- `has_tool_calls` already short-circuits the legitimate case, which is what the 70 `unverifiable` rows above are; - the contrast is measured rather than assumed: 35.2% against 0%, 0%, 0% from three other models over the same window and the same traffic. So the empty-200 trigger should ship **first**, and the encoding trigger second. It is cheaper to detect, unambiguous to attribute, and the incident it prevents is ongoing rather than historical. ### What this changes in the design above - **Detection** gains a signature that is not about bytes: `verdict == "malformed"` with `detail == "empty response"`, which `verification.py` already produces. No new checker. - **Attribution** for this signature is `has_tool_calls == False` plus a 2xx status. Both are already in hand at the call site. - **The trip threshold** wants a *rate*, not a count -- one empty response is noise, 35% over 100 calls is a provider falling over. The rejection detector's novelty-or-rate rule still applies, but here the rate arm is the one that fires. - **Scope stays provider-and-model.** Unlike the mojibake case, this was one model degrading while three others on the same two providers were clean, so tripping the whole provider would have been wrong. Record the provider, trip the model. That is the opposite conclusion from the encoding case, and the reason is measurable in the table above -- which is why the signature, not the plan, should decide the scope. ## Open, and worth settling before writing code 1. **Is there any pre-fix traffic worth mining?** No -- no completions are stored, so this detector starts from zero and its first real data arrives after `8518114` is deployed. Stated plainly rather than discovered later. 2. **Should a trip be visible to the client?** A 5xx breaker skip is invisible because another model serves the request. Same here, but the decision row should record why the provider was skipped. 3. **Is a header check worth having as its own guard?** A provider that omits charset is a standing hazard for anything else that consumes its stream. Cheap to assert at poll time; unclear whether it earns a warning of its own or just a note in the provider row. ## Tests The pattern `tests/test_task_set.py` sets -- prove the harness before trusting the measurement: - each detection signature, positive and negative, against fixed byte strings including the exact `c3 a2 c2 80 c2 94` sequence; - the attribution rule both ways: non-ASCII prompt -> not attributable even with `U+FFFD` present; - a legitimate non-ASCII answer (Japanese, accented prose) to a non-ASCII prompt produces no verdict at all; - the breaker trips on provider scope and clears on the next clean response; - `feedback.py` still folds nothing from a `mangled` row.