Files
6krrt/plans/mangled-output-detection.md
adlee-was-taken b11f4133fd feat(admin): set the routing default from the profiles page; fold the allowlist
Three UI changes and two plan corrections.

Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.

Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.

Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.

Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 20:49:36 -04:00

13 KiB

Degraded output as a routing signal

Status: planned -- not built; prerequisite landed in 8518114, second trigger measured 2026-09-08

Two measured instances of one thing: a provider returning garbage the router can detect for free, with no way to act on it. Mojibake was the first and turned out to be ours. Empty-response-on-200 is the second and is not -- see "The empty-200 incident" below, which is the stronger case for building this and changes what the trigger has to watch.

Why this exists, and the thing it must not become

The trigger was mojibake in agent output. Investigating it found the cause in the router, not the models: the SSE proxy decoded upstream bytes as latin-1 (requests falls back to ISO-8859-1 for any text/* without a charset, and OpenRouter's text/event-stream omits one) and re-encoded them as UTF-8, double-encoding every non-ASCII character in every streamed completion. 8518114 fixes that.

Read that again before building anything here. Had a per-model "mangled output" score existed during that period, it would have recorded a failure against whichever model happened to be routed, and because the bug was triggered by one provider's Content-Type header it would have systematically scored every OpenRouter row worse than every NeuralWatt row, for a reason with nothing to do with the models. CLAUDE.md already lists four cases of "scored the rig rather than the model"; this would have been the fifth, and the first to bias a whole provider.

So the governing constraint is not detection quality. It is attribution: a mangled-output signal is only worth having if it can tell the model's fault from the pipeline's. Everything below is shaped by that.

What it is not

  • Not a proficiency term. Three reasons, each sufficient:
    • Wrong grain. proficiency is per (model, category). Encoding breakage is not category-dependent, so the signal would be diluted across 11 buckets and need roughly 11x the samples to become visible.
    • It reverses a decision already made. Structural verdicts were deliberately demoted to diagnostics because "a structural failure is not a verified client outcome". An automatic checker feeding the score puts that straight back.
    • Wrong instrument. proficiency is a slow mean behind a 20-pseudo-observation prior. Encoding breakage is a step change -- a deploy, a proxy, a quantisation swap, or a router bug. A slow mean is late to notice it and slow to forgive it.
  • Not a new objective or weight. Ranking stays quality-first with cost as the tiebreak inside quality_tolerance. There is no knob to tune here.
  • Not a reason to store completions. This router stores no raw task text, and that does not change. Detection is inline; only a verdict is persisted.

What it is

The same shape as circuit_breaker.py: a passive availability skip, not a score. Mangled output is a health signal, usually transient and usually provider-side, so the right response is to route away now and re-probe soon -- exactly what the breaker already does for 5xx, cooldown and backoff included, clearing on the next clean response.

1. Detection (free, structural, inline)

A new mangled verdict in verification.py, checked only when the response is otherwise checkable. Three signatures, in descending confidence:

signature example confidence
double-encoded UTF-8 c3 a2 c2 80 c2 94 (em dash); any c3 83/c3 82 lead pair high
replacement character U+FFFD high
non-ASCII absent from the input smart quotes injected into ASCII source low -- diagnostic only

Only the first two may trip anything. The third is recorded and surfaced but must never gate routing on its own: a model correctly answering in Japanese, or correctly quoting a non-ASCII identifier, produces it.

2. Attribution (the load-bearing part)

verify_response(text, finish_reason, has_tool_calls) has no view of the prompt today, so this needs one new argument -- the input's charset profile, computed once per request. The rule:

  • prompt was pure ASCII and the completion carries U+FFFD or a double-encoded sequence -> attributable
  • prompt already carried non-ASCII -> not attributable, because a model echoing what it was given is behaving correctly

This is the same lesson as has_tool_calls and client_capped: the checker is only safe once it knows what is not the model's fault. Both of those were learned by over-applying a verdict, once in each direction.

Non-attributable rows are still written, with model_attributable = 0 -- the treatment degraded-classification outcomes already get. The record survives; it just does not steer.

3. Storage (no migration)

verifications already carries kind, verdict, model_attributable and applied_at. A mangled value in the existing verdict column and a kind='structural' row is the whole change. Verdict in verification.py gains one literal.

feedback.py is untouched and must stay untouched: its existing AND model_attributable = 1 plus the client-outcome filter already keep this out of proficiency.

4. Routing response

Trip the breaker. Two decisions inside that:

  • Scope: provider, not model. The incident that motivated this was uniform across every OpenRouter model because the fault was one layer up. A per-model breaker would have tripped them one at a time, taking eleven separate incidents to describe one. Record the model_id on the row for forensics; key the trip on provider. circuit_breaker._store is keyed (model_id, provider) and lives in memory, so this needs a provider-level entry or a reserved key -- a small addition, not a redesign.

    Scope is per-signature, not global to this feature. The empty-200 incident below went the other way: one model at 35% malformed while three others across both providers sat at 0%, so tripping the provider would have taken out three healthy models. Encoding faults are provider-shaped because they come from transport; content faults are model-shaped because they come from the model. Whichever signature fires decides the key.

  • Threshold: novelty or rate, never mere presence. One mangled response is noise. The reactive rejection detector in metrics.py already had to solve this exact problem and its answer applies unchanged: a signature not seen in the 24h baseline warns at n>=2; a familiar one warns at the configured count. A single fixed count cannot both catch a new incident and stay quiet on a known one.

Recovery is automatic on the next clean response, as with 5xx.

5. Surfacing

  • /metrics: a mangled_output_warnings series alongside the existing detectors, per provider, over the same 24h window.
  • Admin: the verdict shows up in the proficiency page's client-outcomes table without any new work, since it renders verdict already.

Measured, on the live providers

Both probed directly; the difference is the whole story.

provider SSE Content-Type r.encoding corrupted?
neuralwatt text/event-stream; charset=utf-8 utf-8 no
openrouter text/event-stream ISO-8859-1 yes

So this was OpenRouter-only, and never touched NeuralWatt. End to end through the router against a free OpenRouter model, same prompt, same model, only the dispatcher differing:

before   'cafeI\x81 naA-ve a\x80\x94 a\x80\x9cquoteda\x80\x9d a\x86\x92 done'
after    'cafe(acute) naive (em dash) (open quote)quoted(close quote) (arrow) done'

(the second line is the exact prompt string round-tripping byte-perfect; rendered here as descriptions so this file stays ASCII.)

Three consequences worth carrying into the build:

  1. The provider-bias argument is measured, not hypothetical. A per-model score during this window would have marked down every OpenRouter row and left every NeuralWatt row untouched, entirely because of a header. That settles the scope question: trip on the provider.
  2. The corrupted output is VALID UTF-8. Both runs decoded strictly without error -- the mangled one is well-formed, just wrong. A decode-failure check would catch none of this. Detection has to be the double-encoding signature (c3 a2, c3 83, c3 82 lead pairs), not "does it parse".
  3. It correlates with when OpenRouter came back. #45 reintroduced OpenRouter as an allowlist provider and traffic shifted onto it, which is when the mojibake became visible. Any pre-fix OpenRouter streamed traffic is unusable as a baseline; NeuralWatt traffic over the same period is clean.

The empty-200 incident, 2026-09-08

Measured on the live router while onlycheaps was the default profile and opencode was hammering a free OpenRouter model. Structural verdicts over a 45-minute window:

model malformed unverifiable total % malformed
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 38 70 108 35.2%
qwen3.6-35b 0 18 18 0%
deepseek-v4-flash 0 8 8 0%
z-ai/glm-5.3-flash 0 2 2 0%

All 38 were malformed: empty response. Every one arrived as HTTP 200. Nothing was logged as an upstream failure, because from the router's point of view nothing failed.

Three things follow, and they are the reason this plan is worth building rather than filing.

1. The circuit breaker cannot see this at all. It trips on status_code >= 400 (both paths -- CLAUDE.md's "on a 5xx" is stale). A provider that degrades by returning an empty completion with a success status is invisible to it, so the router kept dispatching: 108 calls in 25 minutes, 38 of them empty, with no backoff and nothing to re-probe. This is a strictly worse failure mode than the 429 case the breaker does handle, because there is no eventual recovery -- only the operator noticing.

2. The signal already exists and has nowhere to go. verify_response caught every one of them, correctly, for free. Structural verdicts are diagnostics only (deliberately -- see CLAUDE.md), so the router detected the degradation 38 times and could do nothing with it. That gap is exactly this plan's subject; the mojibake case just happened to arrive first.

3. Attribution is far easier here than for encoding. The encoding signal needed the input's charset profile to tell a model echoing non-ASCII from a model mangling it. This one needs no comparison at all:

  • an empty completion on a 200, on a turn with no tool calls, is unambiguous -- has_tool_calls already short-circuits the legitimate case, which is what the 70 unverifiable rows above are;
  • the contrast is measured rather than assumed: 35.2% against 0%, 0%, 0% from three other models over the same window and the same traffic.

So the empty-200 trigger should ship first, and the encoding trigger second. It is cheaper to detect, unambiguous to attribute, and the incident it prevents is ongoing rather than historical.

What this changes in the design above

  • Detection gains a signature that is not about bytes: verdict == "malformed" with detail == "empty response", which verification.py already produces. No new checker.
  • Attribution for this signature is has_tool_calls == False plus a 2xx status. Both are already in hand at the call site.
  • The trip threshold wants a rate, not a count -- one empty response is noise, 35% over 100 calls is a provider falling over. The rejection detector's novelty-or-rate rule still applies, but here the rate arm is the one that fires.
  • Scope stays provider-and-model. Unlike the mojibake case, this was one model degrading while three others on the same two providers were clean, so tripping the whole provider would have been wrong. Record the provider, trip the model. That is the opposite conclusion from the encoding case, and the reason is measurable in the table above -- which is why the signature, not the plan, should decide the scope.

Open, and worth settling before writing code

  1. Is there any pre-fix traffic worth mining? No -- no completions are stored, so this detector starts from zero and its first real data arrives after 8518114 is deployed. Stated plainly rather than discovered later.
  2. Should a trip be visible to the client? A 5xx breaker skip is invisible because another model serves the request. Same here, but the decision row should record why the provider was skipped.
  3. Is a header check worth having as its own guard? A provider that omits charset is a standing hazard for anything else that consumes its stream. Cheap to assert at poll time; unclear whether it earns a warning of its own or just a note in the provider row.

Tests

The pattern tests/test_task_set.py sets -- prove the harness before trusting the measurement:

  • each detection signature, positive and negative, against fixed byte strings including the exact c3 a2 c2 80 c2 94 sequence;
  • the attribution rule both ways: non-ASCII prompt -> not attributable even with U+FFFD present;
  • a legitimate non-ASCII answer (Japanese, accented prose) to a non-ASCII prompt produces no verdict at all;
  • the breaker trips on provider scope and clears on the next clean response;
  • feedback.py still folds nothing from a mangled row.