Files
6krrt/plans/mangled-output-detection.md
adlee-was-taken b11f4133fd feat(admin): set the routing default from the profiles page; fold the allowlist
Three UI changes and two plan corrections.

Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.

Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.

Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.

Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 20:49:36 -04:00

274 lines
13 KiB
Markdown

# Degraded output as a routing signal
Status: planned -- not built; prerequisite landed in `8518114`, second trigger measured 2026-09-08
Two measured instances of one thing: **a provider returning garbage the
router can detect for free, with no way to act on it.** Mojibake was the
first and turned out to be ours. Empty-response-on-200 is the second and is
not -- see "The empty-200 incident" below, which is the stronger case for
building this and changes what the trigger has to watch.
## Why this exists, and the thing it must not become
The trigger was mojibake in agent output. Investigating it found the cause
in the router, not the models: the SSE proxy decoded upstream bytes as
latin-1 (requests falls back to ISO-8859-1 for any `text/*` without a
charset, and OpenRouter's `text/event-stream` omits one) and re-encoded them
as UTF-8, double-encoding every non-ASCII character in every streamed
completion. `8518114` fixes that.
**Read that again before building anything here.** Had a per-model
"mangled output" score existed during that period, it would have recorded a
failure against whichever model happened to be routed, and because the bug
was triggered by one provider's Content-Type header it would have
systematically scored every OpenRouter row worse than every NeuralWatt row,
for a reason with nothing to do with the models. CLAUDE.md already lists
four cases of "scored the rig rather than the model"; this would have been
the fifth, and the first to bias a whole provider.
So the governing constraint is not detection quality. It is attribution:
**a mangled-output signal is only worth having if it can tell the model's
fault from the pipeline's.** Everything below is shaped by that.
## What it is not
- **Not a proficiency term.** Three reasons, each sufficient:
- *Wrong grain.* `proficiency` is per `(model, category)`. Encoding
breakage is not category-dependent, so the signal would be diluted
across 11 buckets and need roughly 11x the samples to become visible.
- *It reverses a decision already made.* Structural verdicts were
deliberately demoted to diagnostics because "a structural failure is not
a verified client outcome". An automatic checker feeding the score puts
that straight back.
- *Wrong instrument.* `proficiency` is a slow mean behind a
20-pseudo-observation prior. Encoding breakage is a step change -- a
deploy, a proxy, a quantisation swap, or a router bug. A slow mean is
late to notice it and slow to forgive it.
- **Not a new objective or weight.** Ranking stays quality-first with cost
as the tiebreak inside `quality_tolerance`. There is no knob to tune here.
- **Not a reason to store completions.** This router stores no raw task
text, and that does not change. Detection is inline; only a verdict is
persisted.
## What it is
The same shape as `circuit_breaker.py`: a **passive availability skip**, not
a score. Mangled output is a health signal, usually transient and usually
provider-side, so the right response is to route away now and re-probe soon
-- exactly what the breaker already does for 5xx, cooldown and backoff
included, clearing on the next clean response.
### 1. Detection (free, structural, inline)
A new `mangled` verdict in `verification.py`, checked only when the response
is otherwise checkable. Three signatures, in descending confidence:
| signature | example | confidence |
|---|---|---|
| double-encoded UTF-8 | `c3 a2 c2 80 c2 94` (em dash); any `c3 83`/`c3 82` lead pair | high |
| replacement character | `U+FFFD` | high |
| non-ASCII absent from the input | smart quotes injected into ASCII source | low -- diagnostic only |
Only the first two may trip anything. The third is recorded and surfaced but
must never gate routing on its own: a model correctly answering in Japanese,
or correctly quoting a non-ASCII identifier, produces it.
### 2. Attribution (the load-bearing part)
`verify_response(text, finish_reason, has_tool_calls)` has no view of the
prompt today, so this needs one new argument -- the input's charset profile,
computed once per request. The rule:
- prompt was pure ASCII and the completion carries `U+FFFD` or a
double-encoded sequence -> **attributable**
- prompt already carried non-ASCII -> **not attributable**, because a model
echoing what it was given is behaving correctly
This is the same lesson as `has_tool_calls` and `client_capped`: the checker
is only safe once it knows what is not the model's fault. Both of those were
learned by over-applying a verdict, once in each direction.
Non-attributable rows are still written, with `model_attributable = 0` --
the treatment degraded-classification outcomes already get. The record
survives; it just does not steer.
### 3. Storage (no migration)
`verifications` already carries `kind`, `verdict`, `model_attributable` and
`applied_at`. A `mangled` value in the existing `verdict` column and a
`kind='structural'` row is the whole change. `Verdict` in `verification.py`
gains one literal.
`feedback.py` is untouched and must stay untouched: its existing
`AND model_attributable = 1` plus the client-outcome filter already keep
this out of `proficiency`.
### 4. Routing response
Trip the breaker. Two decisions inside that:
- **Scope: provider, not model.** The incident that motivated this was
uniform across every OpenRouter model because the fault was one layer up.
A per-model breaker would have tripped them one at a time, taking eleven
separate incidents to describe one. Record the `model_id` on the row for
forensics; key the trip on `provider`.
`circuit_breaker._store` is keyed `(model_id, provider)` and lives in
memory, so this needs a provider-level entry or a reserved key -- a small
addition, not a redesign.
**Scope is per-signature, not global to this feature.** The empty-200
incident below went the other way: one model at 35% malformed while three
others across both providers sat at 0%, so tripping the provider would
have taken out three healthy models. Encoding faults are provider-shaped
because they come from transport; content faults are model-shaped because
they come from the model. Whichever signature fires decides the key.
- **Threshold: novelty or rate, never mere presence.** One mangled response
is noise. The reactive rejection detector in `metrics.py` already had to
solve this exact problem and its answer applies unchanged: a signature not
seen in the 24h baseline warns at n>=2; a familiar one warns at the
configured count. A single fixed count cannot both catch a new incident
and stay quiet on a known one.
Recovery is automatic on the next clean response, as with 5xx.
### 5. Surfacing
- `/metrics`: a `mangled_output_warnings` series alongside the existing
detectors, per provider, over the same 24h window.
- Admin: the verdict shows up in the proficiency page's client-outcomes
table without any new work, since it renders `verdict` already.
## Measured, on the live providers
Both probed directly; the difference is the whole story.
| provider | SSE `Content-Type` | `r.encoding` | corrupted? |
|---|---|---|---|
| neuralwatt | `text/event-stream; charset=utf-8` | `utf-8` | no |
| openrouter | `text/event-stream` | `ISO-8859-1` | **yes** |
So this was **OpenRouter-only**, and never touched NeuralWatt. End to end
through the router against a free OpenRouter model, same prompt, same
model, only the dispatcher differing:
```
before 'cafeI\x81 naA-ve a\x80\x94 a\x80\x9cquoteda\x80\x9d a\x86\x92 done'
after 'cafe(acute) naive (em dash) (open quote)quoted(close quote) (arrow) done'
```
(the second line is the exact prompt string round-tripping byte-perfect;
rendered here as descriptions so this file stays ASCII.)
Three consequences worth carrying into the build:
1. **The provider-bias argument is measured, not hypothetical.** A per-model
score during this window would have marked down every OpenRouter row and
left every NeuralWatt row untouched, entirely because of a header. That
settles the scope question: trip on the provider.
2. **The corrupted output is VALID UTF-8.** Both runs decoded strictly
without error -- the mangled one is well-formed, just wrong. A
decode-failure check would catch none of this. Detection has to be the
double-encoding signature (`c3 a2`, `c3 83`, `c3 82` lead pairs), not
"does it parse".
3. **It correlates with when OpenRouter came back.** #45 reintroduced
OpenRouter as an allowlist provider and traffic shifted onto it, which is
when the mojibake became visible. Any pre-fix OpenRouter streamed traffic
is unusable as a baseline; NeuralWatt traffic over the same period is
clean.
## The empty-200 incident, 2026-09-08
Measured on the live router while `onlycheaps` was the default profile and
opencode was hammering a free OpenRouter model. Structural verdicts over a
45-minute window:
| model | malformed | unverifiable | total | % malformed |
|---|---|---|---|---|
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` | **38** | 70 | 108 | **35.2%** |
| `qwen3.6-35b` | 0 | 18 | 18 | 0% |
| `deepseek-v4-flash` | 0 | 8 | 8 | 0% |
| `z-ai/glm-5.3-flash` | 0 | 2 | 2 | 0% |
All 38 were `malformed: empty response`. Every one arrived as **HTTP 200**.
Nothing was logged as an upstream failure, because from the router's point
of view nothing failed.
Three things follow, and they are the reason this plan is worth building
rather than filing.
**1. The circuit breaker cannot see this at all.** It trips on
`status_code >= 400` (both paths -- CLAUDE.md's "on a 5xx" is stale). A
provider that degrades by returning an empty completion with a success
status is invisible to it, so the router kept dispatching: 108 calls in 25
minutes, 38 of them empty, with no backoff and nothing to re-probe. This is
a strictly worse failure mode than the 429 case the breaker does handle,
because there is no eventual recovery -- only the operator noticing.
**2. The signal already exists and has nowhere to go.**
`verify_response` caught every one of them, correctly, for free. Structural
verdicts are diagnostics only (deliberately -- see CLAUDE.md), so the router
detected the degradation 38 times and could do nothing with it. That gap is
exactly this plan's subject; the mojibake case just happened to arrive
first.
**3. Attribution is far easier here than for encoding.** The encoding
signal needed the input's charset profile to tell a model echoing non-ASCII
from a model mangling it. This one needs no comparison at all:
- an empty completion on a **200**, on a turn with **no tool calls**, is
unambiguous -- `has_tool_calls` already short-circuits the legitimate
case, which is what the 70 `unverifiable` rows above are;
- the contrast is measured rather than assumed: 35.2% against 0%, 0%, 0%
from three other models over the same window and the same traffic.
So the empty-200 trigger should ship **first**, and the encoding trigger
second. It is cheaper to detect, unambiguous to attribute, and the incident
it prevents is ongoing rather than historical.
### What this changes in the design above
- **Detection** gains a signature that is not about bytes:
`verdict == "malformed"` with `detail == "empty response"`, which
`verification.py` already produces. No new checker.
- **Attribution** for this signature is `has_tool_calls == False` plus a
2xx status. Both are already in hand at the call site.
- **The trip threshold** wants a *rate*, not a count -- one empty response
is noise, 35% over 100 calls is a provider falling over. The rejection
detector's novelty-or-rate rule still applies, but here the rate arm is
the one that fires.
- **Scope stays provider-and-model.** Unlike the mojibake case, this was one
model degrading while three others on the same two providers were clean,
so tripping the whole provider would have been wrong. Record the provider,
trip the model. That is the opposite conclusion from the encoding case,
and the reason is measurable in the table above -- which is why the
signature, not the plan, should decide the scope.
## Open, and worth settling before writing code
1. **Is there any pre-fix traffic worth mining?** No -- no completions are
stored, so this detector starts from zero and its first real data
arrives after `8518114` is deployed. Stated plainly rather than
discovered later.
2. **Should a trip be visible to the client?** A 5xx breaker skip is
invisible because another model serves the request. Same here, but the
decision row should record why the provider was skipped.
3. **Is a header check worth having as its own guard?** A provider that
omits charset is a standing hazard for anything else that consumes its
stream. Cheap to assert at poll time; unclear whether it earns a warning
of its own or just a note in the provider row.
## Tests
The pattern `tests/test_task_set.py` sets -- prove the harness before
trusting the measurement:
- each detection signature, positive and negative, against fixed byte
strings including the exact `c3 a2 c2 80 c2 94` sequence;
- the attribution rule both ways: non-ASCII prompt -> not attributable even
with `U+FFFD` present;
- a legitimate non-ASCII answer (Japanese, accented prose) to a non-ASCII
prompt produces no verdict at all;
- the breaker trips on provider scope and clears on the next clean
response;
- `feedback.py` still folds nothing from a `mangled` row.