The audit prompted by 8518114: every place the router encodes, decodes or
slices content, classified by what it can actually break.
plans/text-integrity-audit.md has the table. Two findings needed code.
Truncation could cut a grapheme cluster. Five head/tail slices -- the
classifier input clamp, the local-checker elision, and three sites in
context_prune -- sliced str directly. None could produce mojibake, because
Python indexes codepoints, but all of them could split a cluster: "cafe" +
U+0301 sliced at 4 drops the accent and leaves a bare combining mark at the
head of the tail. The same goes for ZWJ emoji sequences, variation
selectors and the regional-indicator pairs that make flags.
Severity is well below the SSE bug -- a stray mark, not a mangled document
-- but the SHAPE is the one this project keeps paying for: three of those
sites produce the prompt sent to the provider, so a bad cut is the router
corrupting the model's input and then reading the model's output as though
the model were solely responsible. src/textcut.py moves the cut to the
nearest boundary instead, shrinking rather than growing so a caller's
length stays a ceiling.
Reverting textcut fails 3 of the 5 new call-site tests, which is the point
of having them separate from the unit tests: a correct helper nobody calls
prevents nothing.
The six router-generated SSE writes are safe and now say so in the audit.
They look exactly like the bug that was just fixed and differ by one
keyword -- json.dumps defaults to ensure_ascii=True, so the payload is pure
ASCII before it is encoded. Anyone passing ensure_ascii=False there to save
bytes reintroduces a charset decision on an output path.
Every fixture in these files uses \u escapes rather than literal non-ASCII.
Files about text corruption should not silently change meaning if they ever
round-trip through something that mangles encodings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
5.1 KiB
5.1 KiB