conversation-identity: the router learns which conversation each request belongs to #99

Merged
alee merged 18 commits from feat/conversation-identity into main 2026-09-22 22:39:07 +00:00
Owner

conversation-identity — the router learns which conversation each request belongs to

This plan makes the router learn which conversation every request belongs to, from the client, instead of guessing from a hash of the system prompt. An opencode plugin (router-link.js) stamps every request with the opencode session id, the agent name, and the parent session id; the router uses the session id as its session_key, records the agent/parent on route_decisions, and lets POST /outcome attribute a test result to an exact conversation.

Why: on 2026-09-18 one session_key carried a ~100k-context main conversation and dozens of short sub-agent conversations at once. Everything session-scoped read a blend: the Wave 2 incumbent lookup, the classifier's session_history step, the cache-rate and switch measurement. This plan is the foundation that un-blends them.


The Contract (single home — see docs/api.md for the canonical copy)

Headers, all optional, sent by the client to POST /v1/chat/completions:

header meaning validation
X-Router-Conversation id of ONE conversation (opencode sessionID; each sub-agent session has its own) 1-128 chars, ^[A-Za-z0-9._:-]+$
X-Router-Agent agent name that issued the request 1-64 chars, same charset
X-Router-Parent conversation id of the session that spawned this one, best effort 1-128 chars, same charset

[H2] Duplicate-header semantics: starlette Headers is a multi-dict — .get() returns the first value for a name and iteration yields lowercased names. resolve_identity therefore applies first-wins: for a duplicated header name, the first occurrence is used and later ones ignored, whether the input is a real Headers object or a plain dict. The lowercased-key normalization makes first-wins well-defined across both input shapes.

Rules:

  1. An absent or invalid header value is treated as absent. It is never an error. Why: a bad header must not fail a paid request.
  2. session_key = "c:" + <conversation id> when X-Router-Conversation is valid, otherwise session_fingerprint(messages) exactly as today. Why the prefix: a fingerprint is 16 hex characters with no colon, so the two namespaces cannot collide.
  3. parent_key = "c:" + <parent id> when X-Router-Parent is valid, otherwise NULL. agent = the header value, otherwise NULL. Both stored on route_decisions only.
  4. POST /outcome accepts an optional body field conversation_id with the conversation-id validation above. Resolution order in _find_outcome_row: request_id, then conversation_id, then source, then the unambiguous-most-recent fallback. The conversation_id step matches session_key = 'c:' || <conversation id> first in energy_observations, then in local_energy_observations.session_key, so local-dispatch answers are attributable too. When conversation_id is present and matches no row, the result is "not found" (404) — it never falls through to source or the fallback.
  5. Ids and agent names are opaque labels. They never appear alongside message text in any log line or column.

Decision recorded: this plan adds no config knob (header presence is the opt-in; removing the plugin is the off switch, no router restart). The adoption counter is the deliberate exception — its window is config-driven because a counter is observation (read-only), not behavior.

Accepted risk: the router has no auth, so any local client can name any conversation id. Loopback bind is the existing boundary.


Manual acceptance (after merge and restart)

  1. systemctl --user restart llm-router.service (Python changes need it).
  2. Install the plugin: command cp -f deploy/opencode-plugin/router-link.js ~/.config/opencode/plugins/ && rm ~/.config/opencode/plugins/router-outcome.js, then restart opencode.
  3. Run an Atlas plan (or any prompt that spawns two sub-agents), then open the live DB read-only and run:
    SELECT session_key, agent, COUNT(*) FROM route_decisions WHERE observed_at > <start> GROUP BY 1,2.
    Pass = several distinct c:... keys for one run, each with its own agent, and parent_key set on the sub-agent rows.
  4. Run one test command in opencode. Pass = exactly ONE new client_outcome row for that conversation and no 409 in the router log.
  5. If no c: keys appear, the chat.headers hook is not reaching the openai-compatible provider. The router is unharmed (it fell back to fingerprints); report it, because the plugin then needs a different transport for the headers.
  6. Adoption: hit /metrics (or the admin dashboard) and confirm the c: conversation counter reports a non-zero value after step 3.

CLAUDE.md lines proposed (NOT applied — this plan must not edit CLAUDE.md)

Under the Module map / Dispatcher table, add:

| `identity.py` | Pure conversation-identity resolution from client headers (X-Router-Conversation/Agent/Parent), first-wins on duplicates | Imported by dispatcher only; reads request headers, never logs message text |

Under Other entrypoints, update the plugin note (replaces router-outcome):

| `deploy/opencode-plugin/router-link.js` | opencode plugin stamped on the router: X-Router-Conversation/Agent/Parent on chat.headers, conversation_id on tool.execute.after | Replaces the old router-outcome.js; install one file, not two |

And a one-line contract pointer near the module map:

- Conversation headers (X-Router-Conversation / X-Router-Agent / X-Router-Parent) and `POST /outcome` conversation_id: single home is `docs/api.md` (see the Contract).

(These are suggestions for the maintainer to apply when merging — they are intentionally left out of this PR's diff.)


Follow-on plans (each waits on this plan landing)

  • Per-conversation model pinning and escalation on a failed outcome (Wave 5.1, "route once").
  • Reviewer diversity: exclude the implementer's model family for review agents, using parent_key.
  • Per-agent cost and cache series in /metrics, to measure what the 13 OMO agents' system prompts cost (the adoption counter is the seed of this).
  • Receipts: per-todo model, cost and cache rate appended to the plan's notepad.

Adversarial consensus (resolved in the debate; do NOT re-open)

  • 404 over 200-with-unattributed: when conversation_id is present and matches nothing, return 404. Matches codebase doctrine (report_outcome) and is the load-bearing guarantee against silent misattribution.
  • Do NOT split outcome-attribution from routing: both halves need the same foundation (identity module + schema) and share the same cold-start risk; splitting adds merge cycles without reducing scope.
  • Do NOT cut parent_key: zero-cost nullable column with a real roadmap consumer (reviewer diversity).
  • Do NOT "fix" the session-dir heuristic: a category error (filesystem path ≠ conversation topology); out of scope and not made worse by this plan.
  • Contract rule 5 vs TUI surfacing: decision_row carries no message text and tui_model surfaces only opaque labels, so it does not violate the no-message-text rule.
  • Plugin dependency does not gate the router release: both upgrade orders are safe — old router + new plugin (OutcomeReport extra='ignore' drops unknown conversation_id), or new router + old plugin (byte-identical fallback). Either order merges independently.

Evidence summary

  • Full pytest suite (worktree, branch feat/conversation-identity): 2241 passed, 1 failed — baseline was 2186 passed, 1 failed, so +55 new tests, zero new failures, zero new skips.
  • The single failure is the pre-existing test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity date-rot/environmental test that also fails on clean origin/main (seeds 2026-09-13 data outside the trailing cache-rate window on the 2026-09-20 clock). Documented in .omo/evidence/conversation-identity/baseline.txt; not caused by — and not fixed by — this plan.
  • Node plugin suite: node --test deploy/opencode-plugin/ → 16 tests, 16 pass, 0 fail.
  • Full evidence (counts, git status, [H8] shape-check, [H4] cold-start note): .omo/evidence/conversation-identity/final.txt.
# conversation-identity — the router learns which conversation each request belongs to This plan makes the router learn which conversation every request belongs to, **from the client**, instead of guessing from a hash of the system prompt. An opencode plugin (`router-link.js`) stamps every request with the opencode session id, the agent name, and the parent session id; the router uses the session id as its `session_key`, records the agent/parent on `route_decisions`, and lets `POST /outcome` attribute a test result to an exact conversation. **Why:** on 2026-09-18 one `session_key` carried a ~100k-context main conversation and dozens of short sub-agent conversations at once. Everything session-scoped read a blend: the Wave 2 incumbent lookup, the classifier's `session_history` step, the cache-rate and switch measurement. This plan is the foundation that un-blends them. --- ## The Contract (single home — see docs/api.md for the canonical copy) **Headers, all optional, sent by the client to `POST /v1/chat/completions`:** | header | meaning | validation | |---|---|---| | `X-Router-Conversation` | id of ONE conversation (opencode `sessionID`; each sub-agent session has its own) | 1-128 chars, `^[A-Za-z0-9._:-]+$` | | `X-Router-Agent` | agent name that issued the request | 1-64 chars, same charset | | `X-Router-Parent` | conversation id of the session that spawned this one, best effort | 1-128 chars, same charset | **[H2] Duplicate-header semantics:** starlette `Headers` is a multi-dict — `.get()` returns the **first** value for a name and iteration yields lowercased names. `resolve_identity` therefore applies **first-wins**: for a duplicated header name, the first occurrence is used and later ones ignored, whether the input is a real `Headers` object or a plain `dict`. The lowercased-key normalization makes first-wins well-defined across both input shapes. **Rules:** 1. An absent or invalid header value is treated as absent. It is never an error. Why: a bad header must not fail a paid request. 2. `session_key` = `"c:" + <conversation id>` when `X-Router-Conversation` is valid, otherwise `session_fingerprint(messages)` exactly as today. Why the prefix: a fingerprint is 16 hex characters with no colon, so the two namespaces cannot collide. 3. `parent_key` = `"c:" + <parent id>` when `X-Router-Parent` is valid, otherwise NULL. `agent` = the header value, otherwise NULL. Both stored on `route_decisions` only. 4. `POST /outcome` accepts an optional body field `conversation_id` with the conversation-id validation above. Resolution order in `_find_outcome_row`: `request_id`, then `conversation_id`, then `source`, then the unambiguous-most-recent fallback. The `conversation_id` step matches `session_key = 'c:' || <conversation id>` first in `energy_observations`, then in `local_energy_observations.session_key`, so local-dispatch answers are attributable too. When `conversation_id` is present and matches no row, the result is "not found" (404) — it never falls through to `source` or the fallback. 5. Ids and agent names are opaque labels. They never appear alongside message text in any log line or column. **Decision recorded:** this plan adds **no config knob** (header presence is the opt-in; removing the plugin is the off switch, no router restart). The adoption counter is the deliberate exception — its window is config-driven because a counter is observation (read-only), not behavior. **Accepted risk:** the router has no auth, so any local client can name any conversation id. Loopback bind is the existing boundary. --- ## Manual acceptance (after merge and restart) 1. `systemctl --user restart llm-router.service` (Python changes need it). 2. Install the plugin: `command cp -f deploy/opencode-plugin/router-link.js ~/.config/opencode/plugins/ && rm ~/.config/opencode/plugins/router-outcome.js`, then restart opencode. 3. Run an Atlas plan (or any prompt that spawns two sub-agents), then open the live DB read-only and run: `SELECT session_key, agent, COUNT(*) FROM route_decisions WHERE observed_at > <start> GROUP BY 1,2`. Pass = several distinct `c:...` keys for one run, each with its own agent, and `parent_key` set on the sub-agent rows. 4. Run one test command in opencode. Pass = exactly ONE new `client_outcome` row for that conversation and no 409 in the router log. 5. If no `c:` keys appear, the `chat.headers` hook is not reaching the openai-compatible provider. The router is unharmed (it fell back to fingerprints); report it, because the plugin then needs a different transport for the headers. 6. Adoption: hit `/metrics` (or the admin dashboard) and confirm the `c:` conversation counter reports a non-zero value after step 3. --- ## CLAUDE.md lines proposed (NOT applied — this plan must not edit CLAUDE.md) Under the **Module map / Dispatcher** table, add: ``` | `identity.py` | Pure conversation-identity resolution from client headers (X-Router-Conversation/Agent/Parent), first-wins on duplicates | Imported by dispatcher only; reads request headers, never logs message text | ``` Under **Other entrypoints**, update the plugin note (replaces router-outcome): ``` | `deploy/opencode-plugin/router-link.js` | opencode plugin stamped on the router: X-Router-Conversation/Agent/Parent on chat.headers, conversation_id on tool.execute.after | Replaces the old router-outcome.js; install one file, not two | ``` And a one-line contract pointer near the module map: ``` - Conversation headers (X-Router-Conversation / X-Router-Agent / X-Router-Parent) and `POST /outcome` conversation_id: single home is `docs/api.md` (see the Contract). ``` (These are suggestions for the maintainer to apply when merging — they are intentionally left out of this PR's diff.) --- ## Follow-on plans (each waits on this plan landing) - Per-conversation model pinning and escalation on a failed outcome (Wave 5.1, "route once"). - Reviewer diversity: exclude the implementer's model family for review agents, using `parent_key`. - Per-agent cost and cache series in `/metrics`, to measure what the 13 OMO agents' system prompts cost (the adoption counter is the seed of this). - Receipts: per-todo model, cost and cache rate appended to the plan's notepad. --- ## Adversarial consensus (resolved in the debate; do NOT re-open) - **404 over 200-with-unattributed**: when `conversation_id` is present and matches nothing, return 404. Matches codebase doctrine (`report_outcome`) and is the load-bearing guarantee against silent misattribution. - **Do NOT split outcome-attribution from routing**: both halves need the same foundation (identity module + schema) and share the same cold-start risk; splitting adds merge cycles without reducing scope. - **Do NOT cut `parent_key`**: zero-cost nullable column with a real roadmap consumer (reviewer diversity). - **Do NOT "fix" the session-dir heuristic**: a category error (filesystem path ≠ conversation topology); out of scope and not made worse by this plan. - **Contract rule 5 vs TUI surfacing**: `decision_row` carries no message text and `tui_model` surfaces only opaque labels, so it does not violate the no-message-text rule. - **Plugin dependency does not gate the router release**: both upgrade orders are safe — old router + new plugin (OutcomeReport `extra='ignore'` drops unknown `conversation_id`), or new router + old plugin (byte-identical fallback). Either order merges independently. --- ## Evidence summary - Full pytest suite (worktree, branch `feat/conversation-identity`): **2241 passed, 1 failed** — baseline was 2186 passed, 1 failed, so **+55 new tests, zero new failures, zero new skips**. - The single failure is the **pre-existing** `test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity` date-rot/environmental test that also fails on clean origin/main (seeds 2026-09-13 data outside the trailing cache-rate window on the 2026-09-20 clock). Documented in `.omo/evidence/conversation-identity/baseline.txt`; not caused by — and not fixed by — this plan. - Node plugin suite: `node --test deploy/opencode-plugin/` → **16 tests, 16 pass, 0 fail**. - Full evidence (counts, git status, [H8] shape-check, [H4] cold-start note): `.omo/evidence/conversation-identity/final.txt`.
alee added 7 commits 2026-09-21 00:32:23 +00:00
alee added 9 commits 2026-09-21 01:55:22 +00:00
alee added 2 commits 2026-09-22 22:18:54 +00:00
alee merged commit fe8c8134b6 into main 2026-09-22 22:39:07 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#99