From 91e8bba3c4ffe09680de231578c70858ce333595 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Thu, 3 Sep 2026 23:28:28 -0400 Subject: [PATCH 01/10] docs(plans): named routing profiles spec The mechanism is already half-built: auto:batch IS a profile -- a named variant of auto that changes one hard-filter parameter (latency_tolerance), shipped and documented for months. This generalises that hard-coded special case rather than inventing anything. Three call sites do most of the work already: the wants_routing membership test (dispatcher.py:2937) becomes a parse; the latency ternary (:2955) becomes one case of the general mechanism rather than surviving beside it; and routing.select_candidates already takes exclude_models, task_category, latency_tolerance and the rest. Exactly one primitive is missing. select_candidates has a denylist and no allowlist, while every profile the user named -- locality, onlycheaps, bigboybritches -- is naturally an allowlist. One `restrict_to` parameter threaded through rejection_reason -> is_eligible -> select_candidates covers it, with None meaning unrestricted and its own rejection string so a 422 names the profile that emptied the set. One decision is left open for the user and flagged inline: whether a profile is a candidate filter only, or may also override objective.*. Recommendation is filter-only -- a pure-filter `locality` already works because restricting to local rows leaves no cloud row to lose to, which is what the 0.767-vs-0.95 dormancy finding actually blocks. Also reconciles expand-local-llm-usage's stale checkboxes: todos 1-12 shipped via PR #22 and todo 13 was executed by hand, verified against six deliverables. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- plans/named-routing-profiles.md | 175 ++++++++++++++++++++++++++++++++ 1 file changed, 175 insertions(+) create mode 100644 plans/named-routing-profiles.md diff --git a/plans/named-routing-profiles.md b/plans/named-routing-profiles.md new file mode 100644 index 0000000..46f8c28 --- /dev/null +++ b/plans/named-routing-profiles.md @@ -0,0 +1,175 @@ +# Named routing profiles + +**Status: FINAL on architecture and scope. ONE open decision (§3) that is the +user's, flagged inline.** Written 2026-09-03 against `main` at `c2893b5`. + +## What this is + +Today the router has exactly one behaviour: classify, hard-filter, rank +quality-first with cost as a tiebreak. The user wants named alternatives — +`locality`, `onlycheaps`, `bigboybritches` — alongside a `default` that stays +the feature-rich automatic router it is now. + +**The mechanism already exists and is half-built.** `auto:batch` is a profile: +a named variant of `auto` that changes one hard-filter parameter +(`latency_tolerance`), documented in `README` and shipped for months. This plan +generalises that one hard-coded special case into a named, configurable set. +It is deliberately NOT a new ranking path. + +## 1. Where it slots in, precisely + +Three call sites, all already doing most of the work: + +- **`src/dispatcher.py:2937`** — + `wants_routing = bare in (ROUTER_MODEL, ROUTER_MODEL_BATCH)`. An exact + membership test against two constants. This becomes a *parse*: split + `auto:`, look the profile up, reject an unknown one. +- **`src/dispatcher.py:2955`** — + `latency = BATCH if requested == ROUTER_MODEL_BATCH else INTERACTIVE`. This + is the existing profile, expressed as a ternary. It must become one case of + the general mechanism, not survive alongside it. +- **`src/routing.py:313` `select_candidates`** — already takes + `exclude_models: set[str]`, `task_category`, `latency_tolerance`, + `allowed_access_levels`, `min_tool_proficiency`. A profile supplies or + overrides these. **All the filtering machinery is already there.** + +`routing.py` is a pure module with injected dependencies and +`rejection_reason` is the single copy of the rules (`is_eligible` is a thin +wrapper over it). Keep that: a profile must not introduce a second place where +eligibility is decided. + +## 2. The one genuinely missing primitive + +`select_candidates` has a **denylist** (`exclude_models`) and no **allowlist**. +Every profile the user named is naturally an allowlist: + +- `locality` → only `provider = 'ollama-local'` rows +- `onlycheaps` → only rows under a cost bar +- `bigboybritches` → only frontier rows + +Add **one** parameter, `restrict_to: set[str] | None = None`, threaded through +`rejection_reason` → `is_eligible` → `select_candidates`, with `None` meaning +"unrestricted" (NOT an empty set, which must mean "nothing eligible"). Give it +its own `rejection_reason` string so a 422 says *which* profile emptied the +candidate set — the catalog-staleness incident is the precedent: an empty +candidate set with an opaque reason cost ~19 hours to diagnose. + +Prefer expressing profiles as **predicates over catalog columns** (provider, +cost, tier) rather than hard-coded model-id lists, so a profile does not go +stale the moment the poller adds a row. A literal id list stays available for +the cases where that is genuinely what the user means. + +## 3. THE OPEN DECISION — filter only, or objective overrides too? + +**This is the user's call and must not be made silently.** + +A profile is unambiguously a candidate-set filter. The question is whether it +may ALSO override `objective.*` (`quality_tolerance`, +`max_energy_per_request`). + +It matters concretely, and we have the measurement: +`plans/local-dispatch-inert-and-test-coupling.md` established that the local +model scores 0.767 on `file_summarization` against cloud's 0.95 — a 0.183 gap +against a `quality_tolerance` of 0.1 — so local **never routes** under +quality-first ranking. + +- A **pure-filter** `locality` still works: restrict the set to local rows and + the best local model wins by default, because there is no cloud row left to + lose to. This covers the user's stated case. +- A **mixed** profile that *prefers* local without excluding cloud does NOT + work as a pure filter. It reproduces the dormancy exactly. + +So: pure filters satisfy the three named profiles. Objective overrides are only +needed for preference-shaped profiles nobody has asked for yet. + +**Recommendation: ship filter-only.** It is the smaller change, it satisfies +every named use case, and it keeps one ranking rule in the system. Add +overrides later if a real preference-shaped profile appears. If the user wants +overrides now, they must be scoped strictly to the profile and never mutate +global config — and `quality_tolerance` in particular describes measurement +noise on 2-3 samples, so overriding it means asserting a preference through a +knob that means something else. + +## 4. Naming and selection + +Profiles are selected the way `auto:batch` already is: the model field. + +``` +auto -> default profile +auto:batch -> MUST keep working, byte-identical behaviour +auto:locality -> named profile + -> passthrough, unchanged +``` + +`auto:batch` is load-bearing: it is in `README`, `docs/`, and possibly user +clients. **Reimplement it as a profile whose definition sets +`latency_tolerance: batch`, and pin its behaviour with a test that would fail +if the general mechanism changed it.** Do not leave the ternary in place beside +the new path — two mechanisms for one behaviour is how they drift. + +An unknown profile (`auto:nonsense`) must fail **loudly** with a 422 naming the +valid profiles. It must NOT silently fall back to `default`: a user who +mistypes `auto:locallity` would otherwise get cloud routing and a surprising +bill, with nothing indicating why. + +Config lives under a new top-level `profiles:` key. Config is strict +(`extra="forbid"`), so every field needs declaring; an unknown key is an error +by design. + +## 5. Observability + +`route_decisions` must record which profile served each decision, or the +proficiency loop cannot tell a `bigboybritches` win from a `default` one and +will read profile-forced selections as evidence of model quality. This is the +same reasoning that made the local fallback record +`kind='local_dispatch_fallback'` rather than `chat`. + +Add a `profile` column (code-side migration in `ensure_route_decisions`, as +every prior column addition did — `config/schema.sql` is +`CREATE TABLE IF NOT EXISTS` and does not alter a live DB). Surface it in the +admin decisions table and its filters, which already filter by kind, category +and tier. + +## 6. Interactions to get right + +- **`eligible_categories`** (local dispatch) is a per-MODEL filter; a profile + is a per-REQUEST filter. They compose as AND. A `locality` profile must not + bypass `eligible_categories` — that gate is what keeps a summarization-grade + model away from code, and `nemotron-mini:4b` detecting 0/4 bugs is why. +- **The local-dispatch fallback** (`kind='local_dispatch_fallback'`) is a + degraded path, not a profile. Leave it alone. Under `locality` the local + model is the primary choice, so the fallback should simply never fire. +- **Circuit breaker and admin overrides** already feed `exclude_models`. + Profiles must AND with them, never replace them: a profile must not resurrect + a model an operator deprecated or the breaker has marked down. +- **`exploration`** (epsilon-greedy) picks from eligible candidates. Confirm it + reads the profile-filtered set, or exploration will select models the profile + excluded. + +## Non-goals + +- No new ranking algorithm. Quality-first with cost as tiebreak stays. +- No change to `objective.quality_tolerance`'s global value. +- No per-category preference logic (that was §1(a) of the previous plan and is + explicitly out of scope here). +- No changes to the local-dispatch fallback path. +- Do not touch the classifier: profiles select candidates, not categories. + +## Success criteria + +- `auto` and `auto:batch` behave **exactly** as today, pinned by tests that + would catch a regression in the general mechanism. +- At least the three named profiles ship, defined as predicates over catalog + columns where possible. +- An unknown profile returns 422 naming the valid ones; it never silently + degrades to `default`. +- A profile that empties the candidate set returns a rejection reason naming + the profile. +- `route_decisions.profile` is populated and visible in the admin decisions + table. +- Profiles AND with circuit-breaker exclusions, admin overrides, and + `eligible_categories` — proven by a test for each, not by inspection. +- `rejection_reason` remains the single copy of the eligibility rules. +- Full suite green with `local_energy.enabled` both true and false (the Plan 5 + invariant must not regress). +- The user's `config/config.yaml` tariff lines remain uncommitted and verbatim. -- 2.49.1 From e975c70e03ac03c7d1b3b2d2ac3f50a4adaebf3e Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Thu, 3 Sep 2026 23:29:41 -0400 Subject: [PATCH 02/10] docs(plans): profiles are filter-only (user decision) A profile restricts the candidate set; ranking stays quality-first with cost as the tiebreak. Filter-only is sufficient for all three named profiles because a pure-filter 'locality' leaves no cloud row to lose to -- the 0.767-vs-0.95 dormancy only blocks MIXED sets that prefer local without excluding cloud. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- plans/named-routing-profiles.md | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) diff --git a/plans/named-routing-profiles.md b/plans/named-routing-profiles.md index 46f8c28..d4e5021 100644 --- a/plans/named-routing-profiles.md +++ b/plans/named-routing-profiles.md @@ -1,7 +1,7 @@ # Named routing profiles -**Status: FINAL on architecture and scope. ONE open decision (§3) that is the -user's, flagged inline.** Written 2026-09-03 against `main` at `c2893b5`. +**Status: FINAL — decision-complete. §3 was answered by the user on 2026-09-03 +(filter only).** Written 2026-09-03 against `main` at `c2893b5`. ## What this is @@ -61,7 +61,15 @@ the cases where that is genuinely what the user means. ## 3. THE OPEN DECISION — filter only, or objective overrides too? -**This is the user's call and must not be made silently.** +**DECIDED 2026-09-03 by the user: FILTER ONLY.** A profile restricts the +candidate set and nothing else. Ranking stays quality-first with cost as the +tiebreak, one rule for the whole system. Do NOT implement objective overrides, +and do NOT add a `quality_tolerance` field to a profile definition — if a +future preference-shaped profile needs one, that is a separate plan with its +own justification. + +The reasoning is preserved below because it explains WHY filter-only is +sufficient, which is not obvious. A profile is unambiguously a candidate-set filter. The question is whether it may ALSO override `objective.*` (`quality_tolerance`, -- 2.49.1 From 60c054bcf157d3d65a0efafa3cc11204a529b4ba Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:19:31 -0400 Subject: [PATCH 03/10] docs(plans): capability-aware ceiling warnings + reactive rejection detector From a live incident on 2026-09-04: two image requests 422'd with "tier >= 1; context >= 242486 tokens; interactive; vision-capable model". Admin overrides had kimi-k3/-fast/-flex deprecated, and those are the only vision-capable rows with enough context (782,324). Every remaining active vision model tops out at 192,500. Incident #3 recurring through a dimension the existing detector does not model. The existing machinery is fine and must not be rebuilt: warnings ARE surfaced (under coverage.warnings, not a top-level key -- a first pass at this analysis got that wrong), context_ceilings already applies admin deprecations, and demand_ceiling_warnings already has the correct demand-relative shape. Gap 1: ceilings bucket by (tier, latency_tolerance) only. During the incident the tier-1 interactive ceiling was still 782,324 across all models, so every check stayed silent while the vision-capable ceiling had collapsed to 192,500. Add vision and json_mode sub-ceilings -- the two hard filters that fail closed on NULL and can independently empty the set. Gap 2, and the more valuable half: nothing reads route_decisions' recorded rejections. A rejection-rate warning would have caught this in minutes without modelling any capability dimension, and would catch the next failure through a dimension nobody predicted. Predictive checks only catch what you thought of. Non-goals are explicit: observability only, no routing changes, and do NOT auto-revert admin overrides -- those deprecations are deliberate operator cost decisions. Warn, do not act. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- plans/capability-aware-ceiling-warnings.md | 136 +++++++++++++++++++++ 1 file changed, 136 insertions(+) create mode 100644 plans/capability-aware-ceiling-warnings.md diff --git a/plans/capability-aware-ceiling-warnings.md b/plans/capability-aware-ceiling-warnings.md new file mode 100644 index 0000000..1aa90d8 --- /dev/null +++ b/plans/capability-aware-ceiling-warnings.md @@ -0,0 +1,136 @@ +# Capability-aware ceiling warnings, and a reactive rejection detector + +**Status: FINAL — decision-complete.** Written 2026-09-04 from a live incident. + +## The incident + +On 2026-09-04 two image requests failed with: + +``` +422 tier >= 1; context >= 242486 tokens; interactive; + vision-capable model (request carries image(s)) +``` + +Cause: admin availability overrides had `kimi-k3`, `kimi-k3-fast` and +`kimi-k3-flex` deprecated. Those are the ONLY vision-capable rows with enough +context (782,324). Every remaining active vision model tops out at 192,500, +below the request's 242,486. The intersection of the hard filters was empty. + +This is [incident #3](../docs/incidents.md) recurring through a dimension the +existing detector does not model. Resolved by re-activating `kimi-k3`; verified +by recomputing the candidate set through `routing.select_candidates`, which now +returns exactly one row. + +## What already exists — do not rebuild it + +Establish this before writing code, because a first pass at this analysis got +it wrong by checking for a top-level `warnings` key: + +- **`/metrics` surfaces warnings under `coverage.warnings`**, not at top level. + `/admin/api/snapshot` carries the same block. The plumbing is done. +- **`metrics.context_ceilings` already applies admin deprecations** via + `exclude_models`, which was the fix after incident #3, and delegates to + `routing.select_candidates` so the ceiling matches live routing. +- **`metrics.demand_ceiling_warnings` already compares ceiling against observed + demand**, which is the *correct* shape. Do not replace it with an absolute + tier-ordering check: `ceiling(1) >= ceiling(2) >= ceiling(3)` is a theorem + (tier is a capability floor, so the eligible set shrinks monotonically), so + such a warning fires always and means nothing. `docs/incidents.md` #3 records + this trap; do not re-enter it. +- **`route_decisions` already records every rejection** — `selected_model IS + NULL` with the full `rejected_reason` string. + +## Gap 1 — ceilings are blind to capability-gated subsets + +`context_ceilings` buckets by `(tier, latency_tolerance)` only. During the +incident the tier-1 interactive ceiling across all models was still **782,324** +(deepseek and others are large and active), so every existing check was silent +while the *vision-capable* ceiling had collapsed to 192,500. + +Measured during the incident: + +| subset | ceiling | +|---|---| +| all active models | 782,324 | +| **vision-capable only** | **192,500** | + +The hard filters that can independently empty the candidate set, from +`routing.rejection_reason`, are `require_vision` and `require_json_mode` — +both **fail closed on NULL** (an unconfirmed capability is treated as absent), +which is what makes them able to shrink the set sharply. + +**Add capability sub-ceilings.** Extend the bucket key, or compute a small +number of additional named buckets: `vision` and `json_mode`. Do NOT add a +bucket per capability combination — that is a combinatorial explosion for +dimensions that do not interact in practice. Two extra series, compared against +the demand actually observed for requests carrying images / requesting JSON +(`route_decisions.images`, `.json_mode` are already recorded), is enough. + +Reuse `demand_ceiling_warnings`' existing demand-relative shape; only the +bucketing changes. + +## Gap 2 — nothing watches actual rejections + +This is the more valuable half, and it is simpler. + +Every hard-filter rejection is already persisted with its reason. Nothing reads +them. A detector that says *"N requests were rejected in the last hour"* would +have caught this incident within minutes, **without modelling any capability +dimension at all** — and would equally catch the next failure through a +dimension nobody predicted. + +Predictive checks (Gap 1) only catch dimensions you thought of. A reactive +check catches everything, at the cost of firing after the first failure rather +than before. Ship both; the reactive one is the safety net. + +Add a warning in `metrics.scoring_coverage`'s existing `warnings` list: + +- Count `route_decisions` rows with `selected_model IS NULL AND rejected_reason + IS NOT NULL` in a recent window (1 hour and 24 hours are both useful; pick + one, state which). +- Group by a **normalized** reason, not the literal string. The raw strings + embed the request's token count (`context >= 242486 tokens`), so grouping by + the literal yields n=1 per row and hides the pattern. Normalize by replacing + digit runs with a placeholder before grouping. +- Report the count, the normalized reason, and the most recent timestamp, so an + operator can tell a live problem from an old one. +- **Zero rejections must produce no warning.** A rejection is not inherently an + error — a genuinely impossible request should 422. The signal is a *rate*. + +## Non-goals + +- Do not change `routing.py`. Nothing about the hard filters is wrong; the + candidate set was correctly empty. This plan is observability only. +- Do not auto-revert admin overrides or "helpfully" re-activate models. The + deprecations were deliberate operator cost decisions. Warn; do not act. +- Do not add a bucket per capability combination. +- Do not replace `demand_ceiling_warnings` with an absolute tier-ordering + comparison (see above — it is a theorem). +- Do not touch the `local_vision` fallback. It could not have rescued this + request anyway: that path receives the raw unpruned message list, and 242,486 + tokens against `num_ctx 8192` was never going to fit. Worth a doc note, not a + code change. + +## Success criteria + +- A test reproducing the incident state — `kimi-k3*` excluded via + `admin_model_overrides`, a vision request above 192,500 tokens — produces a + warning naming the vision subset. The same state produces **no** warning from + the existing `(tier, latency_tolerance)` checks, proving the gap was real and + is now closed. +- A test with recent NULL-selection rows produces a rejection-rate warning + reporting a normalized reason and a count; a test with zero such rows produces + none. +- Both warnings appear in `coverage.warnings` on `/metrics` and in + `/admin/api/snapshot`, and are visible on the admin dashboard's warnings bell. +- `demand_ceiling_warnings` keeps its demand-relative shape; no absolute + tier-ordering check is introduced. +- Full suite green with `local_energy.enabled` both true and false. +- The user's `config/config.yaml` tariff lines remain uncommitted and verbatim. + +## Related, out of scope but worth recording + +`coverage.warnings` currently reports **metered usage at 146% of the 6.25 kWh +plan allowance** (9.15 kWh, next reset 2026-09-06). That warning is working as +designed and is almost certainly why provider credits ran out on 2026-09-02/03. +It needs an operator decision about the plan, not a code change. -- 2.49.1 From 877f70699600251614f33cf961043ed60b89e467 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:52:49 -0400 Subject: [PATCH 04/10] feat(routing): add restrict_to allowlist primitive Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus --- src/routing.py | 13 +++++++-- tests/test_routing.py | 65 +++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 75 insertions(+), 3 deletions(-) diff --git a/src/routing.py b/src/routing.py index 591708b..d1cdd85 100644 --- a/src/routing.py +++ b/src/routing.py @@ -73,6 +73,7 @@ def rejection_reason( require_vision: bool = False, require_json_mode: bool = False, task_category: str | None = None, + restrict_to: set[str] | None = None, ) -> str | None: """Why this row is not a candidate, or None if it is one. @@ -159,6 +160,11 @@ def rejection_reason( if eligible is not None and (task_category is None or task_category not in eligible): return "category_ineligible" + # Profile allowlist. None means unrestricted; an empty set means nothing is + # eligible, and a non-empty set means only those model ids may be selected. + if restrict_to is not None and row.get("model_id") not in restrict_to: + return "profile_excluded" + return None @@ -291,11 +297,10 @@ def apply_flex_preference( ) is not None: return selected_row, False, False, selected_row.get("cost") - if pref == "prefer-flex": + if pref == "prefer-flex" and latency_tolerance == INTERACTIVE: # A flex sibling is eligible under the current hard filters unless the # latency filter excludes flex rows — which is exactly interactive. - if latency_tolerance == INTERACTIVE: - return selected_row, False, False, selected_row.get("cost") + return selected_row, False, False, selected_row.get("cost") # prefer-flex under batch, or force-flex (which bypasses interactive). flex_forced = pref == "force-flex" and latency_tolerance == INTERACTIVE @@ -324,6 +329,7 @@ def select_candidates( require_vision: bool = False, require_json_mode: bool = False, task_category: str | None = None, + restrict_to: set[str] | None = None, ) -> list[dict]: """Apply every hard filter, preserving input order.""" return [ @@ -342,6 +348,7 @@ def select_candidates( require_vision=require_vision, require_json_mode=require_json_mode, task_category=task_category, + restrict_to=restrict_to, ) ] diff --git a/tests/test_routing.py b/tests/test_routing.py index afb1058..9ed0ba3 100644 --- a/tests/test_routing.py +++ b/tests/test_routing.py @@ -890,6 +890,71 @@ def test_category_ineligible_reason_is_single_token(): assert " " not in reason +# --- restrict_to profile allowlist ---------------------------------------- + + +def test_restrict_to_none_is_unrestricted(): + row = _row(model_id="m1") + assert _eligible(row, restrict_to=None) is True + + +def test_restrict_to_allows_model_in_set(): + row = _row(model_id="m1") + assert _eligible(row, restrict_to={"m1"}) is True + + +def test_restrict_to_excludes_model_not_in_set(): + row = _row(model_id="m2") + assert _eligible(row, restrict_to={"m1"}) is False + + +def test_restrict_to_empty_set_excludes_every_model(): + row = _row(model_id="m1") + assert _eligible(row, restrict_to=set()) is False + + +def test_restrict_to_reason_is_profile_excluded(): + row = _row(model_id="m2") + reason = _reason(row, restrict_to={"m1"}) + assert reason == "profile_excluded" + + +def test_restrict_to_is_single_token_reason(): + row = _row(model_id="m2") + reason = _reason(row, restrict_to={"m1"}) + assert " " not in reason + + +def test_select_candidates_applies_restrict_to_allowlist(): + rows = [_row(model_id="m1"), _row(model_id="m2"), _row(model_id="m3")] + selected = select_candidates( + rows, required_context_tokens=10_000, required_tier=2, + latency_tolerance="interactive", allowed_access_levels=["public"], + exclude_stale=True, exclude_deprecated=True, restrict_to={"m1"}) + assert [r["model_id"] for r in selected] == ["m1"] + + +def test_select_candidates_empty_restrict_to_yields_no_candidates(): + rows = [_row(model_id="m1"), _row(model_id="m2")] + selected = select_candidates( + rows, required_context_tokens=10_000, required_tier=2, + latency_tolerance="interactive", allowed_access_levels=["public"], + exclude_stale=True, exclude_deprecated=True, restrict_to=set()) + assert selected == [] + + +def test_restrict_to_and_exclude_models_compose_as_and(): + # An allowlist of {m1, m2} plus a circuit-open exclusion of m1 should + # leave only m2. + rows = [_row(model_id="m1"), _row(model_id="m2"), _row(model_id="m3")] + selected = select_candidates( + rows, required_context_tokens=10_000, required_tier=2, + latency_tolerance="interactive", allowed_access_levels=["public"], + exclude_stale=True, exclude_deprecated=True, + exclude_models={"m1"}, restrict_to={"m1", "m2"}) + assert [r["model_id"] for r in selected] == ["m2"] + + def test_select_candidates_drops_outside_category(): # Even a perfect row (tier 1, zero cost) is dropped for the wrong category. rows = [ -- 2.49.1 From bd73302ae4ee675f6a99c5116df2882f4ef55543 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:52:51 -0400 Subject: [PATCH 05/10] feat(config): add named routing profiles configuration Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus --- src/config.py | 53 ++++++++++++++++++++ tests/test_config.py | 114 +++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 167 insertions(+) create mode 100644 tests/test_config.py diff --git a/src/config.py b/src/config.py index 6ca7c9d..8b25f77 100644 --- a/src/config.py +++ b/src/config.py @@ -194,6 +194,58 @@ class FlexPreference(Enum): force_flex = "force-flex" +class RoutingProfile(StrictModel): + """A named candidate-set filter requested through ``auto:``. + + Profiles restrict the candidate set; they do NOT change ranking. All + fields are optional, and any absent field falls back to the request's + own value or the global default. This keeps profiles small rewrite + rules rather than a parallel config layer. + """ + + provider: Optional[str] = None + min_tier: Optional[int] = None + max_tier: Optional[int] = None + latency_tolerance: Optional[Literal["interactive", "batch"]] = None + max_cost_per_1m_completion: Optional[float] = None + allowed_model_ids: Optional[set[str]] = None + + @field_validator("min_tier", "max_tier") + @classmethod + def tier_in_range(cls, v: Optional[int]) -> Optional[int]: + if v is not None and v not in (1, 2, 3): + raise ValueError( + f"profiles[...].{cls.__name__}.tier must be in {{1, 2, 3}}, got {v}" + ) + return v + + @model_validator(mode="after") + def min_not_above_max(self) -> "RoutingProfile": + if self.min_tier is not None and self.max_tier is not None: + if self.min_tier > self.max_tier: + raise ValueError( + f"profiles[...].{self.__class__.__name__}.min_tier ({self.min_tier}) " + f"must be <= max_tier ({self.max_tier})" + ) + return self + + @field_validator("max_cost_per_1m_completion") + @classmethod + def cost_positive(cls, v: Optional[float]) -> Optional[float]: + if v is not None and v <= 0: + raise ValueError( + "profiles[...].max_cost_per_1m_completion must be > 0, or null" + ) + return v + + @field_validator("allowed_model_ids") + @classmethod + def allowed_model_ids_non_empty(cls, v: Optional[set[str]]) -> Optional[set[str]]: + if v is not None and not v: + raise ValueError("profiles[...].allowed_model_ids must be non-empty when set") + return v + + class RoutingConfig(StrictModel): allowed_access_levels: list[str] default_latency_tolerance: str @@ -704,6 +756,7 @@ class RouterConfig(StrictModel): logging: LoggingConfig local_energy: LocalEnergyConfig = LocalEnergyConfig() local_dispatch_models: list[LocalDispatchModel] = [] + profiles: dict[str, RoutingProfile] = {} @model_validator(mode="after") def local_energy_needs_tariff_when_enabled(self) -> "RouterConfig": diff --git a/tests/test_config.py b/tests/test_config.py new file mode 100644 index 0000000..3eb01d8 --- /dev/null +++ b/tests/test_config.py @@ -0,0 +1,114 @@ +"""Load-time validation for ``RouterConfig`` and its nested models. + +This file covers core config shape tests that are not tied to the service +endpoints in ``test_config_endpoints.py``. +""" + +from __future__ import annotations + +import copy +from pathlib import Path + +import pytest +import yaml + +from config import RouterConfig + +ROOT = Path(__file__).resolve().parent.parent + + +@pytest.fixture +def raw() -> dict: + with open(ROOT / "config" / "config.yaml") as fh: + return yaml.safe_load(fh) + + +# --- routing profiles ------------------------------------------------------- + + +def test_profiles_empty_dict_defaults(raw): + """A missing ``profiles`` section defaults to an empty dict.""" + cfg = copy.deepcopy(raw) + cfg.pop("profiles", None) + loaded = RouterConfig(**cfg) + assert loaded.profiles == {} + + +def test_a_valid_locality_profile_loads(raw): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"locality": {"provider": "ollama-local"}} + loaded = RouterConfig(**cfg) + assert loaded.profiles["locality"].provider == "ollama-local" + + +def test_a_profile_with_all_fields_loads(raw): + cfg = copy.deepcopy(raw) + cfg["profiles"] = { + "onlycheaps": { + "min_tier": 1, + "max_tier": 2, + "latency_tolerance": "batch", + "max_cost_per_1m_completion": 0.50, + "allowed_model_ids": {"foo", "bar"}, + } + } + loaded = RouterConfig(**cfg) + profile = loaded.profiles["onlycheaps"] + assert profile.min_tier == 1 + assert profile.max_tier == 2 + assert profile.latency_tolerance == "batch" + assert profile.max_cost_per_1m_completion == 0.50 + assert profile.allowed_model_ids == {"foo", "bar"} + + +def test_profile_with_unknown_key_is_rejected(raw): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"locality": {"provider": "ollama-local", "quality_tolerance": 0.2}} + with pytest.raises(ValueError, match="quality_tolerance"): + RouterConfig(**cfg) + + +@pytest.mark.parametrize( + "tier_key,tier_value", + [ + ("min_tier", 0), + ("min_tier", 4), + ("max_tier", 0), + ("max_tier", 4), + ], +) +def test_profile_tier_out_of_range_is_rejected(raw, tier_key, tier_value): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"bigboybritches": {tier_key: tier_value}} + with pytest.raises(ValueError, match="tier"): + RouterConfig(**cfg) + + +def test_profile_min_tier_above_max_tier_is_rejected(raw): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"bad": {"min_tier": 3, "max_tier": 1}} + with pytest.raises(ValueError, match="min_tier.*max_tier"): + RouterConfig(**cfg) + + +@pytest.mark.parametrize("latency", ["realtime", "", "BATCH"]) +def test_profile_bad_latency_tolerance_is_rejected(raw, latency): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"locality": {"latency_tolerance": latency}} + with pytest.raises(ValueError, match="latency_tolerance"): + RouterConfig(**cfg) + + +@pytest.mark.parametrize("cost", [0, -1.5]) +def test_profile_nonpositive_cost_is_rejected(raw, cost): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"locality": {"max_cost_per_1m_completion": cost}} + with pytest.raises(ValueError, match="max_cost_per_1m_completion"): + RouterConfig(**cfg) + + +def test_profile_empty_allowed_model_ids_is_rejected(raw): + cfg = copy.deepcopy(raw) + cfg["profiles"] = {"locality": {"allowed_model_ids": []}} + with pytest.raises(ValueError, match="allowed_model_ids"): + RouterConfig(**cfg) -- 2.49.1 From 320a8c96e84f06021181fef7a6a33d34d0a8fa3c Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:52:53 -0400 Subject: [PATCH 06/10] feat(dispatcher): resolve auto: built-in/config profiles Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus --- src/dispatcher.py | 165 ++++++++++++++++++++++++++++----- tests/test_chat_completions.py | 70 ++++++++++++++ 2 files changed, 211 insertions(+), 24 deletions(-) diff --git a/src/dispatcher.py b/src/dispatcher.py index 7888c23..1fd704c 100644 --- a/src/dispatcher.py +++ b/src/dispatcher.py @@ -46,6 +46,7 @@ import time from datetime import datetime, timezone from pathlib import Path from statistics import median +from collections.abc import Sequence from typing import Any, Literal, Optional import requests @@ -63,7 +64,7 @@ import local_energy import logs import session_cache from capabilities import detect_capabilities, iter_image_url_values -from config import FlexPreference, RouterConfig, load_config +from config import FlexPreference, RouterConfig, RoutingProfile, load_config from context_prune import ( _text_only, estimate_tokens, @@ -84,8 +85,6 @@ from metrics import ( verdict_mix, ) from routing import ( - BATCH, - INTERACTIVE, apply_flex_preference, capability_gate_reason, rank_candidates, @@ -104,7 +103,17 @@ from verification import ( # Virtual model names that mean "you pick". Anything else is taken as a real # model id and dispatched as asked. ROUTER_MODEL = "auto" -ROUTER_MODEL_BATCH = "auto:batch" + +# Built-in routing profiles. Profiles are selectable as `auto:`; they +# restrict the candidate set and may override the effective latency_tolerance. +# Unknown profiles raise 422, so unknown built-ins do not silently fall back. +BUILTIN_PROFILES: dict[str, RoutingProfile] = { + "default": RoutingProfile(), + "batch": RoutingProfile(latency_tolerance="batch"), + "locality": RoutingProfile(provider="ollama-local"), + "bigboybritches": RoutingProfile(min_tier=3), + "onlycheaps": RoutingProfile(max_cost_per_1m_completion=0.5), +} # Observations from the fixed reference workload in seed_energy.py. Only # these steer routing; organic traffic is logged for accounting but varies @@ -219,6 +228,13 @@ class TaskRequest(BaseModel): "requires a JSON-mode-capable model (routing.require_json_mode).", ) # Overrides, mostly for testing the router without the classifier in the loop. + profile: Optional[str] = Field( + None, + description=( + "Named routing profile (e.g. 'locality', 'batch') to apply for this " + "request. Defaults to the 'default' profile." + ), + ) task_category: Optional[str] = None task_tier: Optional[int] = Field(None, ge=1, le=3) required_context_tokens: Optional[int] = Field(None, ge=0) @@ -258,6 +274,7 @@ class Candidate(BaseModel): class RouteResponse(BaseModel): classification: Classification latency_tolerance: str + profile: str = "default" flex_preference: str = "auto" flex_swapped: bool = False flex_forced: bool = False @@ -286,6 +303,76 @@ def _ms(started: float) -> int: return int((time.perf_counter() - started) * 1000) +def _all_profile_names() -> list[str]: + names = set(BUILTIN_PROFILES) | set(cfg.profiles) + return sorted(names) + + +def _resolve_profile(requested: str) -> tuple[str, RoutingProfile]: + if requested == ROUTER_MODEL: + name = "default" + else: + name = requested[len(ROUTER_MODEL) + 1 :] # after "auto:" + return _resolve_profile_name(name) + + +def _resolve_profile_name(name: str) -> tuple[str, RoutingProfile]: + """Resolve a bare profile name, returning (name, profile_obj) or 422.""" + builtin = BUILTIN_PROFILES.get(name) + configured = cfg.profiles.get(name) + if builtin is None and configured is None: + valid = _all_profile_names() + raise HTTPException( + 422, + f"Unknown routing profile {name!r}. Valid profiles: {valid}", + ) + if configured is not None and builtin is not None: + merged = builtin.model_copy( + update={ + k: v + for k, v in configured.model_dump().items() + if v is not None + } + ) + return name, merged + return name, (configured if configured is not None else builtin) + + +def _effective_latency(profile: RoutingProfile, req: Optional[TaskRequest]) -> str: + return ( + profile.latency_tolerance + or (req.latency_tolerance if req is not None else None) + or cfg.routing.default_latency_tolerance + ) + + +def _restrict_to_from_profile(profile: RoutingProfile, rows: Sequence[dict]) -> set[str] | None: + allowed: set[str] = {r["model_id"] for r in rows} + if profile.provider is not None: + allowed &= {r["model_id"] for r in rows if r.get("provider") == profile.provider} + if profile.min_tier is not None or profile.max_tier is not None: + def tier_ok(r: dict) -> bool: + tier = r.get("tier") + if tier is None: + return False + if profile.min_tier is not None and tier < profile.min_tier: + return False + if profile.max_tier is not None and tier > profile.max_tier: + return False + return True + allowed &= {r["model_id"] for r in rows if tier_ok(r)} + if profile.max_cost_per_1m_completion is not None: + allowed &= { + r["model_id"] + for r in rows + if r.get("cost_per_1m_completion") is not None + and r["cost_per_1m_completion"] <= profile.max_cost_per_1m_completion + } + if profile.allowed_model_ids is not None: + allowed &= profile.allowed_model_ids + return allowed if allowed != {r["model_id"] for r in rows} else None + + def _db() -> sqlite3.Connection: conn = sqlite3.connect(cfg.database.path) conn.row_factory = sqlite3.Row @@ -338,7 +425,8 @@ def ensure_route_decisions(conn: sqlite3.Connection) -> None: request_id TEXT, exploration INTEGER DEFAULT 0, pinch_original_tokens INTEGER, - pinch_final_tokens INTEGER + pinch_final_tokens INTEGER, + profile TEXT ) """ ) @@ -360,6 +448,7 @@ def ensure_route_decisions(conn: sqlite3.Connection) -> None: ("exploration", "INTEGER DEFAULT 0"), ("pinch_original_tokens", "INTEGER"), ("pinch_final_tokens", "INTEGER"), + ("profile", "TEXT"), ): if name not in existing: conn.execute(f"ALTER TABLE route_decisions ADD COLUMN {name} {decl}") @@ -834,8 +923,15 @@ def _to_candidate(row: dict) -> Candidate: return Candidate(**fields) -def route(req: TaskRequest) -> RouteResponse: - latency_tolerance = req.latency_tolerance or cfg.routing.default_latency_tolerance +def route( + req: TaskRequest, + *, + profile: str = "default", + profile_obj: Optional[RoutingProfile] = None, +) -> RouteResponse: + if profile_obj is None: + profile_obj = BUILTIN_PROFILES["default"] + latency_tolerance = _effective_latency(profile_obj, req) if req.task_category and req.task_tier and req.required_context_tokens is not None: classification = Classification( @@ -882,7 +978,8 @@ def route(req: TaskRequest) -> RouteResponse: ), task_category=classification.task_category, ) - eligible = select_candidates(rows, **filters) + restrict_to = _restrict_to_from_profile(profile_obj, rows) + eligible = select_candidates(rows, restrict_to=restrict_to, **filters) if logs.enabled_for_debug(): # "No model satisfies the hard filters" is otherwise a dead end with no # explanation. One line per drop, naming the filter and its numbers, is @@ -976,6 +1073,7 @@ def route(req: TaskRequest) -> RouteResponse: return RouteResponse( classification=classification, latency_tolerance=latency_tolerance, + profile=profile, flex_preference=flex_pref.value, flex_swapped=flex_swapped, flex_forced=flex_forced, @@ -1056,6 +1154,7 @@ def persist_route_decision( exploration=0, pinch_original_tokens=None, pinch_final_tokens=None, + profile: Optional[str] = None, ) -> Optional[int]: """Record one routing decision to route_decisions, best-effort and gated. @@ -1088,6 +1187,7 @@ def persist_route_decision( derived_runners = None candidates = None derived_latency = latency_tolerance + derived_profile = profile if isinstance(classification, RouteResponse): clf = classification.classification @@ -1095,6 +1195,7 @@ def persist_route_decision( derived_runners = classification.runners_up candidates = classification.candidates_considered derived_latency = classification.latency_tolerance + derived_profile = classification.profile flex_preference = classification.flex_preference flex_swapped = int(classification.flex_swapped) flex_forced = int(classification.flex_forced) @@ -1151,8 +1252,8 @@ def persist_route_decision( est_cost_usd, est_proficiency, rejected_reason, session_key, tools, images, json_mode, streamed, flex_preference, flex_swapped, flex_forced, exploration, - request_id, pinch_original_tokens, pinch_final_tokens - ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) + request_id, pinch_original_tokens, pinch_final_tokens, profile + ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) """, ( observed_at, @@ -1183,6 +1284,7 @@ def persist_route_decision( None, pinch_original_tokens, pinch_final_tokens, + derived_profile, ), ) decision_id: Optional[int] = int(cursor.lastrowid) @@ -1222,6 +1324,7 @@ def persist_route_decision( "exploration": exploration, "pinch_original_tokens": pinch_original_tokens, "pinch_final_tokens": pinch_final_tokens, + "profile": derived_profile, } ) return decision_id @@ -1898,7 +2001,8 @@ def route_endpoint(req: TaskRequest): """Classify and pick a model without calling it.""" logs.new_trace() started = time.perf_counter() - decision = route(req) + profile_name, profile_obj = _resolve_profile_name(req.profile or "default") + decision = route(req, profile=profile_name, profile_obj=profile_obj) log_decision( decision, tools=req.tools_present, @@ -2859,9 +2963,11 @@ def list_models(): conn.close() data = [ - {"id": ROUTER_MODEL, "object": "model", "owned_by": "router"}, - {"id": ROUTER_MODEL_BATCH, "object": "model", "owned_by": "router"}, + {"id": f"{ROUTER_MODEL}:{name}", "object": "model", "owned_by": "router"} + for name in _all_profile_names() ] + # `auto` (default profile) is exposed as a bare alias for convenience. + data.insert(0, {"id": ROUTER_MODEL, "object": "model", "owned_by": "router"}) data += [ {"id": r["model_id"], "object": "model", "owned_by": r["provider"]} for r in rows @@ -2934,7 +3040,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): # check, the same unstripped id would have gone upstream and drawn a 400 # from NeuralWatt, which only knows the bare form. bare = requested.rsplit("/", 1)[-1] - wants_routing = bare in (ROUTER_MODEL, ROUTER_MODEL_BATCH) + wants_routing = bare == ROUTER_MODEL or bare.startswith(f"{ROUTER_MODEL}:") if wants_routing or ( bare != requested and (_model_exists(bare) or _local_dispatch_config_for(bare) is not None) @@ -2952,7 +3058,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): ) if wants_routing: - latency = BATCH if requested == ROUTER_MODEL_BATCH else INTERACTIVE + profile_name, profile_obj = _resolve_profile(requested) tools_present = caps.tools_present # The turn before the last: a short follow-up inherits its complexity, # so the classifier sees "Context: \n---\nMessage: " @@ -3009,17 +3115,20 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): task=_last_user_text(messages), # context omitted: the override branch (task_category + # task_tier + required_context_tokens all provided) - # never reads req.context — it skips classify(). - latency_tolerance=latency, + # never reads req.context -- it skips classify(). + latency_tolerance=_effective_latency(profile_obj, None), tools_present=tools_present, has_images=caps.has_images, require_json_mode=caps.require_json_mode, task_category=cached.task_category, task_tier=cached.task_tier, required_context_tokens=measured, - ) + ), + profile=profile_name, + profile_obj=profile_obj, ) else: + # Cache miss: classify as today, then re-route on measured context # if the conversation outgrows the classifier's own estimate. After # the decision, a successful (non-fallback) classification is @@ -3029,14 +3138,16 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): TaskRequest( task=_last_user_text(messages), context=prev_context, - latency_tolerance=latency, + latency_tolerance=_effective_latency(profile_obj, None), tools_present=tools_present, has_images=caps.has_images, require_json_mode=caps.require_json_mode, # Take whichever is larger: what the classifier thinks it # needs, or what the conversation actually measures. required_context_tokens=None, - ) + ), + profile=profile_name, + profile_obj=profile_obj, ) # The classifier's own verdict, before the re-route below rewrites # the Classification's source to 'override'. @@ -3050,15 +3161,17 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): task=_last_user_text(messages), # context omitted: the override branch (task_category + # task_tier + required_context_tokens all provided) - # never reads req.context — it skips classify(). - latency_tolerance=latency, + # never reads req.context -- it skips classify(). + latency_tolerance=_effective_latency(profile_obj, None), tools_present=tools_present, has_images=caps.has_images, require_json_mode=caps.require_json_mode, task_category=decision.classification.task_category, task_tier=decision.classification.task_tier, required_context_tokens=measured, - ) + ), + profile=profile_name, + profile_obj=profile_obj, ) if ( cfg.session_cache.enabled @@ -3092,6 +3205,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): flex_preference=cfg.routing.default_flex_preference.value, flex_swapped=0, flex_forced=0, + profile=None, ) return _local_vision_response( fallback, streaming=bool(body.get("stream")) @@ -3211,6 +3325,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): flex_preference=cfg.routing.default_flex_preference.value, flex_swapped=0, flex_forced=0, + profile=None, pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None, pinch_final_tokens=pinch_stats.get("final_tokens") @@ -3305,6 +3420,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): flex_preference=decision.flex_preference, flex_swapped=decision.flex_swapped, flex_forced=decision.flex_forced, + profile=decision.profile, ) _write_request_id(fallback_row, result["request_id"]) return _local_dispatch_response( @@ -3686,7 +3802,8 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks): def dispatch_endpoint(req: TaskRequest): logs.new_trace() started = time.perf_counter() - decision = route(req) + profile_name, profile_obj = _resolve_profile_name(req.profile or "default") + decision = route(req, profile=profile_name, profile_obj=profile_obj) log_decision( decision, tools=req.tools_present, diff --git a/tests/test_chat_completions.py b/tests/test_chat_completions.py index 062358e..34b455f 100644 --- a/tests/test_chat_completions.py +++ b/tests/test_chat_completions.py @@ -212,6 +212,36 @@ def test_auto_routes_and_reports_the_model_it_actually_used(router): assert resp.json()["model"] == CHEAP +def test_auto_locality_profile_routes_to_locality_provider(router): + client, _, _ = router + resp = client.post( + "/v1/chat/completions", + json={"model": "auto:locality", "messages": _messages()}, + ) + assert resp.status_code == 422, "no ollama-local rows in fixture" + + +def test_routed_chat_named_profile_persists_profile(router): + client, calls, _ = router + resp = client.post( + "/v1/chat/completions", + json={"model": "auto:batch", "messages": _messages()}, + ) + assert resp.status_code == 200 + assert calls[0]["body"]["model"] == CHEAP + + db_path = dispatcher.cfg.database.path + conn = sqlite3.connect(db_path) + conn.row_factory = sqlite3.Row + rows = conn.execute( + "SELECT * FROM route_decisions ORDER BY id" + ).fetchall() + conn.close() + assert len(rows) == 1 + assert rows[0]["profile"] == "batch" + assert rows[0]["kind"] == "chat" + + def test_a_provider_prefixed_router_name_still_routes(router): """opencode sends `llm-router/auto`; only the virtual names are stripped.""" client, calls, _ = router @@ -682,6 +712,46 @@ def test_a_json_object_request_without_a_json_capable_model_422s(router): assert not calls +def test_auto_batch_sets_latency_tolerance_to_batch(router): + """auto:batch must resolve to the built-in batch profile.""" + client, calls, _ = router + resp = client.post( + "/v1/chat/completions", + json={"model": "auto:batch", "messages": _messages()}, + ) + assert resp.status_code == 200 + assert resp.headers["X-Router-Model"] == CHEAP + + +def test_unknown_profile_returns_422_naming_valid_profiles(router): + """A misspelled profile must not silently fall back to default.""" + client, calls, _ = router + resp = client.post( + "/v1/chat/completions", + json={"model": "auto:nosuchprofile", "messages": _messages()}, + ) + assert resp.status_code == 422 + detail = resp.json()["detail"] + assert "nosuchprofile" in detail + assert "default" in detail + assert "batch" in detail + assert not calls + + +def test_v1_models_lists_all_profiles(router): + """Every valid profile appears as auto: in the models list.""" + client, _, _ = router + resp = client.get("/v1/models") + assert resp.status_code == 200 + ids = {m["id"] for m in resp.json()["data"]} + assert "auto" in ids + assert "auto:default" in ids + assert "auto:batch" in ids + assert "auto:locality" in ids + assert "auto:bigboybritches" in ids + assert "auto:onlycheaps" in ids + + def test_a_plain_text_request_is_unaffected_by_the_gates(router): """No images, no response_format: routing is exactly as before.""" client, calls, _ = router -- 2.49.1 From 5aded249652404b5165f77749e7082959ccfb31b Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:52:56 -0400 Subject: [PATCH 07/10] feat(db): migrate route_decisions.profile and surface in metrics Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus --- config/schema.sql | 4 +- src/metrics.py | 2 +- tests/test_metrics_endpoint.py | 48 +++++++++ tests/test_route_decisions.py | 184 ++++++++++++++++++++++++++++++++- 4 files changed, 235 insertions(+), 3 deletions(-) diff --git a/config/schema.sql b/config/schema.sql index 1d21739..82f5bfc 100644 --- a/config/schema.sql +++ b/config/schema.sql @@ -245,8 +245,10 @@ CREATE TABLE IF NOT EXISTS route_decisions ( request_id TEXT, -- provider completion id exploration INTEGER DEFAULT 0, -- 0/1 whether rollout/exploration pinch_original_tokens INTEGER, -- estimated tokens before pruning - pinch_final_tokens INTEGER -- tokens sent after pruning (may + pinch_final_tokens INTEGER, -- tokens sent after pruning (may -- equal original when no pruning) + profile TEXT -- routing profile name (e.g., + -- 'default', 'locality') ); CREATE INDEX IF NOT EXISTS idx_verifications_model ON verifications (model_id, provider); diff --git a/src/metrics.py b/src/metrics.py index c40e29a..370427e 100644 --- a/src/metrics.py +++ b/src/metrics.py @@ -424,7 +424,7 @@ def recent_decisions( runner_up_models, est_cost_usd, est_proficiency, rejected_reason, session_key, tools, images, json_mode, streamed, flex_preference, flex_swapped, flex_forced, - pinch_original_tokens, pinch_final_tokens + pinch_original_tokens, pinch_final_tokens, profile FROM route_decisions ORDER BY id DESC LIMIT ? diff --git a/tests/test_metrics_endpoint.py b/tests/test_metrics_endpoint.py index 4d7e11b..0e427ec 100644 --- a/tests/test_metrics_endpoint.py +++ b/tests/test_metrics_endpoint.py @@ -381,3 +381,51 @@ async def test_decision_event_stream_replays_then_streams_live(): assert _sse_frame(live)["id"] == 2 finally: events.clear() + + +@pytest.mark.anyio +async def test_persist_route_decision_publishes_profile_in_sse_payload( + tmp_path, monkeypatch, +): + import asyncio + + events.clear() + db_path = tmp_path / "profile-sse.db" + conn = sqlite3.connect(db_path) + conn.executescript(SCHEMA_SQL) + conn.close() + + monkeypatch.setattr(dispatcher.cfg.database, "path", str(db_path)) + monkeypatch.setattr(dispatcher.cfg.logging, "log_route_decisions", True) + + sse_queue: asyncio.Queue[dict] = asyncio.Queue() + try: + events.subscribe_sse(sse_queue, replay=True) + + dispatcher.persist_route_decision( + "route", + classification=dispatcher.RouteResponse( + classification=dispatcher.Classification( + task_category="coding_general", task_tier=2, + required_context_tokens=100, confidence=0.9, + ), + latency_tolerance="interactive", + profile="batch", + selected=dispatcher.Candidate( + model_id="cheap", provider="neuralwatt", + tier=2, latency_class="standard", reasoning_mode="default", + context_variant="full", effective_context_window=128000, + composite=0.5, cost_score=0.5, proficiency_score=0.9, + ), + candidates_considered=1, + ), + latency_tolerance="interactive", + selected_model="cheap", + selected_provider="neuralwatt", + ) + + decision = await asyncio.wait_for(sse_queue.get(), timeout=2.0) + assert decision["profile"] == "batch" + finally: + events.unsubscribe_sse(sse_queue) + events.clear() diff --git a/tests/test_route_decisions.py b/tests/test_route_decisions.py index c26d7fa..3588f59 100644 --- a/tests/test_route_decisions.py +++ b/tests/test_route_decisions.py @@ -62,6 +62,7 @@ ROUTE_DECISIONS_COLUMNS = [ "exploration", "pinch_original_tokens", "pinch_final_tokens", + "profile", ] @@ -257,6 +258,130 @@ def test_persist_writes_pinch_columns(tmp_path, monkeypatch): conn.close() +def test_persist_route_decision_writes_profile_from_response(tmp_path, monkeypatch): + """A RouteResponse with profile='locality' writes that profile to the row.""" + db_path = tmp_path / "profile.db" + conn = sqlite3.connect(db_path) + conn.executescript(SCHEMA_SQL) + conn.close() + + monkeypatch.setattr(dispatcher.cfg.database, "path", str(db_path)) + monkeypatch.setattr(dispatcher.cfg.logging, "log_route_decisions", True) + + from dispatcher import Candidate + + dispatcher.persist_route_decision( + "route", + classification=dispatcher.RouteResponse( + classification=Classification( + task_category="coding_general", task_tier=2, + required_context_tokens=100, confidence=0.9, + ), + latency_tolerance="interactive", + profile="locality", + selected=Candidate( + model_id="deepseek-v4-flash", provider="neuralwatt", + tier=2, latency_class="standard", reasoning_mode="default", + context_variant="full", effective_context_window=128000, + composite=0.5, cost_score=0.5, proficiency_score=0.9, + ), + candidates_considered=1, + ), + latency_tolerance="interactive", + selected_model="deepseek-v4-flash", + selected_provider="neuralwatt", + ) + + conn = sqlite3.connect(db_path) + conn.row_factory = sqlite3.Row + row = conn.execute("SELECT profile FROM route_decisions").fetchone() + assert row is not None + assert row["profile"] == "locality" + + # A bare Classification (no RouteResponse) should leave profile at None. + dispatcher.persist_route_decision( + "route", + classification=Classification( + task_category="coding_general", task_tier=2, + required_context_tokens=100, confidence=0.9, + ), + latency_tolerance="interactive", + selected_model="deepseek-v4-flash", + selected_provider="neuralwatt", + ) + rows = conn.execute("SELECT profile FROM route_decisions ORDER BY id").fetchall() + assert len(rows) == 2 + assert rows[1]["profile"] is None + conn.close() + + +def _schema_minus_profile_column() -> str: + """schema.sql with the `profile` column and its comments stripped from + route_decisions, yielding the previous-todo schema for migration tests.""" + lines = [] + inside_route_decisions = False + skip_profile_block = False + for line in SCHEMA_SQL.splitlines(): + stripped = line.strip() + if stripped.startswith("CREATE TABLE IF NOT EXISTS route_decisions"): + inside_route_decisions = True + if inside_route_decisions: + if stripped.startswith("pinch_final_tokens"): + lines.append(" pinch_final_tokens INTEGER") + skip_profile_block = True + continue + if skip_profile_block: + if stripped == ");": + skip_profile_block = False + inside_route_decisions = False + lines.append(line) + continue + continue + lines.append(line) + return "\n".join(lines) + + +def test_ensure_route_decisions_migrates_profile_column(tmp_path): + """A live table without `profile` gets the column once, idempotently.""" + conn = sqlite3.connect(tmp_path / "migrate.db") + conn.executescript(_schema_minus_profile_column()) + assert "profile" not in { + r[1] for r in conn.execute("PRAGMA table_info(route_decisions)") + } + + dispatcher.ensure_route_decisions(conn) + assert "profile" in { + r[1] for r in conn.execute("PRAGMA table_info(route_decisions)") + } + + # Second call must be a no-op and not duplicate the column. + dispatcher.ensure_route_decisions(conn) + cols = [r[1] for r in conn.execute("PRAGMA table_info(route_decisions)")] + assert cols.count("profile") == 1 + + # The new column accepts NULL and string values. + conn.execute( + """ + INSERT INTO route_decisions ( + observed_at, kind, profile + ) VALUES (?, ?, ?) + """, + ("2026-01-01T00:00:00+00:00", "route", "locality"), + ) + conn.execute( + """ + INSERT INTO route_decisions ( + observed_at, kind + ) VALUES (?, ?) + """, + ("2026-01-01T00:00:00+00:00", "route"), + ) + conn.commit() + profiles = [r[0] for r in conn.execute("SELECT profile FROM route_decisions ORDER BY id")] + assert profiles == ["locality", None] + conn.close() + + def test_persist_ensure_on_write_fixes_live_db_missing_table(tmp_path, monkeypatch): """A live router.db without route_decisions gets it on the WRITE path. @@ -549,6 +674,58 @@ def test_route_endpoint_persists_one_row(decision_router): assert "session_dir" not in r.keys() +def test_route_endpoint_returns_named_profile(decision_router): + client, db_path = decision_router + resp = client.post( + "/route", json={"task": "write me a function", "profile": "locality"} + ) + assert resp.status_code == 200 + data = resp.json() + assert data["profile"] == "locality" + + rows = _rows(db_path) + assert len(rows) == 1 + assert rows[0]["profile"] == "locality" + assert rows[0]["kind"] == "route" + + +def test_route_endpoint_unknown_profile_returns_422(decision_router): + client, db_path = decision_router + resp = client.post( + "/route", json={"task": "write me a function", "profile": "nosuchprofile"} + ) + assert resp.status_code == 422 + detail = resp.json()["detail"] + assert "nosuchprofile" in detail + assert "default" in detail + assert "batch" in detail + + rows = _rows(db_path) + assert len(rows) == 0 + + +def test_dispatch_endpoint_returns_named_profile(decision_router): + client, db_path = decision_router + resp = client.post( + "/dispatch", + json={ + "task": "write me a function", + "profile": "batch", + "task_category": "coding_general", + "task_tier": 2, + "required_context_tokens": 100, + }, + ) + assert resp.status_code == 200 + data = resp.json() + assert data["route"]["profile"] == "batch" + + rows = _rows(db_path) + assert len(rows) == 1 + assert rows[0]["profile"] == "batch" + assert rows[0]["kind"] == "dispatch" + + def test_route_endpoint_override_has_classifier_ms_null(decision_router): client, db_path = decision_router resp = client.post( @@ -605,6 +782,7 @@ def test_routed_chat_persists_one_row(decision_router): assert r["selected_provider"] == "neuralwatt" assert r["classification_source"] == "classifier" assert r["session_key"] is not None + assert r["profile"] == "default" # The session key is a hash — never a directory, never content. assert len(r["session_key"]) == 16 assert "/" not in (r["session_key"] or "") @@ -640,6 +818,7 @@ def test_streamed_routed_chat_persists_one_row(decision_router): assert rows[0]["kind"] == "chat" assert rows[0]["selected_model"] == CHEAP assert rows[0]["streamed"] == 1 + assert rows[0]["profile"] == "default" def test_streamed_routed_chat_writes_request_id_back(decision_router): @@ -674,6 +853,7 @@ def test_passthrough_persists_one_row_with_no_nameerror(decision_router): assert r["selected_provider"] == "neuralwatt" assert r["classification_source"] is None assert r["session_key"] is not None + assert r["profile"] is None, "non-routed passthrough must write profile=None" def test_passthrough_records_pinch_columns_when_pruned(decision_router, monkeypatch): @@ -751,6 +931,7 @@ def test_local_vision_success_persists_one_local_row(decision_router, monkeypatc assert r["selected_provider"] == "local" assert r["rejected_reason"] is None assert r["images"] == 1 + assert r["profile"] is None, "non-routed local_vision must write profile=None" def test_no_candidate_422_still_persists_a_rejection_row(decision_router): @@ -770,6 +951,7 @@ def test_no_candidate_422_still_persists_a_rejection_row(decision_router): assert r["selected_model"] is None assert r["rejected_reason"] is not None assert "vision" in r["rejected_reason"] + assert r["profile"] == "default" # --- failure modes: best-effort, config-gated ------------------------------- @@ -1098,7 +1280,7 @@ def test_exploration_disabled_keeps_winner_and_zero_flag(decision_router, monkey def test_exploration_max_tier_excludes_tier_three(decision_router, monkeypatch): - client, db_path = decision_router + _, db_path = decision_router _seed_proficiency_outcomes(db_path, {CHEAP: 50, EXPLORABLE_DEAR: 0}) monkeypatch.setattr(dispatcher.cfg.exploration, "enabled", True) monkeypatch.setattr(dispatcher.cfg.exploration, "epsilon", 1.0) -- 2.49.1 From 0e3df6e61eb9e5ad42d82ee7eea459df9588bac2 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Fri, 4 Sep 2026 09:53:00 -0400 Subject: [PATCH 08/10] feat(admin): add Profile column and filter to decisions page Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus --- admin/frontend/decisions.html | 25 +++++++++++++++++++------ 1 file changed, 19 insertions(+), 6 deletions(-) diff --git a/admin/frontend/decisions.html b/admin/frontend/decisions.html index bbf59a0..016b5aa 100644 --- a/admin/frontend/decisions.html +++ b/admin/frontend/decisions.html @@ -205,6 +205,9 @@ header.navbar{padding-top:2px!important;padding-bottom:2px!important} +