Enable circuit_breaker, pinch, and pinch.relevance by default (with budget-gate) #10

Merged
alee merged 4 commits from neuralwatt-router-service into main 2026-08-30 16:57:20 +00:00
9 changed files with 102 additions and 31 deletions

View File

@@ -35,7 +35,7 @@ code.
| `events.py` | In-memory decision-event broker | Pure stdlib (`queue`, `collections.deque`). Thread-safe. No Textual import. The SSE endpoint lives in `dispatcher.py` and calls `events.subscribe()`/`publish_decision()`. | | `events.py` | In-memory decision-event broker | Pure stdlib (`queue`, `collections.deque`). Thread-safe. No Textual import. The SSE endpoint lives in `dispatcher.py` and calls `events.subscribe()`/`publish_decision()`. |
| `config.py` | Pydantic models + YAML loader | All config models inherit `StrictModel` (`extra="forbid"`). Unknown keys fail at load. | | `config.py` | Pydantic models + YAML loader | All config models inherit `StrictModel` (`extra="forbid"`). Unknown keys fail at load. |
| `capabilities.py` | Request-side capability detection (`tools`, `images`, `json_mode`, `reasoning`) | Reads from the OpenAI-format request body, not from the classifier. | | `capabilities.py` | Request-side capability detection (`tools`, `images`, `json_mode`, `reasoning`) | Reads from the OpenAI-format request body, not from the classifier. |
| `context_prune.py` | Relevance-based context pruning (`pinch`) | Ships **disabled** (`pinch.enabled: false`). Imports `PinchConfig` from `config.py` (no circular import: `config.py` doesn't import `context_prune`). | | `context_prune.py` | Relevance-based context pruning (`pinch`) | Ships **enabled** by default (`pinch.enabled: true`; `pinch.relevance.enabled: true`; budget-gated: the embed only fires when the conversation exceeds `budget_tokens`). Imports `PinchConfig` from `config.py` (no circular import: `config.py` doesn't import `context_prune`). |
| `logs.py` | Structured logging (logfmt, journald, ContextVar trace ids) | `logs.bind()` exists for StreamingResponse generators that lose the ContextVar. | | `logs.py` | Structured logging (logfmt, journald, ContextVar trace ids) | `logs.bind()` exists for StreamingResponse generators that lose the ContextVar. |
### Pure scoring/routing modules (no I/O) ### Pure scoring/routing modules (no I/O)
@@ -161,6 +161,47 @@ PYTHONPATH=src python -m router_cli "Refactor this Django view"
PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
``` ```
## Gitea / `tea` CLI (not `gh`)
This repo is hosted on **self-hosted Gitea** (`git.adlee.work/alee/6krrt`).
The GitHub CLI `gh` is **not installed** and will fail — use `tea` instead.
- **Login:** `tea login` as `alee`. The `alee` login is NOT the default tea
login — always pass `--remote origin` so tea resolves the right context.
- **Remote:** `origin` resolves to `git.adlee.work/alee/6krrt`.
- **List open PRs:** `tea pulls --remote origin`
- **PR metadata (JSON):**
```
tea pulls <PR_INDEX> --remote origin --fields index,state,draft,title,mergeable,base,head,body -o json
```
Use the `base` and `head` from this output for local diff commands.
- **PR creation / update / maintenance:** use the `tea pr` family (e.g. `tea pr create --base main --title ... --description ...`). Confirm against `tea pr --help` for the current flag set.
- **Post a comment to a PR** (uses the `tea api` proxy to Gitea's REST API):
```
tea api --remote origin "/repos/{owner}/{repo}/issues/<PR_INDEX>/comments" -F body=@-
```
- **CRITICAL:** use `-F` (typed field), **not** `-f`. `-f key=@file` writes the literal string `@file`; only `-F` reads stdin or the contents of a file.
- The endpoint **must** include the `/repos/` prefix — omitting it returns a 404.
- `tea api` already prints JSON to stdout — do **not** pass `-o json` on top of that (you would write the body to a file literally named `json`).
- Fallback if stdin is awkward: `-F body=@/tmp/review-body.md`.
- **Getting a PR diff:** prefer local `git diff origin/<base>...origin/<head>`
(base and head from the PR JSON). In tea v0.14.1 the `diff`/`patch` fields
return empty even when requested — do not rely on them.
- **Gitea source links** (not GitHub `blob` links):
`https://git.adlee.work/alee/6krrt/src/commit/<FULL-SHA>/<path>#L<start>-L<end>`
(segment is `src/commit`, not `blob`).
- **Code-review plugin:** this repo ships a project-local fork of the code-review
plugin (`code-review-tea`) that already wires these `tea` calls. Prefer
invoking it (e.g. `/code-review-tea:code-review`) over hand-rolling.
## What's NOT built yet (open items) ## What's NOT built yet (open items)
1. **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. 1. **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty.

View File

@@ -166,12 +166,11 @@ pinch:
# they carry the bulk of a long agent session's tokens and are least needed # they carry the bulk of a long agent session's tokens and are least needed
# in full by the time the next turn is answered. # in full by the time the next turn is answered.
# #
# It does NOT touch the classifier's input, and it defaults off — enable it # On by default. Tool results can only be dropped safely because tool
# only if long sessions are shipping more prompt tokens than you want to pay # outputs are idempotent enough for a placeholder; a wrong guess here loses
# for. Tool results can only be dropped safely because tool outputs are # context, so start conservative (large budget, small reduction) and watch
# idempotent enough for a placeholder; a wrong guess here loses context, so # route_decisions / pinch stats on real traffic before widening it.
# start conservative (large budget, small reduction). enabled: true
enabled: false
budget_tokens: 50000 budget_tokens: 50000
keep_last_turns: 4 keep_last_turns: 4
max_summarize_chars: 4000 max_summarize_chars: 4000
@@ -180,7 +179,7 @@ pinch:
# project — and specifically requires pinch.enabled too, since this has no # project — and specifically requires pinch.enabled too, since this has no
# effect otherwise. Ship it, watch route_decisions / pinch stats on real # effect otherwise. Ship it, watch route_decisions / pinch stats on real
# traffic, then decide the default. # traffic, then decide the default.
enabled: false enabled: true
# An EMBEDDING model, not a chat model — this must not point at # An EMBEDDING model, not a chat model — this must not point at
# classifier.model or verification.model. Pull one on the same Ollama: # classifier.model or verification.model. Pull one on the same Ollama:
# ollama pull nomic-embed-text # ollama pull nomic-embed-text
@@ -207,13 +206,11 @@ session_cache:
staleness_minutes: 20 staleness_minutes: 20
circuit_breaker: circuit_breaker:
# Passive availability circuit breaker. Off by default, matching every other # Passive availability circuit breaker, on by default. When enabled, a model
# new-and-unproven knob in this project. When enabled, a model that returns # that returns 5xx is temporarily excluded from routing with exponential
# 5xx is temporarily excluded from routing with exponential backoff; recovery # backoff; recovery is passive (the next real request becomes the probe once
# is passive (the next real request becomes the probe once the cooldown # the cooldown passes). It has a low-risk failure mode even when wrong.
# passes). Unlike most new knobs this one has a low-risk failure mode even enabled: true
# when wrong, so it's a reasonable candidate to flip on sooner.
enabled: false
initial_cooldown_seconds: 30 initial_cooldown_seconds: 30
max_cooldown_seconds: 600 max_cooldown_seconds: 600
backoff_multiplier: 2.0 backoff_multiplier: 2.0

View File

@@ -306,12 +306,12 @@ class PinchRelevanceConfig(StrictModel):
When enabled (and ``pinch.enabled`` is also true), the dispatcher embeds When enabled (and ``pinch.enabled`` is also true), the dispatcher embeds
the current-turn query with the old tool-result candidates and trims the the current-turn query with the old tool-result candidates and trims the
least relevant first, so a relevant-but-old result survives. Off by least relevant first, so a relevant-but-old result survives. On by
default; any failure reverts to uniform trimming. This must point at an default; any failure reverts to uniform trimming. This must point at an
EMBEDDING model, never ``classifier.model`` or ``verification.model``. EMBEDDING model, never ``classifier.model`` or ``verification.model``.
""" """
enabled: bool = False enabled: bool = True
# OpenAI-compatible embeddings endpoint on the same local Ollama. # OpenAI-compatible embeddings endpoint on the same local Ollama.
model: str = "nomic-embed-text" model: str = "nomic-embed-text"
base_url: str = "http://localhost:11434/v1" base_url: str = "http://localhost:11434/v1"
@@ -343,7 +343,7 @@ class PinchConfig(StrictModel):
are summarized or dropped (they carry the bulk of a long session's tokens). are summarized or dropped (they carry the bulk of a long session's tokens).
""" """
enabled: bool = False enabled: bool = True
budget_tokens: int = 50000 budget_tokens: int = 50000
# How many recent user turns (plus their assistant replies and tool results) # How many recent user turns (plus their assistant replies and tool results)
# are protected from pruning. # are protected from pruning.
@@ -406,10 +406,10 @@ class CircuitBreakerConfig(StrictModel):
When enabled, a model that returns 5xx is temporarily skipped by routing When enabled, a model that returns 5xx is temporarily skipped by routing
(with exponential backoff). Recovery is passive: a real request that would (with exponential backoff). Recovery is passive: a real request that would
have picked it becomes the probe once the cooldown passes. Off by default. have picked it becomes the probe once the cooldown passes. On by default.
""" """
enabled: bool = False enabled: bool = True
initial_cooldown_seconds: int = 30 initial_cooldown_seconds: int = 30
max_cooldown_seconds: int = 600 max_cooldown_seconds: int = 600
backoff_multiplier: float = 2.0 backoff_multiplier: float = 2.0

View File

@@ -59,6 +59,7 @@ from capabilities import detect_capabilities, iter_image_url_values
from config import FlexPreference, RouterConfig, load_config from config import FlexPreference, RouterConfig, load_config
from context_prune import ( from context_prune import (
extract_text, extract_text,
estimate_tokens,
order_by_relevance, order_by_relevance,
prune_context, prune_context,
trim_candidates, trim_candidates,
@@ -1818,7 +1819,7 @@ def _embed_for_relevance(
return order_by_relevance(embeddings[0], embeddings[1:]) return order_by_relevance(embeddings[0], embeddings[1:])
def _relevance_order_for(messages: list[dict], cfg) -> Optional[list[int]]: def _relevance_order_for(messages: list[dict], cfg, budget_tokens: int) -> Optional[list[int]]:
"""Compute the relevance_order for prune_context, or None (uniform). """Compute the relevance_order for prune_context, or None (uniform).
When pinch.relevance is off, or candidate count is below min_candidates, When pinch.relevance is off, or candidate count is below min_candidates,
@@ -1831,6 +1832,11 @@ def _relevance_order_for(messages: list[dict], cfg) -> Optional[list[int]]:
candidates, _, _ = trim_candidates(messages, cfg.pinch.keep_last_turns) candidates, _, _ = trim_candidates(messages, cfg.pinch.keep_last_turns)
if len(candidates) < rel.min_candidates: if len(candidates) < rel.min_candidates:
return None return None
# Skip the embed when the conversation is under budget — prune_context would
# be a no-op, so ranking is pointless.
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages)
if orig_tokens <= budget_tokens:
return None
query = _last_user_text(messages) query = _last_user_text(messages)
candidate_texts = [_text_only(messages[i]) for i in candidates] candidate_texts = [_text_only(messages[i]) for i in candidates]
return _embed_for_relevance(query, candidate_texts, cfg) return _embed_for_relevance(query, candidate_texts, cfg)
@@ -2275,7 +2281,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks):
budget_tokens=cfg.pinch.budget_tokens, budget_tokens=cfg.pinch.budget_tokens,
keep_last_turns=cfg.pinch.keep_last_turns, keep_last_turns=cfg.pinch.keep_last_turns,
max_summarize_chars=cfg.pinch.max_summarize_chars, max_summarize_chars=cfg.pinch.max_summarize_chars,
relevance_order=_relevance_order_for(messages, cfg), relevance_order=_relevance_order_for(messages, cfg, budget_tokens=cfg.pinch.budget_tokens),
) )
logs.debug( logs.debug(
"pinch", "pinch",
@@ -2507,7 +2513,7 @@ def chat_completions(body: dict[str, Any], background: BackgroundTasks):
budget_tokens=cfg.pinch.budget_tokens, budget_tokens=cfg.pinch.budget_tokens,
keep_last_turns=cfg.pinch.keep_last_turns, keep_last_turns=cfg.pinch.keep_last_turns,
max_summarize_chars=cfg.pinch.max_summarize_chars, max_summarize_chars=cfg.pinch.max_summarize_chars,
relevance_order=_relevance_order_for(messages, cfg), relevance_order=_relevance_order_for(messages, cfg, budget_tokens=cfg.pinch.budget_tokens),
) )
logs.debug( logs.debug(
"pinch", "pinch",

View File

@@ -20,6 +20,7 @@ import threading
from pathlib import Path from pathlib import Path
import pytest import pytest
import yaml
from fastapi import FastAPI from fastapi import FastAPI
from starlette.testclient import TestClient from starlette.testclient import TestClient
@@ -138,7 +139,8 @@ def test_config_POST_creates_backup_before_write(client, tmp_path):
# The backup captured the sentinel comment and the PRE-write value. # The backup captured the sentinel comment and the PRE-write value.
backup_text = backups[0].read_text() backup_text = backups[0].read_text()
assert _SENTINEL in backup_text assert _SENTINEL in backup_text
assert re.search(r"^\s*enabled:\s*false\s*$", backup_text, re.MULTILINE) is not None backup_yaml = yaml.safe_load(backup_text)
assert backup_yaml["circuit_breaker"]["enabled"] is True
def test_config_concurrent_writes_are_atomic_no_zero_byte_backups(tmp_path): def test_config_concurrent_writes_are_atomic_no_zero_byte_backups(tmp_path):

View File

@@ -11,6 +11,8 @@ from __future__ import annotations
import os import os
import sqlite3 import sqlite3
import subprocess
import sys
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
from pathlib import Path from pathlib import Path

View File

@@ -104,6 +104,9 @@ def router(tmp_path, monkeypatch):
# The local vision fallback is off unless a test opts in; without this it # The local vision fallback is off unless a test opts in; without this it
# would fire for every image request and try to reach localhost. # would fire for every image request and try to reach localhost.
monkeypatch.setattr(dispatcher.cfg.local_vision, "enabled", False) monkeypatch.setattr(dispatcher.cfg.local_vision, "enabled", False)
# Isolate the shared module-level session cache: a real cache keyed on the
# same fingerprint leaks a classification from one test into the next.
monkeypatch.setattr(dispatcher.cfg.session_cache, "enabled", False)
monkeypatch.setenv("NEURALWATT_API_KEY", "test-key") monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
calls = [] calls = []

View File

@@ -237,7 +237,7 @@ def test_nonpositive_local_timeout_is_rejected(raw):
def test_the_shipped_config_loads_the_pinch_section(raw): def test_the_shipped_config_loads_the_pinch_section(raw):
loaded = RouterConfig(**raw) loaded = RouterConfig(**raw)
assert loaded.pinch.enabled is False assert loaded.pinch.enabled is True
assert loaded.pinch.budget_tokens > 0 assert loaded.pinch.budget_tokens > 0
assert loaded.pinch.keep_last_turns > 0 assert loaded.pinch.keep_last_turns > 0
@@ -246,7 +246,7 @@ def test_pinch_defaults_when_absent(raw):
cfg = copy.deepcopy(raw) cfg = copy.deepcopy(raw)
cfg.pop("pinch") cfg.pop("pinch")
loaded = RouterConfig(**cfg) loaded = RouterConfig(**cfg)
assert loaded.pinch.enabled is False assert loaded.pinch.enabled is True
assert loaded.pinch.budget_tokens == 50000 assert loaded.pinch.budget_tokens == 50000
assert loaded.pinch.keep_last_turns == 4 assert loaded.pinch.keep_last_turns == 4
@@ -294,7 +294,7 @@ def test_pinch_max_summarize_chars_accepts_default(raw):
def test_pinch_relevance_defaults_load(raw): def test_pinch_relevance_defaults_load(raw):
loaded = RouterConfig(**raw) loaded = RouterConfig(**raw)
assert loaded.pinch.relevance.enabled is False assert loaded.pinch.relevance.enabled is True
assert loaded.pinch.relevance.model == "nomic-embed-text" assert loaded.pinch.relevance.model == "nomic-embed-text"
assert loaded.pinch.relevance.min_candidates == 2 assert loaded.pinch.relevance.min_candidates == 2
@@ -303,7 +303,7 @@ def test_pinch_relevance_defaults_when_pinch_absent(raw):
cfg = copy.deepcopy(raw) cfg = copy.deepcopy(raw)
cfg.pop("pinch") cfg.pop("pinch")
loaded = RouterConfig(**cfg) loaded = RouterConfig(**cfg)
assert loaded.pinch.relevance.enabled is False assert loaded.pinch.relevance.enabled is True
assert loaded.pinch.relevance.timeout_seconds == 10 assert loaded.pinch.relevance.timeout_seconds == 10
@@ -323,7 +323,7 @@ def test_nonpositive_pinch_relevance_min_candidates_is_rejected(raw):
def test_circuit_breaker_defaults_load(raw): def test_circuit_breaker_defaults_load(raw):
loaded = RouterConfig(**raw) loaded = RouterConfig(**raw)
assert loaded.circuit_breaker.enabled is False assert loaded.circuit_breaker.enabled is True
assert loaded.circuit_breaker.initial_cooldown_seconds == 30 assert loaded.circuit_breaker.initial_cooldown_seconds == 30
assert loaded.circuit_breaker.max_cooldown_seconds == 600 assert loaded.circuit_breaker.max_cooldown_seconds == 600
assert loaded.circuit_breaker.backoff_multiplier == 2.0 assert loaded.circuit_breaker.backoff_multiplier == 2.0

View File

@@ -120,7 +120,7 @@ def test_embed_for_relevance_request_exception_returns_none(monkeypatch):
def test_relevance_order_for_returns_none_when_disabled(): def test_relevance_order_for_returns_none_when_disabled():
cfg = _cfg_with_relevance(enabled=False) cfg = _cfg_with_relevance(enabled=False)
assert dispatcher._relevance_order_for([], cfg) is None assert dispatcher._relevance_order_for([], cfg, budget_tokens=100) is None
def test_relevance_order_for_below_min_candidates_skips_embedding(monkeypatch): def test_relevance_order_for_below_min_candidates_skips_embedding(monkeypatch):
@@ -138,7 +138,27 @@ def test_relevance_order_for_below_min_candidates_skips_embedding(monkeypatch):
{"role": "tool", "name": "read", "content": "x" * 6000}, {"role": "tool", "name": "read", "content": "x" * 6000},
{"role": "assistant", "content": "a"}, {"role": "assistant", "content": "a"},
] ]
assert dispatcher._relevance_order_for(messages, cfg) is None assert dispatcher._relevance_order_for(messages, cfg, budget_tokens=100) is None
assert called["n"] == 0
def test_relevance_order_for_skips_embedding_when_under_budget(monkeypatch):
called = {"n": 0}
def fake_post(*a, **k):
called["n"] += 1
return _FakeResp(status_code=200, json_payload={"data": None})
monkeypatch.setattr(dispatcher.requests, "post", fake_post)
cfg = _cfg_with_relevance(enabled=True, min_candidates=2)
# Two short tool-results -> token estimate under budget -> no embed call.
messages = [
{"role": "user", "content": "q"},
{"role": "tool", "name": "read", "content": "short"},
{"role": "tool", "name": "read", "content": "also short"},
{"role": "assistant", "content": "a"},
]
assert dispatcher._relevance_order_for(messages, cfg, budget_tokens=50000) is None
assert called["n"] == 0 assert called["n"] == 0