Files
6krrt/docs/data-model.md
adlee-was-taken 92031dbd0b feat(classifier): record the classifier's own attempt on route_decisions
A declined answer left one number behind and it was in a log line, so
confidence_min could not be tuned from data. route_decisions.confidence cannot
say it either: the chat path re-routes through the override branch, which
hard-codes 1.0, and 43,804 of 43,856 live rows hold exactly that.

Measured on 2026-10-04 under local_decision (qwen3.5:4b): 769 of 2,557 turns
(30%) had no fresh classification, against 0% for local_llm and local_encoder.
The cause was recorded in only 4 of them.

- classifier_confidence, classifier_coverage and classifier_reject on
  route_decisions, filled from a ClassifierAttempt carried on Classification.
  Both accepted and declined answers carry one, so the two distributions can be
  compared around the floor. Reason codes are listed in docs/data-model.md.
- ClassifierRejected (a RuntimeError subclass, messages unchanged) replaces the
  plain RuntimeErrors at the six floor-miss sites and the two local_decision
  refusals, so the number and reason travel out of the raise site.
- The chat path captures the classifier's verdict before the re-route and passes
  it to persist_route_decision (attempt_of), like it already does for source.
- Admin decisions page: the source badge's tooltip shows the attempt, and
  session_history / session_stale get an amber badge instead of neutral grey.
  TUI detail popup and the live event carry the same three keys.
- The /metrics degraded-share warning lists the recorded reasons instead of
  claiming the classifier "has been failing", which was wrong for a classifier
  that answers and is declined.
- degraded_warn_threshold must be in (0, 1] and degraded_warn_min at least 1,
  refused at load: a value above 1 could never fire.

The three columns arrive by ALTER and are NULL on every earlier row; metrics
selects them only when present, so the live DB reads NULL until its restart.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
2026-10-04 22:06:57 -04:00

24 KiB
Raw Permalink Blame History

SQLite schema reference. Back to README.

Decision Table Schema (SQLite)

Three data tables plus one observability table, PRAGMA foreign_keys = ON:

models — one row per served model variant

Column Type Notes
model_id TEXT Full catalog id (e.g. glm-5.2-short-fast-flex or qwen2.5-coder-router:14b)
provider TEXT neuralwatt | ollama-local
base_model_id TEXT Model family (e.g. glm-5.2). Proficiency/leaderboard keys here.
eligible_categories TEXT Comma-joined category allowlist; NULL = unrestricted (cloud rows)
display_name TEXT Human-readable name
cost_per_1m_prompt REAL Listed USD per 1M input tokens
cost_per_1m_completion REAL Listed USD per 1M output tokens
cost_per_1m_prompt_cached REAL Cached prefix price (null if no cache discount)
context_window INTEGER Advertised max tokens
effective_context_window INTEGER advertised × safety_factor − reserve
max_output_tokens INTEGER
tier INTEGER 1–3, set by tier.py pass
supports_tools INTEGER Boolean 0/1
supports_json_mode INTEGER
supports_vision INTEGER
supports_reasoning INTEGER "API accepts reasoning param" — NOT a quality signal
reasoning_default_enabled INTEGER The actual tier-bearing signal
latency_class TEXT standard | flex (-flex: discounted async, held during peak)
reasoning_mode TEXT default | reduced (-fast: reasoning capped)
context_variant TEXT full | short (-short: 200K pool bounded budget)
access_level TEXT public | preview | canary
pricing_tbd INTEGER
deprecated INTEGER
availability TEXT active | deprecated | stale | blocked — the blocked value is an operator stop set via the admin override (see admin-portal.md); it is excluded from routing and pinned requests
last_updated TEXT ISO8601

Serving class: Neuralwatt ships ~6 base models as 19 catalog rows. The id suffixes are three orthogonal dimensions (glm-5.2-short-fast-flex), parsed by poller.parse_serving_class into columns. Rows carry identical catalog pricing, so routing would pick between them arbitrarily without these — the latency_tolerance hard filter resolves it.

Local rows: provider='ollama-local' rows come from config.yaml's local_dispatch_models: section and are refreshed by poller.upsert_local_dispatch_models each poll. They start with the three cost columns NULL; once seed_local_dispatch_energy.py has run, those columns hold measured tariff-priced rates and re-polls never overwrite them.

Access gating: 6 of 19 rows are prose-gated ("Private preview (grant-gated)", "(Canary)"). poller.parse_access_level parses them into access_level and routing.allowed_access_levels (default [public]) excludes them, so dispatch won't earn a 403.

proficiency — one row per (model, provider, category)

Column Type Notes
model_id TEXT
provider TEXT
category TEXT See category list below
leaderboard_score REAL 0–1, from external benchmarks
self_eval_score REAL 0–1, from self-eval harness
outcome_score REAL Accumulated client-reported success rate on real traffic (0–1)
outcome_samples INTEGER Number of client-reported succeeded/failed verifications folded in
self_eval_samples INTEGER Evidence count for benchmark blending threshold
blended_score REAL Expected pass rate on real traffic after empirical-Bayes shrinkage
source TEXT outcome_blended | outcome_prior | blended | self_eval | self_eval_thin | leaderboard
inherited_from TEXT Model this row was copied from, NULL if measured directly
last_updated TEXT ISO8601

Category set (9 categories, defined in config/config.yaml):

Category Example Scoring type
coding_general Merge intervals, parse semver, word wrap Code (execution)
coding_refactor Remove repetition, refactor dispatch chain Code (execution)
debugging Fix closure leak, fix binary search, fix regex Code (execution)
reasoning_math Percent trap, rate trap, counting Exact match
tool_use_agentic Right tool / right args / no tool when empty Structural
docs_writing Docstring quality, must-mention gotchas Judge
summarization Root-cause isolation, buried-lede identification Judge
translation Technical register, hedging/informal tone Judge
general_chat Simple explanations, measured pushback Judge

Scoring kinds (4 types, objective wherever the category admits it):

  • code — runs model-generated Python in a subprocess, scores fraction of checks passing
  • exact — normalizes & compares a single answer
  • tool — structural: which tool was called, which args matched
  • judge — a strong model scores against a rubric (prose categories only)

Blending the benchmark signal: Once self_eval_samples ≥ self_eval_min_samples (default 10): blended = 0.3 × leaderboard + 0.7 × self_eval. Before that, falls back to leaderboard alone. If neither exists, the benchmark component is neutral.

From benchmark to expected pass rate: The files in src/proficiency.py convert the benchmark blend into an expected client success rate on real traffic. A category with no client outcome traffic keeps the benchmark score verbatim. A trafficked category with no per-model outcomes inherits a peer-rate prior. A row with its own outcomes gets an empirical-Bayes blend of the prior and its observed rate, with outcome_prior_strength pseudo-observations (default 20) pulling thin data toward the category mean.

Source labels report provenance, not just confidence:

  • outcome_blended — per-model has real outcome evidence
  • outcome_prior — category is trafficked, but this row has no own outcomes
  • self_eval_thin — benchmark measurement below the sample threshold; a caller wanting to exclude it can

Proficiency inheritance: propagate_to_variants copies evaluated scores to equivalent serving variants (same weights, same reasoning setting, same context pool), but never over a row that was measured directly. A -fast row is not equivalent to its -standard sibling.

"Measured directly" is inherited_from IS NULL, not self_eval_samples > 0 — inheritance copies the sample count too, so sample count alone cannot tell an inherited row from a measured one, and using it meant a variant inherited exactly once and then froze forever. ensure_columns() adds the column and backfills provenance on databases that predate it.

energy_observations — per-request telemetry

Column Type Notes
id INTEGER Autoincrement
model_id / provider TEXT
task_category TEXT
prompt_tokens / completion_tokens INTEGER
energy_kwh REAL Attributed billed figure (noisy, 20× within-model)
energy_btu REAL kwh × 3412.14, dashboard value
avg_power_watts / duration_seconds REAL Pre-attribution product, ~1.8× within-model. duration_seconds is the provider's reported serving time — see router_wall_seconds below, which is a different quantity
attribution_ratio REAL Request's share of shared GPU pool (stable quantized: 0.001, 0.25, 0.5, 0.75)
carbon_g_co2eq REAL Reported by provider
grid_carbon_intensity REAL gCO2/kWh at call time
grid_id TEXT e.g. FI
carbon_source TEXT static_fallback (constant, excluded from routing) or live measurement
cost_usd REAL Billed figure, not tokens × list price
allowance_remaining_usd REAL
service_tier TEXT As billed
cached_prompt_tokens INTEGER usage.prompt_tokens_details.cached_tokens. NULL when the provider reported no count — never a substituted 0
cached_tokens_source TEXT Why the column above is what it is: reported (a number arrived, 0 included), details_no_count (a details block with no cached count), no_details (no details block at all). NULL only on rows written before the column
router_wall_seconds REAL Router-observed wall clock for the request, time.monotonic(). Not duration_seconds — see below
router_ttft_seconds REAL Router-observed time to the first output token. Streaming only; NULL on a buffered row means not applicable
observed_at TEXT ISO8601

Router-observed latency is a second quantity, not a backfill of the first. duration_seconds is what the provider says it spent serving; the two router_* columns are what the router measured end to end, which additionally includes connection setup, network transit, queueing ahead of the first token, and router overhead. They are never written into each other.

The reason for the second measurement is coverage. On the live database duration_seconds is present on 31,243 of 31,309 NeuralWatt rows and on 0 of 3,850 OpenRouter rows, and no request-body opt-in will change that: OpenRouter serves generation timing only from its separate /api/v1/generation?id= endpoint, a second HTTP call per request. OpenRouter carries ~73% of routed decisions, so a latency term in the objective needs a number that exists for every provider. The router can always take one.

Spans, which differ by path:

path router_wall_seconds router_ttft_seconds
streaming /v1/chat/completions connection open → last byte forwarded connection open → first delta carrying content or a tool_calls fragment
buffered /v1/chat/completions just before the POST → complete response body NULL, not applicable
POST /dispatch just before the SDK call → response returned NULL, not applicable

Three things follow from those definitions. The mark is re-taken per attempt, so a failover records only the candidate that actually answered. A role-only opening delta does not count as a first token — it is protocol, not answer — so TTFT is measured against output a user could see. And if a streaming client hangs up early the finally still records the span up to abandonment, which understates that request's latency rather than inflating it.

NULL on rows written before the columns existed, and on seed_energy.py and eval_proficiency.py rows, which do not pass the timings.

Attribution noise: Billed energy_kwh = avg_power_watts × duration × attribution_ratio. The attribution term looks like noise up close (8 identical calls varied 20×), but ranks 750× between models while within-model spread is 1.8× — it's a stable per-model property reflecting serving concurrency. Routing scores on the attributed figures with a median over all seed_reference rows, so repeated sweeps accumulate into a median-across-time.

provider_balance_observations — per-provider prepaid pool snapshots

Column Type Notes
id INTEGER Autoincrement
provider TEXT
balance_usd REAL Current prepaid pool balance
total_credits_usd REAL Lifetime credits purchased
total_usage_usd REAL Lifetime usage billed against the pool
observed_at TEXT ISO8601

Polled by poller.py from each provider's balance endpoint (provider.balance_url). Used by quota_accounts() to compute per-provider burn rate, projected runway, and stale-reading alerts.

local_energy_observations — per-call local hardware draw

Column Type Notes
id INTEGER Autoincrement
model_id TEXT Local Ollama tag (e.g. mistral-nemo:12b or qwen2.5-coder-router:14b), not a cloud model
call_type TEXT Not a closed enum. Current values include classify, verify, local_vision, local_dispatch, file_summarization, diff_checking, and seed_local_dispatch.
request_id TEXT Optional; joins to POST /outcome reports the same way energy_observations.request_id does for cloud rows
session_dir TEXT Optional; used for source-less /outcome attribution
session_key TEXT Optional; namespaced conversation key ("c:" + conversation) from X-Router-Conversation, or NULL for a no-header local row
avg_power_watts REAL Averaged over the call (background nvidia-smi sampler)
duration_seconds REAL Wall-clock time for the local call
energy_kwh REAL avg_power_watts × duration_seconds / 3_600_000
cost_usd REAL energy_kwh × tariff_usd_per_kwh; NULL if tariff not configured
carbon_g_co2eq REAL energy_kwh × 1000 × grid_intensity_g_per_kwh; NULL unless grid intensity configured
meter TEXT nvidia_smi (room for RAPL / smart plug later)
observed_at TEXT ISO8601

Separate table, by design. quota_burn() sums energy_kwh over all of energy_observations against the NeuralWatt plan allowance. A separate table makes it structurally impossible for local electricity to leak into that number. local_energy_summary() in metrics.py queries this table exclusively, over the same 30-day window. /metrics surfaces it as a top-level "local_energy" key, and the admin dashboard has a distinct "Local compute" card. The meter is off by default (local_energy: in config); config load refuses enabled: true without a tariff_usd_per_kwh. Remote-Ollama safety: if any Ollama base URL is non-loopback, metering skips with a log warning. ensure_local_energy_table() in src/local_energy.py also creates the table and index idempotently for live databases predating this schema.

verifications — response quality observations

Column Type Notes
id INTEGER Autoincrement
model_id / provider TEXT
task_category TEXT Optional (prose answers may lack a category)
kind TEXT structural | local_llm | client_outcome
verdict TEXT ok | truncated | malformed | unverifiable | succeeded | failed
detail TEXT Human-readable reason
completion_tokens INTEGER Wasted answer cost, for payoff sum
observed_at TEXT ISO8601
applied_at TEXT Set by feedback.py when folded into proficiency
model_attributable INTEGER 1 = model's fault; 0 = client caused (e.g. tight cap)

Verifications drive the feedback loop: feedback.py reads unanswered failures, applies a 0.0 sample per failure to proficiency, and marks them applied_at for idempotency. unverifiable is recorded but not treated as a failure — it means the checker had nothing to say, not that the model failed.

route_decisions — routing observability

Column Type Notes
id INTEGER Autoincrement
observed_at TEXT ISO8601, UTC
kind TEXT route | dispatch | chat | passthrough | local_vision | local_dispatch_fallback
task_category TEXT
task_tier INTEGER 1–3
required_context_tokens INTEGER
confidence REAL Confidence of the classification that ROUTED the request. On a chat row this is almost always 1.0, because the chat path re-routes through the override branch, which hard-codes it. Use classifier_confidence for what the classifier said
classifier_ms INTEGER Classification latency
classification_source TEXT classifier | override | fallback | cached | session_stale | session_history | classifier_cloud
latency_tolerance TEXT interactive | batch
candidates_considered INTEGER How many survived hard filters
selected_model TEXT Null when no model was selected
selected_provider TEXT neuralwatt, local (vision fallback), or ollama-local (local dispatch)
runner_up_models TEXT JSON array of up to 3 runner-up candidates
est_cost_usd REAL Estimated cost of the selected model
est_proficiency REAL Estimated proficiency for the task category
rejected_reason TEXT Active filters when nothing was selected
session_key TEXT "c:" + conversation id from X-Router-Conversation when the client sent one, else a hashed content fingerprint of the conversation; NULL on a no-header local row
agent TEXT Optional; agent name from X-Router-Agent (namespaced form, see api.md)
parent_key TEXT Optional; "c:" + parent X-Router-Parent conversation key, when the request named one
tools INTEGER 0/1 — request carried a tools array
images INTEGER 0/1 — request carried image parts
json_mode INTEGER 0/1 — request required JSON mode
streamed INTEGER 0/1 — response was streamed
pinch_original_tokens INTEGER Estimated tokens before pruning (includes extra_fixed_tokens)
pinch_final_tokens INTEGER Estimated tokens after pruning
request_id TEXT Provider request id; joins to energy_observations.request_id
exploration INTEGER 0/1 — selected as the least-evidenced candidate for exploration
prefix_divergence_index INTEGER First message position whose bytes differ from the previous turn in this session
prefix_tokens_after_divergence INTEGER Estimated tokens at or after that position in THIS turn
prefix_prev_message_count INTEGER The previous turn's message count
classifier_confidence REAL The primary classifier's confidence in its answer, accepted or not. NULL when it produced none (a timeout) or made no attempt (override, cached)
classifier_coverage REAL local_decision only: the total option mass behind that answer
classifier_reject TEXT NULL when the answer was used; otherwise why it was not (codes below)

What the classifier did, and why its answer was not used

The three classifier_* columns record the primary classifier's own attempt, read before the chat path re-routes on measured context. They exist because a declined answer used to leave one number behind and it was in a log line: classification_source = session_history says a borrowed label routed the turn, not whether the classifier was a hair under its floor or nowhere near it, and confidence could not say so (see its row above).

classifier_reject is NULL when the answer was used, and otherwise one of:

code meaning numbers recorded
below_confidence_min answered, but under classifier.*.confidence_min confidence (and coverage for local_decision)
below_coverage_min local_decision: total option mass under coverage_min coverage, and the would-be confidence
no_logprobs local_decision: Ollama returned none none
timeout the endpoint did not answer in time none
transport_error the endpoint was unreachable or errored none
parse_error the reply could not be parsed none
primary_failed any other failure of the primary none
skipped_gaming_mode, skipped_backoff the local classifier was never asked none

classification_source then says what routed the turn instead: session_history (this session's last real label), fallback (the static guess) or classifier_cloud. Tune confidence_min from the distribution: for local_decision rows, compare classifier_confidence where classifier_reject is NULL against where it is below_confidence_min.

Rows from before these columns existed hold NULL in all three, which is the honest value; the numbers were never stored, so there is nothing to backfill. The admin decisions page shows them as the tooltip on the source badge, the TUI detail popup shows all three, and the /metrics degraded-share warning lists the reasons it saw.

The three prefix_* columns are the prefix-stability probe (pinch.prefix_probe, on by default). The provider bills the longest byte-identical prefix of a prompt at the cached rate, so a divergence early in the payload re-bills everything after it; these say how much of the previous turn's cache this turn kept. Read them together: prefix_divergence_index == prefix_prev_message_count is a healthy append, and anything lower is rewritten history — prefix_tokens_after_divergence is then what it cost.

They are NULL together when there was no previous turn to compare against (first turn of a session, or first after a restart), which is deliberately distinct from a measured 0 — a zero would read as total cache loss at message 0. Hashes only, and not even those: the per-message digests that produce these numbers live in process memory for exactly one turn and never reach the database, so nothing here is reversible to any message. Rows written before the columns existed are NULL and nothing may be inferred for them; the payloads were never stored, by design.

pinch_original_tokens and pinch_final_tokens capture context-pruning outcomes: tokens_saved = original - final, and pruned = original > final. Both are NULL when pinch is disabled or the request predates the columns. request_id is written back after the provider returns a completion id; it is NULL on /route (no upstream call) until backfilled. exploration is 1 when exploration.choose replaced the ranked winner with a least-evidenced candidate.

Rows with kind='local_dispatch_fallback' record a degraded answer after the cloud refused or was exhausted — they are NOT evidence that local was competitive on merit (do not feed them into proficiency analysis). They carry selected_provider='ollama-local'.

route_decisions stores one row per routing decision so "how is routing performing" is answerable: which model was picked, for what category/tier, how long classification took, and — when nothing was selected — which hard filter shut it out. It is an observability table: nothing in routing reads it.

It keys a conversation by session_key -- "c:" + conversation id when the client sent X-Router-Conversation, else a hashed content fingerprint; session_dir, prompts, and answers are deliberately excluded. A test enforces that the write path does not store prompt or answer text. The write is gated by logging.log_route_decisions and is best-effort: a failed write is logged at warning and swallowed so monitoring cannot slow or fail a request.

Watchdog Schema (SQLite)

Four tables written by watchdog.py via src/watchdog_store.py (the code-side inline-create mirrors this schema for live databases). See watchdog.md for how the loop reads them.

watchdog_ticks — one row per detection run

Column Type Notes
id INTEGER Autoincrement
ticked_at TEXT ISO8601 run time
sessions_seen INTEGER Count of sessions examined this run, default 0
outcome TEXT

watchdog_verdicts — one row per session per tick

Column Type Notes
id INTEGER Autoincrement
tick_id INTEGER FK → watchdog_ticks.id, NOT NULL
session_id TEXT NOT NULL
session_root TEXT
agent TEXT
model_id TEXT
provider TEXT
flagged INTEGER 0/1, default 0
dup REAL Duplicate-fraction signal
top INTEGER
top_what TEXT
landed INTEGER
slow INTEGER
coverage REAL
calls_since_landed INTEGER
cost_since_landed_usd REAL
llm_second_opinion TEXT
created_at TEXT NOT NULL

Four indexes (config/schema.sql L482-485):

Index On
idx_watchdog_verdicts_model_id watchdog_verdicts(model_id)
idx_watchdog_verdicts_created_at watchdog_verdicts(created_at)
idx_watchdog_verdicts_flagged watchdog_verdicts(flagged)
idx_watchdog_verdicts_session_root watchdog_verdicts(session_root)

watchdog_alerts — one row per open alert

Column Type Notes
dedup_key TEXT PRIMARY KEY; e.g. opencode-loop:<root session id>
state TEXT NOT NULL
severity TEXT NOT NULL
flagged_ticks INTEGER Default 0
opened_at TEXT NOT NULL
last_fired_at TEXT NOT NULL
resolved_at TEXT

watchdog_channel_settings — per-channel notification config

Column Type Notes
id INTEGER Autoincrement
channel_name TEXT NOT NULL, UNIQUE
enabled INTEGER Default 1
min_severity TEXT Default warning