Two small, contained specs that can run as one wave.
admin-profile-writes-to-overlay: profile CRUD still writes config.yaml
while allowlisted scalars write the overlay. That split is chronological
accident -- profile CRUD shipped in PR #25, the overlay decision came
after. Redirect create/update to the overlay, scope delete to
overlay-defined profiles, and classify base-defined profiles as read-only
alongside built-ins, which sidesteps the fact that a deep merge cannot
express "remove" without tombstones.
Two defects found while specifying it, both present today. An
overlay-defined profile is LIVE in routing but invisible to the portal --
confirmed live: load_config sees it, _persisted_profiles() does not,
because the latter reads the base file only. And delete reads the merged
store for its default-profile guard but writes the base file, so an
overlay-defined profile is found and then 404s on the delete.
provider-literal-cleanup: the 20 hardcoded neuralwatt literals are wrong
whichever provider lands second, so they need not wait on the parked
multi-provider plan. Three are latent defects: _model_exists fails OPEN
and silently strips a real vendor/model id (the native format for
OpenRouter and most aggregators), catalog staleness is blind to a second
provider's frozen catalog, and leaderboard priors are pinned to one
provider -- which would decide the still-open proficiency-sharing question
by accident, so the spec explicitly refuses to resolve it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Provider selection is reopened -- Z.ai was recommended on the strength of
the GLM catalog overlap, which was an argument for the *criterion*, not a
confirmation that Z.ai's API or terms fit. Downgrade the recommendation to
a shortlist and keep the criterion for whatever candidate comes next.
Also corrects a real error. The draft attributed dispatcher.py:2879 to
_check_pinned_capabilities and called it a spurious-422 risk. That function
already takes provider as a parameter and is not a coupling site. 2879 is
_model_exists, and it fails open rather than closed: it decides whether a
`provider/model` string is an opencode alias to strip or a real id, so a
second-provider id resolves to False, gets stripped, and dispatches to
whatever the remainder matches -- silently, on a different provider, at a
different price. `vendor/model` is the native id format for OpenRouter and
most aggregators, so this is the default case, not an edge one.
Records that the 20 hardcoded neuralwatt literals are provider-agnostic
work that need not wait on the provider decision. Re-verified today: still
20 refs across the same six files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Closes the last gap from docs/incidents.md #5. The hourly local snapshots
survive `git clean -fdx` because they live outside the repo, but they do not
survive the disk. This pushes the irreplaceable-and-small state to a private
Gitea repo: .omo plans and evidence ledger, tuned systemd units, opencode
plugins, this project's Claude memory, config/config.local.yaml, and a DAILY
gzipped router.db.
The database is committed daily rather than hourly on purpose: it is binary and
~5MB gzipped, so git cannot delta it. Hourly would grow the repo ~120MB/day
instead of ~5MB. SYNC_DB=0 turns it off entirely.
.env is deliberately NOT synced. There is no usable secret key on this host to
encrypt it to (public keys only), so it would sit in git history in plaintext,
and history is forever even in a private repo. An API key is replaceable by
regenerating it from the provider; 22,821 energy observations are not. It stays
in the local backups only, and the user confirmed that trade.
Verified by fresh clone: 35 plan artifacts, 127 evidence files, 9 systemd
units, 2 opencode plugins, 11 memory files, the tariff, and a 5.1MB db
snapshot. Checked the actual 67-character key VALUE appears in zero off-site
files -- an earlier check grepped for the variable NAME and for 'sk-', which
matches 'task-', and produced 50 false positives. Grep for the secret, not for
its label.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Direct response to docs/incidents.md #5, where `git clean -fdx` truncated
router.db to 0 bytes and deleted .env, .venv and config/config.local.yaml.
Recovery was luck -- a QA copy happened to exist in /tmp from 25 seconds
earlier. There was no backup policy at all.
llm-router-backup.sh + .service + .timer: hourly, keeps 24. Uses sqlite3
.backup rather than cp, because copying a live database with an open writer can
capture a torn page set that passes a size check and fails integrity_check. It
verifies the new snapshot with integrity_check BEFORE rotating, so a failing run
never leaves fewer copies than it started with.
workstation-backup.sh: tarballs what git does not have -- router.db, .env,
config/config.local.yaml, .omo/ (plans and evidence ledger), tuned systemd
units, opencode plugins, and this project's Claude memory -- plus a RESTORE.md
with the clone -> venv -> restore sequence. Excludes .venv and node_modules
(rebuildable from pinned requirements) and Ollama models (~30GB, re-pullable,
recipes in docs/local-models.md).
Backups write OUTSIDE the repository by design. A backup kept inside it, even
gitignored, would have been destroyed by the same command that caused the
incident.
Test-restored before committing: 22,821 observations with integrity=ok, API key
present, tariff intact, 7 systemd units, 2 opencode plugins, 11 memory files.
That test caught a second loss nobody had noticed -- .omo/plans had also been
destroyed by the same clean, taking 35 plan artifacts and 127 evidence files,
since .omo/ is gitignored too. Recovered from the same QA copy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
First incident in this project with real data loss, and the first where
recovery depended on luck rather than design.
An agent ran `git clean -fdx`. The -x flag removes IGNORED files as well as
untracked ones, and everything this deployment needs to run is ignored by
design: router.db (truncated to 0 bytes -- 22,776 energy observations, 17,321
route decisions, 148 proficiency rows), .env (the NeuralWatt API key), .venv
(the virtualenv systemd's ExecStart runs from), and config/config.local.yaml
(the operator's tariff). Router returned 500 with "no such table:
energy_observations" while systemd reported active.
`git clean -fd` is a reasonable thing for an agent to run. Adding -x turns it
from "discard my scratch files" into "delete the deployment", and nothing in
the repo warned about that.
Recovery was luck: the plan running at the time had made a QA copy in /tmp 25
seconds before the wipe, and that copy happened to include .env. Restore was
effectively lossless. A different plan and the entire measurement history would
be gone.
Records the recovery runbook, and the gap it exposes: router.db has no backup
policy despite holding data that either costs money to rebuild (proficiency,
via eval runs) or cannot be rebuilt at all (historical observations).
Also notes what limited the damage -- committing the in-flight plan's 11 files
minutes earlier, and parking the tariff outside the repo -- both of which were
reactions to that same file being clobbered seven times the same day.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
WIP checkpoint committed by Claude while Atlas was still in its final
verification wave. Committed early deliberately: 11 files were sitting
uncommitted with two of them untracked, and untracked files in this tree have
been destroyed twice today by agent git cleanup -- the user's
multi-provider-support plan and a config.local.yaml holding their tariff.
Protecting the work cost nothing; losing it would have cost a whole plan.
Deployment-specific values previously lived as an UNCOMMITTED modification to
the tracked config/config.yaml. That arrangement failed seven times in one
session: three agent checkout/stash/restore calls, two `git commit -am` sweeps
that each needed a history rewrite, one ordinary branch switch, and one
cleanup during this very plan. Once a value is correctly absent from git,
every checkout, switch, pull and rebase wipes it -- that is the intended fix
behaving as designed, which is what makes the arrangement itself the bug.
load_config now deep-merges an optional config/config.local.yaml over the base
before validation, so every consumer inherits it (poller, feedback,
eval_proficiency, seed_local_dispatch_energy, admin, dispatcher). Mappings
deep-merge; lists replace wholesale; the MERGED result is validated once so
extra="forbid" still catches an overlay typo.
The admin portal now writes to the overlay and never to config/config.yaml.
An operator changing a knob in a loopback-only portal is making a local
operational decision, not a project decision -- someone changing a project
default edits config.yaml and commits it through git. This also removes the
shadowing trap by construction rather than guarding against it.
config/config.yaml stops being dirty in normal operation, which removes the
condition behind every clobbering above.
NOTE: final verification wave had not reported when this was committed. The
suite was green at 1155 before it started; Atlas may amend or add commits on
top.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Counts the coupling instead of asserting it: grep -rn neuralwatt src/*.py
returns 20 hardcoded refs across 6 files, tabulated as the actual work list.
Three findings the draft did not have:
- metrics.py:175 computes catalog staleness over provider='neuralwatt' only,
so a second provider whose poller stops would keep the warning green while
its catalog froze -- the silent-and-open failure mode again, by a new door.
- leaderboard.py:125 writes priors to a hardcoded provider, so a curated prior
for a shared model attaches to the NeuralWatt row only. Harmless today
(leaderboards.yaml ships empty) but it decides the proficiency-sharing
question by accident.
- dispatcher.py:2879 resolves a pinned model id against NeuralWatt rows only.
Since capability flags fail closed, a pin on a second provider would 422 and
read as a capability problem rather than provider scoping.
Also adds a risk the draft missed: poller.mark_stale runs only inside main(),
which returns early on RequestException, so with two providers one fetch
failure can skip staleness marking for the other -- or mark rows stale that
were never that provider's. Requires per-provider fetch isolation and a
per-provider sanity floor.
Closes three open questions by inspection: circuit_breaker IS provider-generic
(keyed on the (model_id, provider) tuple), 'no serving class' IS a real state
(parse_serving_class returns schema defaults), and proficiency's PK already
supports per-provider divergence -- though the propagate_to_variants sharing
question survives as a genuine design decision.
Corrects the 'needs ~zero change' list, which metrics.py and leaderboard.py
already contradict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Draft outline for generalizing routing beyond NeuralWatt: the real coupling
is narrow (poller.py's fetch_neuralwatt, and three tangled things in
dispatcher.py -- SSE-comment usage parsing, the kWh cost model, and
single-account refusal handling), while scoring/tiering/routing/proficiency
already operate on provider-agnostic DB rows. Recommends Z.ai as the first
test provider (same GLM weights already served via NeuralWatt, making it a
same-weights cross-provider comparison rather than a disjoint catalog),
with OpenRouter, DeepInfra, Together, and Fireworks as further candidates
for cheap prepaid-credit testing. Not decision-complete -- several open
questions flagged before this graduates to FINAL.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NFq4QnaQx1CPvoEXoNijk2
The decision table declares "id, kind, category, tier, ctx, selected, est $,
flex" and has never had a timestamp -- observed_at is already carried by
decision_row and simply never rendered, so that one is a display fix.
Five columns were added to route_decisions across this session's merges and
none reached the TUI: profile (PR #25), exploration (proficiency branch),
pinch_original_tokens/pinch_final_tokens (PR #18), plus request_id and
session_key. The admin portal got a Profile column; the TUI did not, from the
same data.
Deliberately does NOT add six columns to an eight-column table on a terminal.
profile earns a column because it changes which models were considered at all;
exploration becomes a flag beside flex, because an epsilon-greedy pick is not a
ranking result and must not read as the router's judgement; the rest go to the
existing detail popup as per-decision forensics.
The quota panel section is explicitly SEQUENCED behind
plans/quota-balance-and-burn-rate.md and must not land before it -- a progress
bar against a plan figure that is routinely exceeded (146% observed, nothing
failed) implies a ceiling that does not exist. If that plan has not landed,
skip the panel rather than reimplementing balance/burn in the TUI.
The load-bearing success criterion is a test that diffs
PRAGMA table_info(route_decisions) against what the TUI model exposes and fails
when a column is added without a decision about surfacing it. The rest of this
plan is a one-time catch-up; that test is what stops the drift recurring.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U