From 65dd3138ed04315191611f78e0d8f35d5324b641 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Sun, 6 Sep 2026 20:48:26 -0400 Subject: [PATCH] fix(deploy): redirect HF_HOME into the repo, same crash on a fresh install The live incident this branch is already fixing (#48's gated model default) had a second layer once that was fixed: HuggingFace's default cache (~/.cache/huggingface) falls outside ProtectHome=read-only's ReadWritePaths exception, so the service crash-looped with "OSError: Read-only file system" the moment local_encoder tried to download a model -- even with a working model id. Redirect HF_HOME into the already-writable repo path instead of widening the sandbox to a new home-directory location, keeping the "only the repo is writable" invariant intact. Applied live already; this backports it to the tracked deploy template so a fresh install doesn't hit the same crash. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- .gitignore | 1 + deploy/README.md | 9 +++++++++ deploy/llm-router.service | 7 +++++++ 3 files changed, 17 insertions(+) diff --git a/.gitignore b/.gitignore index f8c5b95..73d7e7e 100644 --- a/.gitignore +++ b/.gitignore @@ -6,6 +6,7 @@ __pycache__/ router.log .venv/ venv/ +.hf-cache/ .omo/ config.yaml.bak.* config/config.local.yaml diff --git a/deploy/README.md b/deploy/README.md index 0ac26c1..42b470d 100644 --- a/deploy/README.md +++ b/deploy/README.md @@ -109,6 +109,15 @@ your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for the repo limits the blast radius on the filesystem, but nothing limits spend. Putting this on a LAN address needs an auth layer first. +If you enable `classifier.mode: local_encoder`, note `HF_HOME` is redirected +into the repo (`%h/llm-router/.hf-cache`) for the same reason — Hugging +Face's default cache lives outside `ReadWritePaths` and the service will +crash-loop trying to download a model into a read-only home directory +otherwise. Run `pip install -r requirements-encoder.txt` before switching to +this mode; the startup check refuses to boot with a clear message if it's +missing, rather than failing opaquely on the first request, but a missing +model still needs the dependency installed first. + ## Using an Ollama on another machine The local LLM does the classifying; it does not have to be on the machine you diff --git a/deploy/llm-router.service b/deploy/llm-router.service index 8d76add..e7182fa 100644 --- a/deploy/llm-router.service +++ b/deploy/llm-router.service @@ -19,6 +19,13 @@ Environment=PYTHONPATH=%h/llm-router/src # Holds NEURALWATT_API_KEY. Create it with: # echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env EnvironmentFile=%h/llm-router/.env +# HuggingFace's default cache (~/.cache/huggingface) falls outside +# ReadWritePaths below; redirect it into the repo instead of widening the +# sandbox to a new home-directory path. Only classifier.mode: local_encoder +# ever touches this. Caught live 2026-09-06: without this, the service +# crash-loops with "OSError: Read-only file system" the moment +# local_encoder needs to download a model. +Environment=HF_HOME=%h/llm-router/.hf-cache ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080 --timeout-graceful-shutdown 5 Restart=always -- 2.49.1