diff --git a/.gitignore b/.gitignore index f8c5b95..73d7e7e 100644 --- a/.gitignore +++ b/.gitignore @@ -6,6 +6,7 @@ __pycache__/ router.log .venv/ venv/ +.hf-cache/ .omo/ config.yaml.bak.* config/config.local.yaml diff --git a/deploy/README.md b/deploy/README.md index 0ac26c1..42b470d 100644 --- a/deploy/README.md +++ b/deploy/README.md @@ -109,6 +109,15 @@ your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for the repo limits the blast radius on the filesystem, but nothing limits spend. Putting this on a LAN address needs an auth layer first. +If you enable `classifier.mode: local_encoder`, note `HF_HOME` is redirected +into the repo (`%h/llm-router/.hf-cache`) for the same reason — Hugging +Face's default cache lives outside `ReadWritePaths` and the service will +crash-loop trying to download a model into a read-only home directory +otherwise. Run `pip install -r requirements-encoder.txt` before switching to +this mode; the startup check refuses to boot with a clear message if it's +missing, rather than failing opaquely on the first request, but a missing +model still needs the dependency installed first. + ## Using an Ollama on another machine The local LLM does the classifying; it does not have to be on the machine you diff --git a/deploy/llm-router.service b/deploy/llm-router.service index 8d76add..e7182fa 100644 --- a/deploy/llm-router.service +++ b/deploy/llm-router.service @@ -19,6 +19,13 @@ Environment=PYTHONPATH=%h/llm-router/src # Holds NEURALWATT_API_KEY. Create it with: # echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env EnvironmentFile=%h/llm-router/.env +# HuggingFace's default cache (~/.cache/huggingface) falls outside +# ReadWritePaths below; redirect it into the repo instead of widening the +# sandbox to a new home-directory path. Only classifier.mode: local_encoder +# ever touches this. Caught live 2026-09-06: without this, the service +# crash-loops with "OSError: Read-only file system" the moment +# local_encoder needs to download a model. +Environment=HF_HOME=%h/llm-router/.hf-cache ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080 --timeout-graceful-shutdown 5 Restart=always