fix(deploy): redirect HF_HOME into the repo, same crash on a fresh install #49

Merged
alee merged 1 commits from fix/deploy-hf-home-cache into main 2026-09-07 02:53:51 +00:00
3 changed files with 17 additions and 0 deletions

1
.gitignore vendored
View File

@@ -6,6 +6,7 @@ __pycache__/
router.log router.log
.venv/ .venv/
venv/ venv/
.hf-cache/
.omo/ .omo/
config.yaml.bak.* config.yaml.bak.*
config/config.local.yaml config/config.local.yaml

View File

@@ -109,6 +109,15 @@ your allowance. `ProtectHome=read-only` plus a `ReadWritePaths` exception for
the repo limits the blast radius on the filesystem, but nothing limits spend. the repo limits the blast radius on the filesystem, but nothing limits spend.
Putting this on a LAN address needs an auth layer first. Putting this on a LAN address needs an auth layer first.
If you enable `classifier.mode: local_encoder`, note `HF_HOME` is redirected
into the repo (`%h/llm-router/.hf-cache`) for the same reason — Hugging
Face's default cache lives outside `ReadWritePaths` and the service will
crash-loop trying to download a model into a read-only home directory
otherwise. Run `pip install -r requirements-encoder.txt` before switching to
this mode; the startup check refuses to boot with a clear message if it's
missing, rather than failing opaquely on the first request, but a missing
model still needs the dependency installed first.
## Using an Ollama on another machine ## Using an Ollama on another machine
The local LLM does the classifying; it does not have to be on the machine you The local LLM does the classifying; it does not have to be on the machine you

View File

@@ -19,6 +19,13 @@ Environment=PYTHONPATH=%h/llm-router/src
# Holds NEURALWATT_API_KEY. Create it with: # Holds NEURALWATT_API_KEY. Create it with:
# echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env # echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
EnvironmentFile=%h/llm-router/.env EnvironmentFile=%h/llm-router/.env
# HuggingFace's default cache (~/.cache/huggingface) falls outside
# ReadWritePaths below; redirect it into the repo instead of widening the
# sandbox to a new home-directory path. Only classifier.mode: local_encoder
# ever touches this. Caught live 2026-09-06: without this, the service
# crash-loops with "OSError: Read-only file system" the moment
# local_encoder needs to download a model.
Environment=HF_HOME=%h/llm-router/.hf-cache
ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080 --timeout-graceful-shutdown 5 ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080 --timeout-graceful-shutdown 5
Restart=always Restart=always