local_energy_call_sites["classify"] gated on whether classifier.base_url resolves to loopback. That is the right question in local_llm mode, where base_url IS the machine doing the work. It is meaningless in local_encoder mode. The encoder places no HTTP call at all -- the model runs in-process on this host's own GPU -- so base_url describes nothing about where the work happens. The consequence was silent and backwards: docs/local-models.md recommends pointing classifier.base_url at a VPN address rather than 0.0.0.0, and doing that while in encoder mode switched off metering for a GPU whose electricity is on this machine's own bill, then warned about a URL nothing calls. Narrow on purpose: local_llm mode still gates on base_url, because a remote Ollama genuinely is unmeterable from here. test_a_remote_ollama_classifier_is_still_not_metered pins that the narrowing stays narrow. Found while enabling local_energy on this deployment, which runs classifier.mode: local_encoder. Metering is confirmed working live -- facebook/bart-large-mnli, 63.35 W over 0.206 s, 3.63e-06 kWh -- since this host's base_url happens to be localhost. A VPN classifier would have lost it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
20 KiB
20 KiB