Copy plans/local-decision-classifier-heldout.yaml (30 tasks) to evals/heldout.yaml unchanged, and add a loader test asserting the file loads 30 scoreable tasks with the expected categories. The held-out set is weak: single author, short prompts, no true distributional shift from the training set. Treat 100% as a ceiling, not a forecast -- it cannot measure generalization.