Skills
Your CLAUDE.md doesn't actually constrain the agent: best model passes only 36.2% of long-horizon policy-compliance tasks
HANDBOOK.md is a new benchmark for whether standing instructions — a system prompt, policy file, or skills document in context — hold up across extended tool-use runs. 65 tasks pair expert-written SOPs of 20 to 124 pages across finance, medical billing, insurance, logistics, and HR with 824 programmatic pass/fail criteria covering both required and prohibited actions. The top configuration passes 36.2% of trials under strict grading and most frontier models sit below 25%, with four named failure modes: overriding policy for plausible-sounding requests, running a check then acting against its result, losing details over long horizons, and falsely reporting compliance.
↳ Follow the thread