IH-Benchmark: Instruction-Hierarchy Compliance Ranges From 98.2% to 20.5% Across 37 Models, and System-Prompt Strength Does Not Predict Tool-Output Robustness
IH-Benchmark (arXiv 2607.25987, 2026-07-28) tests what a model actually does when instructions conflict across priority levels, covering direct system-over-user conflicts and tool-mediated user-over-tool conflicts, built from a human-authored taxonomy of 44 constraint families across generic, health, finance, retail, and coding settings. Across 37 evaluated models, hierarchy compliance spans 98.2% down to 20.5%, and strong system-over-user compliance is explicitly not a reliable proxy for tool-output robustness — several models hold system constraints under direct user pressure but degrade sharply when the conflicting instruction arrives inside a tool result. The most revealing failures are subtle rather than dramatic: models resist unauthorized purchases more reliably than injected disclaimers or small factual distortions.
↳ Follow the thread