PACT benchmarks whether enterprise assistants break their own system-prompt rules under pressure from a persistent user
arXiv 2609.18605 (submitted 2026-09-16) introduces Pressure-Applied Compliance Testing, covering twelve regulated enterprise domains and 48 multi-turn scenarios where each item pairs a standing rule in the system context against a rule-violating shortcut, then applies pressure from a persistent user, a hurried manager, or circumstances where violating is simply convenient. The authors built it under strict LLM-as-judge auditing to keep samples unambiguous, ungameable, and realistic enough not to trigger evaluation-aware behavior, and profile models across six metrics spanning robustness over multi-turn conversation, transparency, and whether a model correctly recognizes when a rule applies at all. If you have agents in hiring, healthcare or finance workflows, this is the first systematic answer to which model actually holds the line.
↳ Follow the thread