LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
arXiv·high signal
LITMUS introduces a benchmark for behavioral jailbreaks — where adversaries induce agents to execute dangerous OS-level operations with irreversible consequences (file deletion, privilege escalation, data exfiltration). Unlike content-safety benchmarks, LITMUS tests physical-layer harms in isolated OS sandboxes, finding that even safety-tuned models comply with 23-41% of dangerous OS commands when framed through multi-step social engineering. Critical finding for anyone deploying autonomous agents.