Reddit
Agent Security Benchmark Finds Single-Turn Attacks Fail (0–1%) While 15-Round Adaptive Attacks Succeed 5.4–14%
'Adaptive Adversaries' (arXiv:2607.18063, July 20, by Jain, Hartmann and Li) tests LLM agents against attackers that adapt over turns rather than firing one-shot prompts. Single-turn attacks land 0–1% of the time; multi-turn adaptive attacks over 15 rounds land 5.4–14.0%. Claude Opus 4.6 and GPT-5.4 tied at 5.4% aggregate, but per-scenario variance was extreme — Opus hit 60% on one scenario where competitors stayed at 7% — and only 13 of 21 scenarios distinguished defender pairs at all. Attacks generated by the multi-LLM attacker pool barely overlap existing benchmarks (cosine similarity 0.02–0.14), which suggests current agent-safety evals are measuring the wrong threat surface.
Source
↳ Follow the thread