AISI Benchmark: Frontier AI Agents Scale Log-Linearly on Multi-Step Cyber Attacks — No Plateau Found, Opus 4.6 Averages 9.8/32 Steps
arXiv / AISI·high signal
A UK AI Safety Institute study (arXiv:2603.11214) evaluated seven frontier models on a 32-step corporate network attack and a 7-step industrial control system attack, tracking performance from August 2024 (GPT-4o) to February 2026 (Opus 4.6). Model capability at 10M tokens jumped from averaging 1.7 steps completed to 9.8 steps, and each 10x increase in inference-time compute yields up to 59% additional steps completed with no observed plateau. The best single run completed 22 of 32 steps — roughly 6 hours of a human expert's 14-hour estimated workload — establishing that adversarial AI capability is scaling predictably with both model generation and compute budget.