Agents
KEX-bench: coding agents reach kernel crashes but mostly fail to turn them into exploit primitives
KEX-bench (arXiv 2609.25591, submitted 2026-09-22) runs coding agents in isolated VMs against 45 tasks built from 40 Linux and Windows kernel CVEs, and a deterministic verifier checks each exploit primitive. With no reference PoC, the best configuration solves 14 of 25 Linux tasks (56%) and 1 of 20 Windows tasks (5%). Given a reference PoC it solves 31 of 45 (68.9%). For builders this puts a number on the gap between 'agent finds a bug' and 'agent builds a working exploit', and shows that handing a PoC to an agent sharply raises its offensive capability.
Source
↳ Follow the thread