Fetching from the wire…
Public story · 2026-07-30 · high
23 frontier models summarized alerts they'd already gotten, but skipped the raw disk data where hidden intrusions live.
Why now: The benchmark posted July 30, arguing incident response, the exact workload agents are pitched for, is the one they're failing at.
Researchers tested 23 frontier AI models against a simulated network breach, per the new SecRespond benchmark, arXiv 2607.26791. Zero of the 23 closed out a single range end to end, a real gap for teams pitching AI agents at incident response.
Ten cyber ranges made up the test, covering four entry-point types, 21 ATT&CK techniques, and five operating systems, evaluated on the OpenCode harness.
Each agent gets a forensic disk snapshot of a breached host, plus alerts, vulnerability scans, and baseline security checks. It has to produce three reports, on intrusion, baseline risk, and vulnerability risk, plus a remediation plan.
Same failure mode across every model. Agents write solid summaries of what the alerts already caught, but skip the disk snapshot as a lead worth chasing on its own. Intrusions that never tripped an alert mostly go unfound.
I'd bet these models keep writing clean reports on what the alerts already flagged, and keep missing what they didn't. Explaining a known incident and hunting for an unknown one are different skills, and right now only the first one works.
Each link below shares sources, entities, or timing with this story.
OpenCode competes with Claude Code / Shared entity: Opencode / Same source domain / Earlier coverage
Linked by a graph relationship (OpenCode competes with Claude Code); both cover Opencode; reported by the same outlet (arxiv.org).
Linked by a graph relationship (OpenCode competes with Claude Code); both cover Opencode; reported by the same outlet (arxiv.org).
OpenCode competes with Claude Code / Shared entity: OpenCode / Shared topic / Earlier coverage
Linked by a graph relationship (OpenCode competes with Claude Code); both cover OpenCode; overlapping topics (agent, check).
AionUi uses OpenCode / Shared entity: OpenCode / Earlier coverage
Linked by a graph relationship (AionUi uses OpenCode); both cover OpenCode; earlier OpenCode coverage from 2026-07-29.
OpenCode competes with Claude Code / Shared entity: OpenCode / Earlier coverage
Linked by a graph relationship (OpenCode competes with Claude Code); both cover OpenCode; earlier OpenCode coverage from 2026-07-27.
Linked by a graph relationship (OpenCode competes with Claude Code); both cover OpenCode; earlier OpenCode coverage from 2026-07-25.
Linked by a graph relationship (OpenCode competes with Claude Code); both cover OpenCode; earlier OpenCode coverage from 2026-07-25.
Linked by a graph relationship (OpenCode competes with Claude Code); both cover OpenCode; earlier OpenCode coverage from 2026-07-21.