Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.09476 ran 24,000 trajectories across 15 LLMs and 6 cowork agents, finding variation across base models (10.1%–94.4%) far exceeds variation across harnesses (73.7%–94.4%). Note the harness floor: 73.7%. No harness tested brought attack success anywhere near zero. SADF measured a 2.6x spread between the best and worst wrappers; ActBench measured that even the best wrapper fails most of the time. Pick your framework carefully and assume it doesn't save you.
Each link below shares sources, entities, or timing with this story.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, assume, between).
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, best).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, attack).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (best, harness).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (attack, between).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, doesn).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, attack).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (agent, attack).