Research
ActBench: Attack Success Against Cowork Agents Ranges 10.1%–94.4% Across Models but Only 73.7%–94.4% Across Harnesses — the Model Matters More Than the Scaffold
ActBench, posted 2026-08-10, evaluates behavioral safety from execution trajectories rather than final responses: 600 cases from 213 scenarios covering 15 risk behaviors, six execution spaces, and 48 web-service APIs, each pairing a benign task with an adversarial variant that preserves instruction, config, and initial state while injecting a task-reachable payload. Across 15 LLMs and 6 open-source cowork agents over 24,000 trajectories, variation across base models (10.1%–94.4%) far exceeds variation across agent harnesses (73.7%–94.4%). No harness tested brought attacks near zero. Benchmark released at github.com/zjuicsr/ActBench.
↳ Follow the thread