Research
HarnessSafe: 328-Case Benchmark Shows Agent Memory, Skills, and Tools Are Distinct Attack Surfaces — Containment Is Carrier-Specific, Not Model-Specific
HarnessSafe evaluates 328 executable attack cases across seven persistent-carrier families (memory, skills, tools, shared artifacts) on mainstream agent harnesses, tracing each as a 'Persistent-Risk Lifecycle' from attacker entry through cross-session persistence to a later benign trigger. The key result: containment depends on the specific carrier AND the harness-model pairing, and end-to-end attack-success rates hide where in the chain an attack actually gets stopped. For anyone running persistent agent memory or skill files, this says your safety posture is per-carrier — hardening one does not generalize.
↳ Follow the thread