Sources
EvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%
Submitted 2026-09-05 by Nanxi Li, Bo Li, Dawn Song, Chaowei Xiao and colleagues, this paper argues static expert-written agent harnesses are the wrong shape, since a harness strict enough for one model over-blocks another. It jointly searches natural-language policies and executable code logic against model behavior analysis, domain specs and adversarial review. On DecodingTrust-Agent it drops attack success rate from 45.6% to 10.0% for a 3.3-point utility cost and wins 14 of 15 test cells; on AgentDojo it holds 82.8% utility at a 0% attack success rate.
↳ Follow the thread