ContainmentBench: Taint-Only Prompt-Injection Enforcement Completes Just 16.4% of Authorized Workflows While a Trusted Ledger Reaches 85.7%
ContainmentBench (arXiv 2607.23999, July 27) argues that terminal attack/policy labels hide what actually happens after a prompt injection lands, and instead measures endpoint compliance, logged propagation, recovery, and authorized-action completion separately. In a pre-specified 17,640-rollout study with Qwen2.5-7B-Instruct, all 600 matched active-tainted pairs comparing taint-only versus intent-aware enforcement produced identical zero committed-harm outcomes — yet 73.5% differed in logged trajectory or retained utility. Taint-only enforcement completed only 0.1642 of authorized tainted workflows; a trusted-ledger policy raised that to 0.8567 and a strong tool-boundary baseline to 0.9233 under the same endpoint outcomes. The study is synthetic and single-model, but the lesson generalizes: 'no harm committed' is not a sufficient statistic for agent security.
↳ Follow the thread