Agents
'Labels Are Not Endpoints': an audit reclassifies 58 MCP agent attack-success labels to benign, dropping verified attacks to zero
arXiv 2608.12880 (2026-08-13) does a treatment-blind reconstruction of an MCP agent security evaluation, collapsing 10,200 execution rows to 180 model-bound requests, 45 semantic requests and 15 observable stimuli, of which 96 requests were structurally interpretable. It then corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to benign completions, taking verified attacks in the locked dataset to zero. If this generalizes, a meaningful share of the agent-security attack-success numbers circulating right now are measuring label leakage rather than compromise — worth checking before you cite an ASR figure.
Source
↳ Follow the thread
No related signals yet.