A deterministic pre-action passport check took unauthorized agent payments from 140 to 0 across 69,297 evaluations
APort Vault (arXiv 2609.22076, 18 Sep) replays 4,371 human-written attacks from a public CTF against a live payment agent across 14 models from 8 labs, five policy configurations and two tracks, for 225,964 total evaluations, with and without a deterministic check implementing the Open Agent Passport spec. At Levels 2 to 4, transfers to recipients the passport did not permit numbered 140 of 76,842 with the model alone and 0 of 69,297 behind the layer, and 105 against 0 on matched model/prompt/track triples, spanning 790 source sessions for a per-session upper bound of 0.38%. The zero was not achieved by refusing to pay: 25,370 payments executed behind the layer while the policy denied 187 of 25,640 evaluated transfer calls. Request rates varied far more across policy configurations than across models, and the dataset, passports and scoring code are released on HuggingFace.
Source
↳ Follow the thread