Fetching from the wire…
Public story · 2026-09-21 · high
It let 25,370 payments through and blocked only transfers to recipients the passport didn't list.
Why now: The paper posted to arXiv on September 18, 2026.
A deterministic pre-action check stopped every unauthorized transfer by a payment agent across 69,297 evaluations, per a paper posted to arXiv. Agents that pay vendors, contractors, or other agents can be tricked into paying the wrong recipient. This test cut that risk from 140 breaches to zero without slowing legitimate payments.
The test replayed 4,371 human-written attacks from a public capture-the-flag exercise against a live payment agent. It spanned 14 models from 8 labs, five policy configurations, and two tracks, for 225,964 evaluations total.
The zero wasn't a refusal count. Agents behind the layer still executed 25,370 payments, while a policy engine denied 187 of 25,640 evaluated transfer calls. The check, built on the Open Agent Passport spec, blocks payments to recipients a passport doesn't list. It doesn't stop the agent from paying at all. On matched model, prompt, and track triples, the comparison was 105 breaches against 0.
Across 790 source sessions, that puts a per-session upper bound on leakage at 0.38%, small enough that a lucky sample could mask it. The paper doesn't say whether that bound holds outside the CTF's attack style. Request rates swung more across the five policy configurations than across the 14 models, which points at policy design as the bigger lever.
The authors released the dataset, the passports, and the scoring code on HuggingFace, so the 0-of-69,297 claim is checkable rather than asserted. Teams treating model alignment as a payment agent's security boundary are leaning on the layer these numbers say failed 140 times.
Each link below shares sources, entities, or timing with this story.
A new analysis of AP2 v0.2 found eight high-severity gaps where signed payment mandates don't cover the steps that set up the transaction.
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
Missing argument logs hid why the agents failed for ten days across 8,199 runs testing 40 open and hosted models.
CTF-ABACUS traced 1,435 agent attempts across 240 challenges and found only 62-87% of flags backed by demonstrated exploitation.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
An attacker stole an AI agent's signing keys through email injection in under five minutes, per a prior incident this design cites.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.