Fetching from the wire…
Agents2026-09-21 · source-backed
APort Vault replays 4,371 human-written attacks from a public CTF against a live payment agent, across 14 models from 8 labs, five policy configurations and two tracks, for 225,964 total evaluations. At Levels 2 through 4, transfers to recipients the passport didn't permit numbered 140 of 76,842 with the model alone and 0 of 69,297 behind the deterministic layer. The zero wasn't achieved by refusing to pay: 25,370 payments executed behind the layer while the policy denied 187 of 25,640 evaluated transfer calls. Request rates varied far more across policy configurations than across models, which is the finding to carry: your policy config matters more than your model choice.
Each link below shares sources, entities, or timing with this story.
The Agent Payments Protocol signs the finished transaction but not the decision behind it, so text in a product description can steer the agent (arXiv 2609.11757). Against the Gemini Flash-Lite models that AP2's sample agents use by default, the attacks fetched another user's...
arXiv 2609.18674 extends CaMeL with a static verification layer. CaMeLoT translates a generated plan into a finite-state transition system labeled with tool calls, provenance and taint, then checks it against CTL policies with nuXmv before execution starts. Unsafe plans get re...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
The failure mode is a well-formed but policy-forbidden call, cancel a booking, change a passenger count, that neither the tool nor the agent's self-report flags (arXiv). In the airline domain tested, the fix wasn't more reasoning. It was cheap, read-only deterministic gates th...
AgentLSD separates adversarial task contamination from prompt injection: injection needs attacker-supplied instructions, contamination works through non-instructional evidence like fake results and decoy endpoints planted in pages, logs, configs and command output. Six models,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.