Fetching from the wire…
Agents2026-09-01 · source-backed
Splitting ASR into covert success, injections leaving no trace in the final response, and overt success, ones a user can spot, follows from the ReAct format where the final response summarizes the most recent action (arXiv 2608.30362). A trace that hands control back to the user task before ending stays invisible. Their ICoA attack steers the agent back to the user task after firing the injection and posts the highest covert rate across four models on AgentDojo, 3.79 to 12.01 points over the strongest baseline. If you are reporting ASR on a tool-using agent, you are measuring the wrong number.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
arXiv 2608.27234 calls the planner exactly once per query to emit a full plan in a declarative DSL, then applies dual-lattice information-flow control over confidentiality and integrity across explicit data flows and control dependencies, storing results as labeled artifacts a...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
Every skill marketplace runs on one assumption: certify each package, and the ecosystem is safe. CompoSkill breaks that assumption by showing composition risk is a path property, not a node property. The attack works black-box. The attacker knows only a role profile. They down...
WebMASLab holds task, tools, and browser fixed and varies only architecture. The Telephone Loop attack exploits cross-agent delegation to create cyclical task loops, averaging 80% success with 0% detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2, and GPT-5.4...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.