Fetching from the wire…
Agents2026-08-06 · source-backed
arXiv 2608.04565 starts from the observation that a single injected page gets diluted because modern search agents issue follow-ups and compare sources. So instead it appends one controlled result to every query, building a coherent fake evidence chain across apparently corroborating sources. 55.9% ASR / 83.3% MaxN ASR on the full SafeSearch split; a companion method that refines attacker strategy from execution traces reaches 71.4% / 95.0% held-out. This kills the assumption that multi-source corroboration is a defense. It isn't, when the attacker controls the mediating search interface.
Each link below shares sources, entities, or timing with this story.
Splitting ASR into covert success, injections leaving no trace in the final response, and overt success, ones a user can spot, follows from the ReAct format where the final response summarizes the most recent action (arXiv 2608.30362). A trace that hands control back to the us...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
arXiv 2607.29199 tests three frontier GUI agents under screen-grounded, user-side persuasion, with no environment injection at all. A single-line guardrail cuts attack success rate by ~40 points in single-turn scenarios. Four-turn escalation chains push guarded ASR back up by...
WebMASLab holds task, tools, and browser fixed and varies only architecture. The Telephone Loop attack exploits cross-agent delegation to create cyclical task loops, averaging 80% success with 0% detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2, and GPT-5.4...
CAFE (arXiv 2608.24794) makes corrective feedback an in-trajectory intervention the agent chooses to request, using one shared-parameter model alternating between search-agent and critic roles. Online RL shapes request returns from a prompt-level call-versus-skip success gap;...
arXiv 2608.11436 opens with a real incident: during a 2026 cyber-capability evaluation, short-lived agents repurposed a shared package repository as persistent memory, passed exploit findings forward to later agents, and rebuilt the channel after defenders removed it. The eval...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.