Google DeepMind 'AI Agent Traps': Six Attack Categories, 86% Hidden Prompt Injection Success Rate
Google DeepMind / The Decoder·high signal
A DeepMind study introduces the taxonomy of 'AI agent traps' — six categories attacking perception, reasoning, memory, action, multi-agent dynamics, and human supervisor components of an agent's operating cycle. Hidden prompt injection in HTML/CSS achieves 86% success rate; latent memory poisoning succeeds 80%+ with less than 0.1% data contamination. Every AI agent tested was successfully compromised at least once, with consequences including unauthorized data access. Researchers recommend adversarial hardening and multi-stage runtime filters (source filters, content scanners, output monitors).