Audio Prompt Injection Piggybacks on User Speech for 69.1% Attack Success Against Gemini 3 Pro
'Piggybacking on Perception' (arXiv 2607.28165, July 30) attacks always-listening multimodal agents by hiding malicious instructions in ambient audio that overlaps user speech, using instruction augmentation and scenario concealment so the injected command is imperceptible. The authors release AudioAgentSecurity — 8 real-world task scenarios and 10 attack patterns — and evaluate 11 agents including Gemini 3 Pro and GPT-4o-audio, averaging 69.10% attack success rate against Gemini 3 Pro. Their CADV defense (acoustic source separation plus cross-modal consistency analysis) detects over 90% of injections where prompt-level defenses fail, and real-world tests with volunteers on a Doubao AI smartphone confirmed the attacks work in dynamic environments.
Source
↳ Follow the thread