ChannelGuard Shows an 'Attack Success 0.000' Multi-Agent Pipeline Owed 54 of 60 Blocks to Azure's Server-Side Filter, Not Its Own Code
Across a 2,100-trace evaluation over eight attack families, five defenses, and three model backends, Hossain et al. found an undefended planner/worker/verifier/synthesizer pipeline that reported perfect 0.000 attack success on tool- and memory-poisoning — but 54 of its 60 blocks came from Azure GPT-5's provider-side filter, and safety silently shifted to the agent model's own alignment on a backend without one. Their training-free ChannelGuard puts embedding-similarity information-bottleneck gates on every inter-agent hop, blocking Tool Poisoning 30/30 identically across Azure GPT-5, Sonnet 4.5, and Haiku 4.5, halving prompt-injection success (0.333 to 0.167) and preserving GSM8K accuracy at 0.867, with no added LLM call. The honest caveat: white-box adaptive paraphrase evades every embedding gate, and the whole study cost $47.36 to run.
↳ Follow the thread