Fetching from the wire…
Public story · 2026-08-04 · high
WebMASLab isolates multi-agent architecture as the sole variable, and one of four models resisted the exploit.
Why now: As of August 4, this is the first study to isolate multi-agent architecture itself as the variable behind a web-agent attack.
An attack called the Telephone Loop breaks multi-agent web crews but leaves single agents alone, per a new WebMASLab study. It hit an 80% success rate with zero detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2 and GPT-5.4. A team routing a browsing task through a crew of delegating agents is exposed to a failure mode a single agent never triggers.
WebMASLab holds the task, the tools and the browser fixed and changes only the number of agents, which isolates architecture itself as the vulnerability. The Telephone Loop exploits cross-agent delegation, trapping the crew in cyclical task loops that never resolve.
Claude Sonnet 4.6 was the exception, catching the attack 92% of the time on its own. Prompt hardening, the obvious fix, didn't generalize: it cut one model's success rate from 100% to 8% but barely moved the rest.
Model choice made the difference here, not prompt patches. Claude Sonnet 4.6 resisted the attack without any patch, while hardening built for one model didn't transfer to the rest. Picking a base model for a multi-agent crew is a security decision, whether the team treats it that way or not.
Each link below shares sources, entities, or timing with this story.
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover ASR, GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover ASR, GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (agent, claude).
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (agent, claude, gpt 5).
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (agent, claude).
Critique uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Critique uses GPT); both cover GPT; overlapping topics (agent, claude).