WeClawArena builds an auditable sandbox for cross-user agent networks, with four attack variants per task
As persistent personal-agent frameworks make agent-to-agent networks a real deployment target, everyday tool use becomes multi-party collaboration over personal workspaces where files, records, tools, and policies are not visible across owners. WeClawArena provides a runtime sandbox and benchmark of 124 base tasks across six cross-user domains, expanded into 620 scenario variants with one benign control and four attack-vector variants each. It records peer messages, tool calls, resource operations, governed decisions, and final workspace states, reporting utility and attack success rate separately so failures can be attributed to task breakdown, privacy leakage, poisoned evidence, or invalid authority paths.
Source
↳ Follow the thread