ToolHazard Auto-Synthesizes Adversarial Agent Environments Instead of Hand-Building Them, and Its Alignment Data Transfers to AgentDojo
Existing indirect-prompt-injection research reuses a handful of manually implemented environments with predefined injection points, which caps how broadly agent security can be studied. ToolHazard uses an Environment Simulator, an Attacker Agent, and a User Simulator to synthesize executable stateful environments, discover viable injection points, generate environment-specific payloads, and build state-grounded long-horizon tasks. The resulting ToolHazard-Bench exposes substantial agent vulnerabilities and shows injection timing and placement materially change attack success; alignment data generated by the framework improves security on both ToolHazard-Bench and AgentDojo without hurting benign task utility.
↳ Follow the thread