Prompt injection success is a function of the attacker's compute budget, not a fixed property of the victim agent
This paper reframes indirect prompt injection as a test-time search over an attack surface induced by the environment, the user task and the injection task, and builds an agentic attacker with a search harness that does environment reconnaissance, structured strategy reasoning, and adaptive evaluation using victim-agent feedback. Increasing attacker test-time compute consistently improves both vulnerability discovery and exploitation, and ablations show explicit strategy management is what prevents redundant search from eating the larger budgets. The practical consequence is that a red-team result reported without an attacker compute budget is uninterpretable, and your agent's measured injection resistance will fall as attacker budgets rise.
↳ Follow the thread