AgentLSD plants fake flags and decoy endpoints in CTF challenges: traps cost a solving agent +20 turns and +2k reasoning tokens even when it still wins
The framework separates adversarial task contamination from prompt injection. Injection needs attacker-supplied instructions; contamination also works through non-instructional evidence like fake results and decoy endpoints planted in pages, logs, configs and command output that the agent inspects. Across six models on 11 web CTF challenges with deterministic trap generation, runtime injection and delivery verification, agents captured 41% of flags clean and no model solved every challenge. Traps raised turn count by roughly 20 and reasoning tokens by about 2k even on runs that still succeeded, with heterogeneous solve-rate damage: some model-challenge pairs were unaffected, others followed decoys or submitted wrong flags. Clean benchmark performance therefore understates how an agent behaves on a real, adversarial surface, and the token overhead is a budgeting item.
↳ Follow the thread