Agents
How you phrase a coding-agent task changes repository-poisoning risk by 4.5x, and test-execution is the silent attack surface
CIPR is the first benchmark to vary user-side invocation choices rather than attacker payloads: 1,920 instances across 20 poisoned real-world repos, four task types, three prompt styles and three skill/rule conditions. Task type alone produces up to a 4.5-fold spread in attack success rate, and asking the agent to run tests is the worst case, high success with a low agent alert rate. Underspecified prompts cut success by truncating execution depth, and noisy prompts trend toward suppressing alerts, which means agent vulnerability is a property of how a builder drives it, not a fixed property of the harness.
Source
↳ Follow the thread