Deterministic confirmation rules are cheaper to forge than LLM judges, and routing between the two makes forgery near-certain
Security testing tools must decide for themselves whether an attack worked, and this study shows the target can forge that verdict. Nine of fifteen confirmation mechanisms in a four-stage AI-assisted pipeline were forgeable, and forgeability was predicted entirely by whether the decision reads attacker-controlled data — a prediction fixed in advance that separated sixteen held-out mechanisms exactly and scored 99.9% across 12,203 public scanner template mechanisms. The counterintuitive part for builders: deterministic rules failed at 2% of attacker-controlled response content against a median of 50% for eight open-weight LLM judges, and routing between a rule and a judge raised forgery to 99%. Moving the decisive evidence to a channel the attacker cannot write dropped attack success from 97% to 0%.
↳ Follow the thread