Research
Internal Harmfulness Scores Rank Successful Jailbreaks BELOW Failed Ones — Outcome AUROC of 0.220 on Wrapped Prompts
This 2026-08-10 audit separates two things safety filters conflate: harmful intent (a property of the prompt) and jailbreak success (an outcome produced later by a specific model, decoder, and judge). Using Active Attention Probing to fix a content-independent measurement coordinate, the authors pair every goal with plain and wrapped versions and generate real completions. On Llama, wrapping raises harmful generation from 0.05 to 0.27 while harmful-intent AUROC falls from 0.936 to 0.803 — attacks get more dangerous as prompts look safer. Among wrapped harmful prompts, outcome AUROC is 0.220, an active reversal, reproduced across three target models, seven attack families, and two judges.
↳ Follow the thread