Fetching from the wire…
Security2026-08-11 · source-backed
arXiv 2608.09624 separates harmful intent (a prompt property) from jailbreak success (an outcome from a specific model, decoder, and judge). On Llama, wrapping a prompt raises harmful generation from 0.05 to 0.27 while harmful-intent AUROC falls from 0.936 to 0.803. Attacks get more dangerous exactly as prompts look safer. Among wrapped harmful prompts, outcome AUROC hits 0.220, an active reversal, not just a failure, reproduced across three target models, seven attack families, and two judges. If your safety filter scores intent and you assume that predicts outcomes, it predicts the opposite.
Each link below shares sources, entities, or timing with this story.
Shared entity: Among / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Among; reported by the same outlet (arxiv.org); overlapping topics (among, audit).
Shared entity: AUROC / Same source domain / Shared topic / Earlier coverage
Both cover AUROC; reported by the same outlet (arxiv.org); overlapping topics (attack, auroc).
Shared entity: AUROC / Same source domain / Earlier coverage / Tension
Both cover AUROC; reported by the same outlet (arxiv.org); earlier AUROC coverage from 2026-08-04.
Shared entity: Among / Same source domain / Earlier coverage / Tension
Both cover Among; reported by the same outlet (arxiv.org); earlier Among coverage from 2026-07-29.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (active, attack, audit); pushes against this story (against).
Shared entity: AUROC / Same source domain / Earlier coverage
Both cover AUROC; reported by the same outlet (arxiv.org); earlier AUROC coverage from 2026-07-30.
Shared entity: Among / Shared topic / Earlier coverage
Both cover Among; overlapping topics (among, model); earlier Among coverage from 2026-07-28.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (attack, intent); pushes against this story (but).