Sources
"The Safeguard Worked. Is the LLM System Safer?" argues refusal and attack-success rates cannot support the claim they are used for
arXiv 2609.00519, posted 2026-09-01, depth-codes the safeguard literature and shows the evidence requirements are asymmetric: a single successful attack establishes that harmful help remains, but no amount of favorable local scoring establishes that little remains, because that also requires evidence about what the surrounding system still allows after the safeguard does its local job. Only a small minority of coded claims supply or derive that system-level evidence, and just one bounds its scoped residual. The practical consequence for anyone shipping a guarded LLM product is that a better refusal rate is not by itself a stronger deployment safety claim.
↳ Follow the thread