Research
TIER Replaces Binary Safety Scores With a Six-Label Behavior Scale Across Four Threat-Implicitness Levels
arXiv 2609.05117 argues current LLM safety benchmarks rely on binary metrics that miss how models respond to harmful prompts of varying implicitness. TIER covers four risk domains and four threat levels from explicit harmful requests to sophisticated jailbreaks, scoring responses on a six-label behavior scale with two independent LLM judges. Across six open-weight LLMs, safety behaviors evolve gradually across threat levels rather than flipping from refusal to compliance, and models with similar Attack Success Rates show distinct response distributions.
↳ Follow the thread