Fetching from the wire…
Public story · 2026-08-24 · high
One model scored 99.3 on safety but refused a third of harmless prompts, and a distillation shortcut crashed another's robustness score to 2.6.
Why now: The paper posted in August 2026, covering 120 open-weight models in one benchmark.
A benchmark called aiXamine ran more than 5,000 tests on 120 language models and found top safety scores carry a hidden cost. One model scored 99.3 on safety alignment, then refused one in three benign, harmless queries, per the aiXamine paper.
That tradeoff matters for anyone picking an open-weight model off a leaderboard. A high safety score can mean careful tuning, or it can mean a model trained to refuse by default, and the ranking doesn't show which.
The paper's second finding is sharper. Testing the same base architecture under different training methods, aiXamine found that off-policy distillation without on-policy correction dropped robustness from 56.9 to 2.6.
A drop that size, on the same base architecture, wouldn't show up in a single safety number.
Anyone deploying an open-weight model needs the refusal rate on benign traffic and the training method behind the safety number. The paper doesn't say how many of the 120 models it tested shipped with that gap unmeasured.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (correction, document); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (alignment, benign, document, safety).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (benign, collapse); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (collapse, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (distillation, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (collapse, model); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (architecture, model); traces where this leads (which means).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (correction, model); pushes against this story (but).