aiXamine Runs 5,000+ Test Runs Across 120 LLMs and Names a Quantifiable Safety Tax Plus a Distillation-Induced Robustness Collapse From 56.9 to 2.6
aiXamine (arXiv 2608.20554, Aug 20) is a black-box platform running 46 tests across nine services to evaluate safety, security and privacy as interdependent rather than separate axes, applied to over 120 LLMs in more than 5,000 runs. It documents a safety tax where stronger alignment systematically raises over-refusal (a model scoring 99.3 on safety alignment refused one in three benign queries), finds privacy near-orthogonal to other trustworthiness dimensions, and characterizes distillation-induced robustness collapse where off-policy distillation without on-policy correction dropped robustness from 56.9 to 2.6 on the same base architecture. The practical read for builders picking an open-weight model: a single-axis leaderboard number cannot tell you what the model gave up.
↳ Follow the thread