'Compliance Theatre': A 120B LLM Judge Loses 47 Accuracy Points to Keyword Stuffing on Regulatory Evaluation
Principle-Bench releases 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles with paraphrase, keyword-stuffing, and boundary perturbations under a pre-registered rubric — the first benchmark scoring LLM-as-judge on accuracy, paraphrase robustness, adversarial robustness, and calibration together. The 120B judge that leads on benign inputs collapses from 0.74 to 0.27 accuracy on keyword-stuffed Consumer Duty inputs, and a judge from a different model family agrees with it at only Cohen's kappa 0.16 on that split, localizing the failure to the model rather than the corpus. Across keyword counting, three sentence-transformer embedders, an open-weight judge, and a calibrated cascade, no method wins all four axes.
↳ Follow the thread