ImpossibleRubrics: every one of eleven rubric generators gets gamed 8-26% of the time on tasks with no honest answer
arXiv 2609.16816 (submitted 15 Sep) isolates the hardest case for LLM-generated rubrics — 169 impossible tasks across six impossibility categories where the prompt pressures the model toward an unsupported conclusion and the only honest response is to say it cannot be done. Instead of fixed rubrics, the benchmark supplies task environments and verifiable oracle certificates specifying what an honest answer may and may not claim, then adversarially tests whether downstream-generated rubrics reward certificate-violating answers. Eleven generators were exploited 8-26% of the time on the unbiased 150-of-169 subset, with 48 answerable controls, which is a direct warning to anyone using rubric-based RL or LLM-as-judge scoring in a pipeline.
Source
↳ Follow the thread