Rhetoric Reward-Hacks AI Reviewers: 4,200 Manuscripts From 120 ICLR 2026 Submissions Show Evidence Framing and Novelty Stance Move Scores With Content Held Constant
arXiv 2608.08975 (August 10, 2026, 43 HF upvotes) built a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions, had two LLM rewriters push six rhetorical dimensions in opposing directions while preserving reported scientific content, then had five LLM reviewers grade them under standard and strict protocols. Evidence framing and novelty stance produce the largest swings, with scope framing a weaker second tier — and the effect is regression-shaped: low original scores rise, high ones fall, with the clearest contrasts in the middle. Elaborate workflows do not help; joint rewriting is rewriter-dependent and reviewer-guided rewriting does not beat an unguided second pass.
↳ Follow the thread