EquiReview-R treats omission and overcritique as separate risks and halves major overcritique in AI paper review
arXiv 2609.03943 (2026-09-03) argues that an AI reviewer producing more specific criticisms is not producing a better review, because missing a consequential weakness and retaining an unsupported allegation require opposite corrections that aggregate metrics hide. Their system resolves existing concerns against localized evidence, searches for missing issues from both independent and review-conditioned angles, and returns stop, continue or defer. On a frozen cohort of unseen papers it meets its prespecified non-inferiority criterion for major omission while cutting major overcritique from 15.5% to 8.1%, with a one-sided omission upper bound of 9.9%. The retrospective finding that nearly all concerns in a high-recall review lack a definitive evidential disposition applies to any LLM-as-critic loop, not just paper review.
Source
↳ Follow the thread