Skills
LLM code reviewers degrade as review rounds accumulate, and MCR-Bench has the state labels to show where
MCR-Bench is the first defect state-aware benchmark for multi-round review: 2,269 real tasks across five languages, each annotated with defect description, type and severity plus cross-round state labels tracking a defect's full trajectory. Mainstream LLMs show limited defect detection and lifecycle tracking, degrading significantly as rounds increase, and miss semantically complex or low-salience defects disproportionately. Error analysis names the two mechanisms, cross-round temporal misalignment and inadequate long-range memory, which is a direct argument for re-anchoring the defect list every round rather than relying on the conversation.
↳ Follow the thread