Cross-Model Code Review Study Goes Viral — Then r/ClaudeAI Corrects the Headline: Claude Alone (91.4%) Still Beats Claude-Reviewing-Codex (89.7%)
A LeadDev article published August 4 by Chris Stokel-Walker covering arXiv:2607.21656 (Xiang et al., 'Cross-Model LLM Code Review') hit 409 upvotes on r/ClaudeAI with the framing that Claude review lifts Codex GPT-5.5 drafts from 71.6% to 89.7% on 116 medium/hard LiveCodeBench tasks. The top comment (93 upvotes) pulled the rest of the abstract: Claude Opus 4.7 working alone scores 91.4%, Claude self-review leaves it unchanged, and Codex reviewing Claude actively degrades it to 82.8%. The practical takeaway for anyone wiring multi-agent review loops is that the pairing is asymmetric and hierarchy-driven — a weaker reviewer on a stronger drafter is a net negative, and the best measured config is no reviewer at all.
↳ Follow the thread