Skills
A cheaper reviewer from a different model family beats self-review by 12 points; same-model self-review falsely rejects 35% of its own correct answers
Across a fixed set of 100 olympiad math problems in an execute-review-revise pipeline, a cross-family mid-tier reviewer lifted final accuracy from 52% to 64% (p = 0.0005) with zero damaged answers, while same-model self-review had the highest error-detection recall of any condition (0.85) and still produced no significant gain. Self-review rejected 2.1x as often for a third the repair rate (15% vs 43%) and falsely rejected 35% of its own correct answers against 2% for the cross-family reviewer. The paper also finds a capability floor: the weakest reviewer changed zero of 100 final answers while doubling token cost, so a reviewer below the executor's level is pure waste, not cheap insurance.
↳ Follow the thread