Sources
A 50-PR code review benchmark puts GPT-5.6 Luna at 28x cheaper than Astra and 74% precision against 96%
Entelligence published a 2026-09-14 benchmark running both models over 50 public PRs, ten each from Cal.com, Sentry, Discourse, Keycloak and Grafana, with identical prompts and a dual-judge scheme where GPT-6 Astra and GPT-5.6 Sol both had to agree before a bug counted. Luna cost $0.0041 per review against Astra's $0.113 and found 69 verified bugs at 74% precision versus Astra's 92 at 96%, but the overlap was partial: Luna caught 25 bugs Astra missed while Astra caught 48 Luna missed. The security split is the operational line, with Luna finding 9 of 24 security bugs to Astra's 19, which is why the authors would not let it review auth or permission code alone.
Source
↳ Follow the thread