Hacker News
Real-SWE Benchmarks Coding Agents on Licensed Private Production Codebases: Fable 5.1 Tops It at 38.8%, GPT-6 Astra 33.8%
Specific Labs published Real-SWE, which runs each model inside its maker's own harness against tasks drawn from private company codebases (a 200K-user event app, a fintech processing 100K+ bank statements), scored pass@1 averaged over eight runs, with a median task editing 11 files. Fable 5.1 in Claude Code leads at 38.8%, then GPT-6 Astra in Codex CLI at 33.8%, Gemini 3.8 Flash 31.2%, GLM 5.3 28.8%, Grok 4.6 and Muse Spark 1.3 tied at 23.8%, Kimi K3 18.8%, GPT-5.6 Sol 16.2%. At $6.96 per rollout Fable costs about $17.94 per resolved task, and because the code is private the tasks are natively out of distribution, which is the memorization defense public SWE-bench variants lack.
↳ Follow the thread