SWE-bench Verified Leaderboard Update: Fable 5 at 0.950, Best Open Model 15 Points Behind
LLM Stats·medium signal
The SWE-bench Verified leaderboard, refreshed in July 2026 across 104 evaluated models, has Claude Fable 5 leading at 0.950. DeepSeek-V4-Pro-Max is the top open-source entry at 0.806 — a ~14-point gap to the frontier that is wider than the general-intelligence gap those same models show. Among models within 10% of the leader, Claude Opus 4.8 is cheapest at $5.00 per Mtok input with a 0.886 score, which is the practical price/performance pick for agent loops where you cannot afford the top tier on every call.