SWE-rebench Expands Beyond Python to Go, Java, Rust and TypeScript — and Its Contamination-Free Spread Is Only 2 Points Wide
The SWE-rebench team announced a multilingual leaderboard slice covering Go, Java, Python, Rust, and TypeScript, evaluating GLM-5.2, DeepSeek-V4 Pro, Qwen3.6-27B and others on real-world tasks. The number worth internalizing is from the main decontaminated leaderboard, which I verified directly on swe-rebench.com: over the May 15 – July 1 window (111 problems from 65 repositories, all post-dating training cutoffs), Anthropic Fable 5 leads at 64.5%, Grok 4.5 at 63.8%, Opus 5 at 63.4%, GLM-5.2 at 62.9%, and GPT-5.6 Sol at 62.3% — a 2.2-point spread across five frontier models, versus the 95.0 Fable 5 posts on SWE-bench Verified. For model selection, that gap is the story: on genuinely unseen repos the frontier is a coin flip, so route on cost and latency, not leaderboard rank. Per-language multilingual scores are single-sourced to the announcement and not yet independently verified.
Source
↳ Follow the thread