Sources
Dan Luu re-graded Senior SWE-Bench ten times and 23% of results flipped
In a post that reached the HN front page this week, Dan Luu dissects Senior SWE-Bench alongside two non-AI benchmarking failures. The headline numbers (Fable 5 at 29.1%, Opus 4.8 at 25.0%, GPT-5.6 Sol at 24.4%) sit inside a scoring scheme with discontinuous thresholds, including a task whose 1-line reference solution forces 'tasteful' answers into 3 lines. Re-running the LLM grader ten times on the same solutions flipped 23% of official results, each task runs exactly once despite large per-run variance, and only 4 of 113 tasks are Rust. If you pick coding models off this leaderboard, the ranking is inside the noise.
Source
↳ Follow the thread