Dispatch
MiniMax-M3 Tops Vendor-Reported Open-Weight SWE-Bench Pro at 59.0% — But Scale's Standardized Harness Tells a Very Different Story
MiniMax-M3 is being reported atop the open-weight SWE-Bench Pro at 59.0%, edging Kimi K2.6's 58.6%, but those figures come from vendor-tuned agent harnesses. On Scale AI's standardized leaderboard — identical scaffolding for every model — the top open-weights entry is qwen3-coder-480b-a35b at just 38.7%, a 10-to-30-point gap. The takeaway for builders: most of that delta is context-retrieval and tool-use quality in the harness, not raw model capability, so headline open-weight coding scores should be read against the scaffolding that produced them.
↳ Follow the thread