Research
SWE-bench Science: Claude Code With Opus-5 Max Scores Under 50% pass@1 on Real Scientific Repos
The OpenMOSS team released SWE-bench Science (arXiv 2608.19799), 119 repository-level tasks drawn from 98 GitHub repos across 20 scientific domains, split into Issue-driven, Expert-exploratory, and Engineering-integration paradigms. The best agent tested, Claude Code with Opus-5 (max), lands below 50% pass@1. A paired ablation that strips explicit scientific guidance while keeping repo and execution context shows domain knowledge is not uniformly helpful: well-grounded facts improve average performance and token efficiency, but poorly aligned guidance anchors the agent and does not raise exact repair success.
↳ Follow the thread