CalibForge Builds 5,431 Terminal Tasks Calibrated to Solver Difficulty, Lifting SWE-bench Pro by 27.7 Points Over Base
CalibForge (arXiv 2608.06352, Aug 6) argues that executable validation only proves a training task is solvable, not that it's learnable, and synthesizes tasks revised through adversarial solver calibration — targeting disagreement across a heterogeneous solver pool, or a designated strong-pass/weak-fail relation. The resulting 5,431 calibrated terminal tasks train models to 32.58% and 47.57% on Terminal-Bench 2.0, with largest gains over base of 24.71 points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Ablations show both calibration strategies beat authoring-plus-validation alone and ordinary single-solver feedback, which is the practical takeaway for anyone generating synthetic agent training data.
↳ Follow the thread