Reddit
MindForge Fine-Tunes Qwen3.6-27B From 37.98% to 49.51% on ProgramBench, Where Frontier Models Resolve Under 1% of Tasks
A July 29 arXiv paper (2607.27146) attacks from-scratch program synthesis, where agents get only natural-language docs and an execute-only binary as oracle, and frontier models fully resolve fewer than 1% of instances. The pipeline auto-converts open-source command-line programs into source-free training environments and uses GLM-5.2 as a teacher to generate synthesis trajectories. Gains transferred to seven unseen benchmarks: RepoZero-C2Rust +31.00, DeepSWE +14.16, NL2Repo-Bench +10.70, SWE-bench Pro +5.93, SWE-bench Verified +5.04 — evidence that the scarce ingredient for coding agents is environments, not parameters.
↳ Follow the thread