Fetching from the wire…
Public story · 2026-07-30 · high
Frontier models solve fewer than 1 percent of these tasks, and MindForge's gains carried over to seven other benchmarks too.
Why now: The paper posted to arXiv in July.
Qwen3.6-27B jumped from solving 37.98% to 49.51% of tasks on a from-scratch coding benchmark, per a MindForge paper posted to arXiv in July.
That matters for anyone building coding agents. Frontier models resolve under 1% of these tasks, so a 27B open model closing that gap changes which teams can compete without frontier-scale budgets.
The task is designed to resist shortcuts. Agents write code from scratch against only natural-language documentation, checking their work against an execute-only binary with no source to read or copy.
MindForge's fix isn't a bigger model. Its system auto-converts open-source command-line programs into training environments. GLM-5.2 then serves as a teacher, generating synthesis trajectories the smaller model learns from, per the paper.
The gains didn't stay confined to that one benchmark. Tested against seven unseen benchmarks, per the paper, results included gains of 31.00 points on RepoZero-C2Rust and 14.16 on DeepSWE. NL2Repo-Bench rose 10.70, SWE-bench Pro climbed 5.93, and SWE-bench Verified gained 5.04 points.
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage / Tension
Both cover GLM, SWE, Verified; overlapping topics (agent, benchmark, coding, evidence); earlier GLM coverage from 2026-07-17.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover Bench, SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, coding).
Shared entities / Shared topic / Earlier coverage
Both cover DeepSWE, SWE, Verified; overlapping topics (agent, benchmark, coding, deepswe, environment); earlier DeepSWE coverage from 2026-05-27.
Both cover Bench, Qwen3, SWE, Verified; overlapping topics (coding, swe bench); earlier Bench coverage from 2026-07-27.
Both cover Bench, GLM, SWE, Verified; overlapping topics (benchmark, coding); earlier Bench coverage from 2026-06-14.
Both cover Bench, GLM, SWE, Verified; overlapping topics (benchmark, coding); earlier Bench coverage from 2026-03-04.
Shared entities / Shared topic / Earlier coverage / Tension
Both cover Bench, Qwen3, SWE; overlapping topics (agent, benchmark, coding); earlier Bench coverage from 2026-04-21.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover Bench, SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (agent, coding).