Fetching from the wire…
Research2026-09-23 · source-backed
arXiv 2609.25804 mines forks automatically from parallel agent attempts and from detours inside single trajectories, where one direction demonstrably led to a better outcome, and asks the model to choose without seeing what came after. Judgment about which hypothesis to pursue is a distinct weakness that end-to-end success metrics hide entirely, because a model that picks badly and recovers still scores as a pass.
Each link below shares sources, entities, or timing with this story.
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
DSEffi-Bench covers 1,000 instances across 10+ libraries with stress-testing harnesses and human-validated references, evaluated on 16 models (arXiv 2608.30248). GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite so...
UOJ-Bench uses real competitive-programming submissions to test generation, error-finding, and repair. In single-attempt evaluation, top models fail to identify errors in over 50% of incorrect submissions (arXiv). Test-time scaling pushes success above 90%, but models also fla...
arXiv 2609.20152 targets a measurement gap: end-to-end voice benchmarks mix recognition and model errors into one number, while LLM benchmarks isolate the model and drop the conditions that make phone calls hard (arXiv). The benchmark runs the model with transcription errors p...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.