Fetching from the wire…
Research2026-09-01 · source-backed
APIFlow-Bench tests long-horizon dependent REST workflows with provenance-sensitive grading that traces a minted canary through the data flow to the response the answer must come from (arXiv 2608.29128). Across 19 models, success falls to 74% on clean 20-subtask chains and 61% including flagged trials. Best-case scores spread only 7 points across models while five-of-five reliability spread 44 points, which is the number to quote at anyone benchmarking on best-of-N. On clean chains, 77% of failing runs reached the correct final state and failed at delivery.
Each link below shares sources, entities, or timing with this story.
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.