Fetching from the wire…
Agents2026-09-21 · source-backed
GameASG-Bench builds 47 browser game-generation tasks across 12 genres, each with an evaluation interface declared before generation. Across nine agent stacks the highest mean runtime check pass rate is 93.2%, but the highest strict task success, requiring every applicable check, is 55.3%. Averaged check rates hide task-level compliance, and the two tested harnesses each got 18 strict successes while overlapping on only ten tasks. Same harness, same model class, different tasks solved. Reasoning effort didn't behave monotonically.
Each link below shares sources, entities, or timing with this story.
This is the cleanest experimental result I've seen in weeks, and it explains a class of debugging pain I've hit personally. Researchers held weights, test cases, decoding parameters and seeds fixed on BFCL v4 and changed exactly one thing: the serving adapter. The tool-call sc...
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Cline released @cline/sdk on May 13, an open-source TypeScript agent runtime that powers their CLI, VS Code, and JetBrains extensions. Running claude-opus-4.7, Cline CLI scores 74.2% on Terminal-Bench 2.0. Claude Code on the same model: 69.4%. Same model. Different harness. Al...
40+ experts and 31,000+ human hours built 60 project-level research tasks across 11 domains, designed so guidance is progressively withdrawn *within the same project*. Across 18 agent-model configurations: 50.91 with full methodological guidance, 29.10 when only the method is...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.