Fetching from the wire…
Agents2026-07-26 · source-backed
Chen et al. extend AppWorld into a 516-task user-in-the-loop benchmark across nine simulated apps, injecting ambiguities and constraints that force the agent to ask clarifying questions, request confirmation, or declare a task infeasible (arXiv 2607.20536). Opus 4.7 gets 48.6% overall, 35.7% compositional, 21.3% under the stricter scenario-level metric. User behavior is simulated by an LLM with designed knowledge boundaries rather than the unconstrained simulators prior work used. The analysis finds correct interaction, not tool execution, determines success, which is a different failure axis than most agent evals measure.
Each link below shares sources, entities, or timing with this story.
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
A solo Claude Opus 4.5 agent spent $9 and 20 minutes building a retro game. It was broken. The same model, wrapped in Anthropic's multi-agent harness, spent $200 over 6 hours and produced a fully playable game with physics, sprite editors, and AI integration. Anthropic's engin...
I've been saying for months that the real gains aren't in switching models. They're in how you set up the environment around the model. Now there's quantitative proof. Stanford IRIS Lab published Meta-Harness, a system that autonomously evolves its own coding harness, system p...
The ARC-AGI-2 breakthroughs reveal a concrete architectural pattern. Symbolica's Agentica achieves 85.28% (with Opus 4.6) using recursive delegation where sub-agents spawn sub-agents, each receiving only relevant state and avoiding context rot. Average 2.6 agents per task, max...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.