Fetching from the wire…
Research2026-09-22 · source-backed
OSWorld-Pro decomposes 300+ tasks into over 2,800 sequentially dependent subgoals grounded in 67,000+ human annotations. Claude Opus 5 reaches 75.7% against 83.4% on the original OSWorld. The subgoal traces separate failure modes that final-state scoring can't distinguish, specifically subgoal-irrelevant actions from click-based grounding mistakes, and those need different fixes. (arXiv 2609.24890)
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
arXiv 2609.17598 studies PRs from OpenAI Codex, Devin, GitHub Copilot, Cursor and Claude Code across 2,807 repositories (Dec 2024 to Jul 2025), combining AIDev with 58,792 cached GitHub API responses. Codex PRs were reverted 6.1% of the time against a human baseline of 11.5% (...
Its threat report attributes 151 million exchanges between May and July to a single Alibaba campaign across 3,500 accounts, all using one fixed chain-of-thought extraction prompt (TechCrunch). It counts more than 12 million DeepSeek exchanges over 14 days. It also says Moonsho...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Cursor stopped being an IDE wrapper and became a model company. Cursor shipped Composer 2, a proprietary coding model trained via reinforcement learning on long-horizon coding tasks. On CursorBench — their own benchmark, caveats acknowledged — it scores 61.3, beating Claude Op...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.