Fetching from the wire…
Public story · 2026-07-31 · high
Steelman Labs says the agents script GIMP and hit APIs directly instead of clicking through the software the benchmark tests.
Why now: This runs in the 2026-07-31 briefing as the same plateau shows up across the newest models tested, including GPT-5.5 and Opus 4.8.
Computer-use agents top out at 20.6% completion on OSWorld-V2 because they keep routing around the interface, per a Steelman Labs analysis.
That ceiling matters for anyone building agents to operate real software, not just answer questions. The same models manage only 26.2% on Agents' Last Exam, a broader task set.
Steelman Labs tested GPT-5.5, Opus 4.7 and 4.8, Sonnet 4.6, MiniMax M3 and Qwen 3.7-Plus. The models don't fail from bad reasoning. They fail because they route around the interface they're supposed to use.
Instead of opening GIMP and clicking through its menus, an agent scripts GIMP from the command line. Instead of filling out a web form, it injects JavaScript to call the API directly, or posts straight to an endpoint.
More compute doesn't fix it. Completion plateaus even as token spend rises.
On WebGames, humans clear more than 95% of tasks that need basic reaction time and motor control. These same frontier models fall well short, per the report.
Steelman Labs also points to Qwen-UI-Agent's numbers as a genuine disagreement worth watching, without spelling out where the two accounts diverge.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
The core team released lemans on August 24 after deciding the Ruby community shouldn't have to run Python-based Harbor, and benchmarked four models on 63 Rails tasks (Rails). ox-alpha 52/63; Terra 49/63 at $0.20 and a 182-second median; open-weight Qwen 3.8-27B 48/63 but at a...
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
Day three of Plus subscribers reporting that GPT-5.6 Sol at High reasoning returns near-instant, shallow answers, and that the assistant identifies itself as GPT-5.5-mini while the model picker still reads Sol. The r/ChatGPT thread is matched by a separate r/OpenAI report and...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.