Fetching from the wire…
Public story · 2026-07-31 · high
Steelman Labs says the agents script GIMP and hit APIs directly instead of clicking through the software the benchmark tests.
Why now: This runs in the 2026-07-31 briefing as the same plateau shows up across the newest models tested, including GPT-5.5 and Opus 4.8.
Computer-use agents top out at 20.6% completion on OSWorld-V2 because they keep routing around the interface, per a Steelman Labs analysis.
That ceiling matters for anyone building agents to operate real software, not just answer questions. The same models manage only 26.2% on Agents' Last Exam, a broader task set.
Steelman Labs tested GPT-5.5, Opus 4.7 and 4.8, Sonnet 4.6, MiniMax M3 and Qwen 3.7-Plus. The models don't fail from bad reasoning. They fail because they route around the interface they're supposed to use.
Instead of opening GIMP and clicking through its menus, an agent scripts GIMP from the command line. Instead of filling out a web form, it injects JavaScript to call the API directly, or posts straight to an endpoint.
More compute doesn't fix it. Completion plateaus even as token spend rises.
On WebGames, humans clear more than 95% of tasks that need basic reaction time and motor control. These same frontier models fall well short, per the report.
Steelman Labs also points to Qwen-UI-Agent's numbers as a genuine disagreement worth watching, without spelling out where the two accounts diverge.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover GPT, Opus, Qwen; overlapping topics (agent, capability, model).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover GPT, Opus, OSWorld; overlapping topics (agent, capability, model).
Linked by a graph relationship (Claude Code uses Opus); both cover GPT, Opus, Python; overlapping topics (model, number).
Anthropic criticizes Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic criticizes Qwen); both cover Opus, Python, Sonnet; overlapping topics (agent, model).
Qwen benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover GPT, JavaScript, Python; overlapping topics (agent, model).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Agent, GPT, Qwen; overlapping topics (agent, model).
Computer uses Gemini / Shared entities / Shared topic
Linked by a graph relationship (Computer uses Gemini); both cover Agent, GPT, Opus, OSWorld; overlapping topics (agent, computer use).
Qwen benchmarked against Claude / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Qwen benchmarked against Claude); both cover Agent, Opus, Sonnet; earlier Agent coverage from 2026-05-17.