Fetching from the wire…
Agents2026-09-17 · source-backed
ERPBench evaluates six screenshot-only computer-use agents against a live reproducible ERP system, scoring against ground-truth database values rather than screen state. Strong general GUI performance does not transfer. The agents reach the right form and save it; the stored record is frequently wrong. That 85%-to-3% spread is the single most damning agent number I've read this month, and it exists only because someone checked the database instead of the screenshot. Every computer-use eval that scores on final screen state is measuring the wrong thing.
Each link below shares sources, entities, or timing with this story.
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure prof...
AgentHijack deployed patches on author-controlled GitHub Pages and a locally hosted CSDN clone against five GUI-agent backends. Across 600 online cases: 84.5% success at the VLM output stage, 47.0% at action parsing, 20.3% end-to-end with verifiable environmental consequences....
Only about half an LLM's accuracy advantage reaches the human consulting it, according to one study this week's companion work on advantage transfer. Pair that with ERPBench's 85%-save/3%-correct gap and AutoTuneBench's 10.6x-to-2.03x correction, and the pattern for the week i...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Salesforce AI Research argues in arXiv 2607.22798 that a screenshot is a lossy rendering of program state — different states produce identical pixels — so the main agent should manipulate files, backends, and the DOM through code, delegating to a GUI subagent only when necessa...
OS-Themis generates structured critiques of GUI agent trajectories rather than binary pass/fail signals, enabling gradient-rich RL training feedback across stochastic environments. Core enabler for next-generation computer-use agents that learn from interaction rather than req...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.