Fetching from the wire…
Research2026-09-16 · source-backed
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure profile is the useful part: weaker agents fail before producing a valid artifact at all, while stronger ones increasingly fail on structural, geometric and construction-process requirements. The gap is design intent, not GUI operation.
Each link below shares sources, entities, or timing with this story.
AgentHijack deployed patches on author-controlled GitHub Pages and a locally hosted CSDN clone against five GUI-agent backends. Across 600 online cases: 84.5% success at the VLM output stage, 47.0% at action parsing, 20.3% end-to-end with verifiable environmental consequences....
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
CONFLICTGUI benchmarks conflict-aware termination, covering instructions that contradict themselves and instructions that contradict what's on screen, built on the observation that real users issue infeasible instructions by ordinary mistake. The result is execution-biased ove...
OS-Themis generates structured critiques of GUI agent trajectories rather than binary pass/fail signals, enabling gradient-rich RL training feedback across stochastic environments. Core enabler for next-generation computer-use agents that learn from interaction rather than req...
PCAS: Policy Compiler for Secure Agentic Systems — The first paper to provide measured enforcement results for agent policy compliance (48% to 93%). Uses dependency graphs and Datalog-derived policy language with a reference monitor intercepting all actions. Three case studies...
At 23,817 stars with 2,954 forks and only 20 open issues, MIT-licensed, pushed September 12 after gaining 514 stars that day. It's an open-source architectural modeling editor exposing its own operations through a CLI and MCP server so a coding agent drives the model directly...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.