Fetching from the wire…
Agents2026-07-18 · source-backed
arXiv 2607.15193 uses a planner-executor split that externalizes plans and replanning as persistent, inspectable, revisable objects, with users making localized corrections via editable plans plus screenshot-grounded intervention. The claim worth testing: a large share of GUI agent failures are structurally repairable if the plan stays visible and intervention can be scoped narrowly instead of restarting the run. Anybody who's watched a browser agent go off the rails at step 7 of 30 knows why that matters.
Each link below shares sources, entities, or timing with this story.
MobileWorldSafety (arXiv 2608.17659) embeds injection attacks in Android apps and evaluates with a two-stage pipeline combining rule-based verification with LLM adjudication, specifically to separate safety failures from capability failures. That distinction is the contributio...
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
arXiv 2608.04755 injected Android permission popups into real GUI tasks across four frontier multimodal LLMs with synchronized screenshots and UI trees. Holding the task fixed and changing only the requesting app flipped grants from 26/32 to 0/32, an App-Trust Bias. Holding th...
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp (arXiv 2608.03327). The authors call it the "adoption gap": the reasoning model used a tool on just 55 of 309 ta...
arXiv 2607.29199 tests three frontier GUI agents under screen-grounded, user-side persuasion, with no environment injection at all. A single-line guardrail cuts attack success rate by ~40 points in single-turn scenarios. Four-turn escalation chains push guarded ASR back up by...
Alibaba Tongyi Lab's technical report describes a foundation GUI agent spanning mobile, computer-use, web and DeepSearch, with a unified action space interleaving GUI operations with CLI execution and emitting batched actions per model turn. 82.1% MobileWorld, 92.2% MobileWorl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.