Fetching from the wire…
Research2026-09-18 · source-backed
arXiv 2609.19843 ran a randomized online shopping experiment with 3,600 GUI agents and 21,600 simulations across six frontier models from three providers. Agents fell for both automatic and reflective interface nudges, and extended reasoning moved the two in opposite directions, cutting susceptibility to automatic default nudges while raising it to reflective social-influence nudges. More reasoning didn't produce a more robust agent, it changed which manipulation worked.
Each link below shares sources, entities, or timing with this story.
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp (arXiv 2608.03327). The authors call it the "adoption gap": the reasoning model used a tool on just 55 of 309 ta...
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 1...
arXiv 2607.29199 tests three frontier GUI agents under screen-grounded, user-side persuasion, with no environment injection at all. A single-line guardrail cuts attack success rate by ~40 points in single-turn scenarios. Four-turn escalation chains push guarded ASR back up by...
The PEEU paper (arXiv:2606.27330, ACL 2026 Main) has a small multimodal model autonomously explore GUI environments and use hindsight to synthesize high-level training data. The 7B reaches 30.6% accuracy, surpassing the much larger 32B, and high-level task training drove stron...
A paper from Stanovsky's group (arXiv:2606.16576) tests whether LLM agents can uncover a hidden deterministic finite automaton through membership and equivalence queries. Performance drops sharply as the automaton grows, and trajectory analysis exposes recurring failures in qu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.