Fetching from the wire…
Agents2026-09-04 · source-backed
CONFLICTGUI benchmarks conflict-aware termination, covering instructions that contradict themselves and instructions that contradict what's on screen, built on the observation that real users issue infeasible instructions by ordinary mistake. The result is execution-biased overcompliance, and it correlates with feasible-task performance. CONFLICTGUARD is inference-time only, pairing a feasibility verification protocol with conditional action modulation, and lifts conflict-task success across five agents without hurting normal performance. arXiv 2609.03438
Each link below shares sources, entities, or timing with this story.
PCAS: Policy Compiler for Secure Agentic Systems — The first paper to provide measured enforcement results for agent policy compliance (48% to 93%). Uses dependency graphs and Datalog-derived policy language with a reference monitor intercepting all actions. Three case studies...
Alibaba Tongyi Lab's technical report describes a foundation GUI agent spanning mobile, computer-use, web and DeepSearch, with a unified action space interleaving GUI operations with CLI execution and emitting batched actions per model turn. 82.1% MobileWorld, 92.2% MobileWorl...
The PlayCoder benchmark is the cold shower the vibe coding movement needed. Researchers tested 10 state-of-the-art code LLMs on generating GUI applications across six categories. The models achieved high compilation rates. The code built and ran. But when they measured whether...
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
CVE-2026-84885, -84886 and -84887 published September 3 against 0.3.1/0.3.2, covering code_agent.py, the OCR HTTP API's ImageData handler via img_bytes, and the model-generated GUI action execution workflow in grounding.py. Same 48-hour window produced the same non-response pa...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.