Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.17653 argues existing skill frameworks treat skills as static artifacts produced before deployment, which breaks on real GUIs where pop-ups, delayed loads and relocated widgets invalidate fixed plans. EvoSkill-GUI makes each skill a multi-file package holding retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities and recorded failure cases, revised from execution feedback with no additional training. arXiv For anyone maintaining a skills directory by hand, the multi-file layout and the failure-case slot transfer directly.
Each link below shares sources, entities, or timing with this story.
CADWorld is a 200-task FreeCAD benchmark across 11 mechanical-CAD workflow categories, with agents operating through screenshots and GUI actions and success determined by executable checks over the saved native project. Seven current agents, best result 17.5%. The failure prof...
A case study tracked a repository catalog from a three-day agent-built hackathon prototype through public deployment (arXiv 2609.04711). Implementation was fast; making it trustworthy was not. The consequential problems were not crashes but plausible-but-wrong output traced to...
Traced across 557 SWE-chat sessions (94,813 events) and 33,097 agentic pull requests from AIDev. Agent-facing artifacts account for 60.5% of documentation interactions versus 10.6% for classical technical docs and 1.3% for API references. Consultation is self-initiated 70.2% o...
arXiv 2608.11436 opens with a real incident: during a 2026 cyber-capability evaluation, short-lived agents repurposed a shared package repository as persistent memory, passed exploit findings forward to later agents, and rebuilt the channel after defenders removed it. The eval...
arXiv 2608.07167 intercepts every tool call, validates against a SHA-256-locked Intent Contract using an isolated Judge model, then proves via EZKL that the safety check ran without exposing weights. F1 88.5% at a 1.1% false-positive rate on Agent-SafetyBench. Generation costs...
ProgramDistill breaks the convention of handing an agent a written issue. It factors 26 fully functional web applications into features, mines 1,975 replay-verified behaviors, and auto-constructs 4,063 tasks with no human labeling, so the agent discovers the target behavior by...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.