Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.20822 pairs each robot manipulation task with a forbidden obstacle and finds the agent collides with it in most cases. The diagnosis generalizes well past robotics: the agent reasons about the obstacle in its traces and the prompt forbids touching it, so neither perception nor instruction failed. The stated constraint simply never became a planning priority. SafeHarness grounds objects as bounding boxes, draws candidate waypoint routes, plans-verifies-replans before executing, and reaches 71.9% task success and 87.5% collision avoidance against the same agent's 31% and 58%. arXiv
Each link below shares sources, entities, or timing with this story.
Traced across 557 SWE-chat sessions (94,813 events) and 33,097 agentic pull requests from AIDev. Agent-facing artifacts account for 60.5% of documentation interactions versus 10.6% for classical technical docs and 1.3% for API references. Consultation is self-initiated 70.2% o...
DEPBENCH collects 203 real-world upgrade tasks across five package ecosystems and five language communities, each containing code-level breakage the upgrade never announced (arXiv 2608.30300). The best completed agent configuration solves 104 of 203, with wide variation across...
arXiv 2608.09885 treats the harness as the thing that evolves with emerging risk rather than a static wrapper around a model you keep re-aligning. Four artifacts with non-overlapping responsibilities: System Prompt, Rule Bank, Safety Memory, Tool Policy. Failures get attribute...
CodeGrep measures a 30B OpenHands agent averaging 23 rounds and 631K tokens per resolved SWE-Bench Verified issue, much of it grep, glob and view_file. A 14B retrieval agent trained end-to-end with GRPO raises resolve rate to 27.0% from 25.8% while cutting 15% of rounds and 19...
An LLM-based multi-agent simulation seeded with real GitHub data from 1,084 active developers branched the same community state into no-agent and agent conditions for 4-week runs (arXiv 2608.03585). Planned tasks up 34%, completed up 39%, median completion time from 45 to 20 m...
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.