Fetching from the wire…
Research2026-09-23 · source-backed
arXiv 2609.25911 trains a LinearSVC on PR titles and bot/dependency flags to decide which bumps deserve an agent. It captured 51.4% of repairs within the top 20% of routed PRs, cut calls per captured repair from 6.90 to 2.68, and a 60-case pilot cut diagnosis tokens 66%. Features drawn from full PR history looked better but leak hindsight. Renovate or Dependabot feeding a coding agent should have a small router in front of it.
Each link below shares sources, entities, or timing with this story.
Version-update PRs now wait three days by default, on the reasoning that most malicious package versions get published and pulled inside a short window. Security updates are exempt and fire immediately; the delay is tunable in dependabot.yml. Copy this for any agent that auto-...
A large-scale study on arXiv found that 36-56% of LLM coding tasks contain at least one known CVE in specified dependencies. Not in the generated code itself. In the packages the model tells you to install. The numbers get worse. 62-75% of those CVEs are rated Critical or High...
SWEADV built 750 adversarial issue descriptions from 150 SWE-bench Verified tasks, five per task across command execution, deserialization, path traversal, DoS and weak hashing (arXiv 2609.15963). Across mini_swe agents on GPT-5-Mini, MiniMax-M2.5 and DeepSeek-R, adversarial i...
Three models were evaluated on 832 Defects4J bugs with hallucination tracked across final patches and the intermediate artifacts guiding them (arXiv 2609.04909). Only 21.0% to 55.9% of generated patches passed the developer-written test suite, and manual analysis of 812 sample...
Someone finally measured the thing benchmarks ignore: are the patches any *better*? Four generations of Claude and DeepSeek models on SWE-bench Lite, measured via CodeQL, CodeScene, CPU execution time, and peak memory (arXiv 2607.18462). Newer models resolve more instances. Bu...
Across the AIDev dataset, nearly half of fixes from Copilot, Devin, Cursor, and Claude are rejected, sorted into incorrect implementation, CI failures, inability to execute the fix, and low priority (arXiv). The fix the authors push: better model guidance on implementation app...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.