Fetching from the wire…
Research2026-09-17 · source-backed
arXiv 2609.17598 studies PRs from OpenAI Codex, Devin, GitHub Copilot, Cursor and Claude Code across 2,807 repositories (Dec 2024 to Jul 2025), combining AIDev with 58,792 cached GitHub API responses. Codex PRs were reverted 6.1% of the time against a human baseline of 11.5% (OR 0.50); Devin hit 14.5% (OR 1.31). Agent code pooled across vendors was less likely than human code to carry a security smell (OR 0.63), driven by fewer hardcoded credentials and eval-style constructs. Review effort concentrated unevenly: Copilot PRs drew the most human reviews and change requests, Claude Code PRs waited longest for a first human review at a 12.6 hour median. Quality here is vendor-specific, not a property of "AI code," which kills most of the arguments people have about this. Pick by measured revert rate on your repo, and budget review latency separately from generation quality.
Each link below shares sources, entities, or timing with this story.
We've been operating on faith here. Everyone tells you to write an AGENTS.md or CLAUDE.md, you write one, and you assume it helps because it feels like it should. Now there's data, and it's more interesting than "yes, write the file." A study of 15,549 agentic pull requests ac...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
Across the AIDev dataset, nearly half of fixes from Copilot, Devin, Cursor, and Claude are rejected, sorted into incorrect implementation, CI failures, inability to execute the fix, and low priority (arXiv). The fix the authors push: better model guidance on implementation app...
Researchers analyzed 61,837 GitHub Actions runs from 2,355 repos triggered by PRs from Claude, Devin, Cursor, Copilot, and Codex. Substantial differences in pass rates across bots. This is the first empirical data on how AI-generated code actually performs under real CI/CD con...
Barry Zhang and Mahesh Murag, the engineers who built Claude Skills at Anthropic, published a talk and engineering post that's gotten 14K+ likes and is reshaping how I think about agent development. The core argument: most agent approaches fail because they lack domain experti...
First major enterprise observability platform to ship a production-grade MCP server. Feeds live logs, metrics, and traces directly into Claude Code, Cursor, Codex, GitHub Copilot, and VS Code. AI coding agents can now investigate production issues using real-time telemetry. MC...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.