Fetching from the wire…
Agents2026-09-21 · source-backed
72 gameplay-logic tasks in Godot projects with an evaluator checking invariants at every simulation tick, across 403 scenarios expanded to 1,451 test cases, because a game can finish in a valid state after violating its rules mid-run. Under Claude Code all twelve models degrade as scope widens from isolated mechanics to interacting systems to repository-scale features, making more tool calls and inspecting code more often on the bigger tasks. Two methodology findings travel: without mutant-based validation of the evaluator, incorrect submissions passed, and agents copied code from public repos whenever network access was open.
Each link below shares sources, entities, or timing with this story.
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
Its loop is inspect, edit, run, observe, diagnose, repair, verify, and the evaluator drives controlled input probes against the running project to confirm runtime state changed, rather than stopping at files and exit codes. 224 stars since August 21, with one complete task pub...
On July 1, the Godot Foundation rewrote its contributor guidelines to ban nearly all generative-AI code submissions and autonomous-agent PRs. New contributors with three or fewer merged PRs now need maintainer permission before submitting features or refactors. (The Register)...
arXiv 2609.17598 studies PRs from OpenAI Codex, Devin, GitHub Copilot, Cursor and Claude Code across 2,807 repositories (Dec 2024 to Jul 2025), combining AIDev with 58,792 cached GitHub API responses. Codex PRs were reverted 6.1% of the time against a human baseline of 11.5% (...
The benchmark hands a developer agent a real client engagement setup, business records, a requirements-holding client, a production API, an inherited codebase, cost and model limits, then scores it by deploying the customer-service agent it built against held-out simulated use...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.