Fetching from the wire…
Research2026-09-13 · source-backed
arXiv 2609.10590 targets the gap where requests come from end users in non-technical language and a developer has to translate them into specs first. ReqEvolve chains automatic requirements engineering with TDD: clarification, specification decomposition, test generation, runtime integration. Across 72 evolution cases in 18 projects it reaches 89.2% pass@1, beating the RE-focused baseline SpecFix by 18.8% (p<0.01, r=0.79) and its own ablation by 32.6%. The ablation delta is the honest signal that the RE stage is doing work and not just adding tokens.
Each link below shares sources, entities, or timing with this story.
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
arXiv 2607.23665 uses differential profiling to find time-critical structures, then routes them to specialized strategies — mixture-of-experts, but over prompts — across four abstraction levels from algorithmic down to API usage. Up to 57.48% opt% on COFFE and Effibench (corre...
The standard criterion for LLM-generated bug reproduction tests, fails on buggy code and passes on the golden fix, turns out to be insufficient. Many F→P tests are "lax": they reproduce the symptom while still admitting plausible-but-wrong patches. Worse, co-generating the tes...
A 400-problem framework built by injecting AST-level corruptions into BigCodeBench reference solutions, so every repair task has a known minimal patch. Adding a preservation instruction lowered average excess Levenshtein distance from 0.195 to 0.131, cut added cognitive comple...
No compiled code. No runtime. No dependencies. Just 18 markdown files in a .claude directory. Matt Pocock's skills repo hit 75,700 stars this week, gaining 6,400 in seven days. It's the #1 trending AI repo on GitHub. The repo solves four problems that every Claude Code user hi...
42. Vellum AI — Opus 4.6 Benchmarks 43. Arcee AI — Trinity Large 44. arXiv — PCAS 45. arXiv — SpargeAttention2 46. arXiv — CUWM 47. arXiv — VESPO 48. arXiv — SAGE 49. The Register — TDD for AI Agents 50. Latent Space — Anita TDD 51. builder.io — TDD + AI
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.