Fetching from the wire…
Research2026-09-10 · source-backed
IdeaAMBIG holds 660 evidence-grounded instances, 163 real gaps mined from reproducibility reports and GitHub issues plus 497 synthetic injections. Across 13 LLMs, best-case defect recovery on real instances is 9.6% while clarification-action success once handed the annotated defect is 80.6%, and an oracle study lifts downstream codification-readiness from 14% to 98%. Localization is the bottleneck, not repair, which is the same shape as the SWE-bench localization work and argues for spending review attention on "what's missing from this spec" rather than "is this implementation right."
Each link below shares sources, entities, or timing with this story.
The July 30 changelog closed the hosted model-catalog and playground service that let developers prototype against multiple LLMs from GitHub directly. If you prototyped against Models endpoints, this is a migration event, not a skim. The surrounding changelog items (Copilot up...
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
This comparison ran the baseline the retrofitted-linear-attention literature skipped. Across multiple LLMs and downstream tasks SWA with sinks matches or beats post-trained linear attention, and on Needle-in-a-Haystack and BABILong it scores 2 to 10 times higher. The recommend...
A new paper on "Contagion Networks" (arXiv:2606.20493) shows that when LLMs serve as evaluators inside multi-agent systems, their systematic biases propagate through the network rather than staying local. A single biased judge can contaminate downstream agent decisions. If you...
GLM-5.1 scored 58.4% on SWE-Bench Pro. Opus 4.6 scored 57.3%. GPT-5.4 scored 57.7%. Read those numbers again. An open-weight, MIT-licensed model now leads the most rigorous coding benchmark we have. This isn't a narrow win on a cherry-picked eval. SWE-Bench Pro tests real-worl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.