Fetching from the wire…
Research2026-09-02 · source-backed
Three arms on the same failed candidate: blind whole-solution resampling, spectrum-based localization then suspect-span infilling, and same-length infilling at a random disjoint span. Across three frozen 26-32B models and 488 failing candidates, localization was available on only 9.0% of failures, and where available it lost to blind resampling at matched attempt count 3:40 (p = 3.0e-9), replicating at -11.3 points in a fourth model (arXiv 2609.00854). Sixteen localized attempts reach 6.8%; one blind attempt reaches 10.1%, because infilling reproduces the removed span verbatim 48.9% of the time. The placebo arm is what makes this convincing.
Each link below shares sources, entities, or timing with this story.
A placebo-controlled July 28 study found blind resampling beats self-repair at 2.5-5.5x lower token cost on MBPP+, because showing a model its own failed attempt makes it reproduce a near-identical program 33-68% of the time versus 2-14% under blind resampling. Real execution...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
FaulT-Bench runs 200 scenarios across eight topologies, five reimplemented from public practitioner labs, spanning genuine faults, false fault reports, wrong device attribution and wrong root-cause claims. SADE, ReAct and Claude Code are all near-saturated on accurate tickets...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
arXiv 2608.23873 starts from a structural observation I hadn't seen framed this cleanly: the serving stack knows which span is user input, tool output or instruction, but the model sees only tokens and infers span identity from text the attacker controls. Semantic Overlays are...
Diffusion LMs decode many tokens per step but pay to interact with all suffix tokens every step, and existing fixes just keep a local window while re-initializing suffix tokens identically each timestep (arXiv 2608.23167). This method splits the suffix into local, middle and t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.