Fetching from the wire…
Skills2026-07-30 · source-backed
A placebo-controlled July 28 study found blind resampling beats self-repair at 2.5-5.5x lower token cost on MBPP+, because showing a model its own failed attempt makes it reproduce a near-identical program 33-68% of the time versus 2-14% under blind resampling. Real execution feedback added nothing over a content-free failure notice. Tested only up to 7B, so treat it as a hypothesis to test on frontier models, not settled.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover July, Tested; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-23.
Shared entity: July / Same source domain / Shared topic / Earlier coverage / Tension
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (beat, code).
Shared entity: July / Same source domain / Shared topic / Earlier coverage
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (added, code).
Shared entity: July / Shared topic / Earlier coverage
Both cover July; overlapping topics (beat, blind, code, model); earlier July coverage from 2026-07-21.
Shared entity: July / Same source domain / Shared topic / Earlier coverage
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (execution, model).
Shared entities / Same source domain
Both cover July, MBPP; reported by the same outlet (arxiv.org).
Shared entity: July / Same source domain / Earlier coverage / Tension
Both cover July; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-28.
Both cover July; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-26.