Fetching from the wire…
Research2026-08-21 · source-backed
Using a single accepted program's outputs as ground truth for LLM-generated tests is standard practice. On external inputs where three accepted implementations agree, generated outputs match the panel only 27.79% and 50.12% of the time. Correct for the inflation and equal-budget independent resampling beats mutation-based evolution by 6.01 to 18.83 points, and a real three-round feedback loop is statistically indistinguishable from a density-matched placebo. arXiv Add this to the pile of self-improvement results that don't survive a proper null.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-16.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-17.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-12.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-07.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-05.