Agents
"One Run Is Not an Idea": winner reversal hits 43.6% when automated research judges an idea by a single implementation
arXiv 2607.26587 (July 29) names the "implementation lottery" — automated research systems score one implementation of an idea, then credit that realization-level score as evidence about the parent mechanism. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance exceeded same-artifact rerun variance by more than 5x and 10x respectively, and the winner from one implementation draw differed from the other-two mean in 25.6% and 43.6% of decisions. Reversal survived card-level filtering under two outcome-blind review rules, which is a direct warning for anyone running agent-driven experiment loops that promote ideas on single-run scores.
Source
↳ Follow the thread