Skills
Test-time scaling fails on open-ended work because selection is near-random, not because candidates are bad
Across five test-time scaling methods and five benchmarks spanning medicine, law, finance, chat, and creative writing, candidate generation kept improving with compute in every domain but reward models correlated with actual quality at roughly rho 0.12, making selection near-random at any budget. Only candidate fusion consistently beat a single-sample baseline, recovering about 40% of the available quality gain. The practical move is to stop spending on best-of-N with a reward-model picker for subjective tasks and to synthesize across candidates instead.
↳ Follow the thread