Research
Evolution Strategies Beat GRPO on Pass@K Without Entropy Collapse, and the Gains Live in a Sparse Subset of Weights
This analysis shows Evolution Strategies achieve broader reasoning coverage than GRPO, with verifier-projected Jensen-Shannon diversity across the ES population theoretically tied to higher Pass@K, and empirically ES improves Pass@1 while reaching higher Pass@K where GRPO shows entropy collapse. The authors propose a sequential GRPO-then-ES schedule that keeps GRPO's Pass@1 strength and adds ES's Pass@K gains. They also find that despite large whole-model parameter drift, task gains come from a sparse subset of larger-magnitude updates, held-out evaluations show no catastrophic forgetting, and larger LLMs need a smaller ES population size.
↳ Follow the thread