Research
MathDuels: Adversarial LLM-vs-LLM Evaluation Defeats Static Benchmark Ceiling Effects
Xu et al. introduce MathDuels, a framework that evaluates LLMs as both problem posers and solvers in adversarial mathematical duels. By having models generate novel problems for each other rather than solving fixed benchmarks, the approach circumvents the ceiling effects plaguing static evaluations like MATH and GSM8K. This dual-role evaluation reveals capability gaps invisible to traditional pass-rate metrics and provides a continuously refreshing evaluation signal.
Source
↳ Follow the thread