DiagEvo Derives a Self-Play Curriculum From the Solver's Own Failure History Instead of External Task Data
Unguided self-play steers question generation by difficulty, learnability, or diversity, which keeps questions varied but never says which unresolved weakness the next round should target, while guided methods import that direction from human examples or document corpora outside the loop. DiagEvo instead extracts recurring error causes from the solver's failure history into a hierarchical memory that groups causes under skill nodes and marks each Active or Mastered by self-consistency, with double-confidence filtering keeping intermediate-difficulty questions. With a 4B diagnostician it beats every baseline in mean accuracy across all nine benchmarks for three solvers, reaching 72.3% on five math benchmarks with Qwen3-8B, 4.5 points above R-Zero.
↳ Follow the thread