Skills
A self-critique loop with a CONFIRMED early stop ends 82-88% of items at equal accuracy, averaging 2.1 generations
EvoResearcher runs generate, self-critique, revise on a frozen model with four meta-rewards for correctness, efficiency, reflection depth, and tool-call diversity, all expressed purely as prompts with no gradient updates. Tested on Big-Bench Hard, GSM8K, and MATH with Qwen2.5-72B, the CONFIRMED early stop terminated 82-88% of items at the same accuracy, costing about 2.1 generations per question. Note the honest negative result, accuracy did not improve over baseline, so the value here is bounding reflection cost rather than raising quality.
↳ Follow the thread