Stop bolting self-critique onto your agent loop: across 36 paired comparisons, Self-Refine and Reflexion never beat plain repeated sampling at equal token cost
A designed experiment ran seven reflection-style methods against a repeated-sampling baseline on 1.5B/3B/7B open models, two math benchmarks, 150 questions each, counting every token spent on critiques, debate turns, and checking. No method was reliably better than repeated sampling at matched cost anywhere; ten were reliably worse, and all 18 self-inspection comparisons came out negative. Reflexion as published never triggered its own retry on the smallest model — it judged itself correct every time and silently degraded into a single chain of thought. Caveat for builders: this covers 1.5B–7B open models, not frontier models, but it means any self-critique layer you add needs a token-matched control before you believe it.
↳ Follow the thread