Self-Trained Verification for Training- and Test-Time Self-Improvement
arXiv·high signal
Addresses the verifier bottleneck that stalls both test-time verification-refinement loops and training-time self-training. Proposes verification methods that prevent score inflation and provide actionable feedback, enabling genuine self-improvement in reasoning models. Directly relevant to anyone building self-improving agent loops.