Research
When Good Verifiers Go Bad: self-improving VLMs can regress on new tasks
This paper (arXiv 2606.14629, cs.CR/cs.AI) shows that verifier-driven self-DPO — a common recipe for self-improving production vision-language models where a frozen verifier scores candidate generations and the top picks are reinforced — can cause measurable regression when the model faces new tasks. It is a direct cautionary result for anyone running self-improvement, RLAIF, or self-DPO loops in production, since the frozen verifier silently entrenches narrow behavior.
Source
↳ Follow the thread