A search agent's critic has to co-evolve with it, because improving either half alone plateaus
CAFE (arXiv 2608.24794, 2026-08-25) makes corrective feedback an in-trajectory intervention the agent chooses to request, using one shared-parameter model that alternates between search-agent and critic roles. Online RL shapes request returns from a prompt-level call-versus-skip success gap and reweights token advantages before and after feedback; offline preference optimization learns feedback from matched successful and unsuccessful trajectories. It beats the evaluated RL-based search agents on seven agentic search benchmarks, holds gains across all six out-of-domain ones, and reduces answer-level hallucinations. The finding builders should take is the ablation: improving only the agent or only the critic eventually plateaus, while alternating updates keeps improving.
↳ Follow the thread