CIPO gives turn-level credit for actually using retrieved evidence, attacking confirmation bias in search agents
Search-agent RL usually rewards only final-answer correctness or vague intermediate progress, which quietly trains prior-driven reasoning: the agent decides from parametric memory and uses retrieval to confirm itself. CIPO (arXiv 2608.06128, submitted 2026-08-06) assigns dense turn-level credit specifically to reasoning actions that were influenced by retrieved information, combined with a global outcome reward to preserve correctness, and needs neither human process annotations nor a separate reward model. Across seven in-domain and out-of-domain benchmarks it reduces the prevalence of prior-driven reasoning while performing well on most tasks. For research-agent builders this is the clearest statement yet that 'it cited a source' and 'the source changed its answer' are different things worth measuring separately.
Source
↳ Follow the thread