RLVR Works for a 0.8B Search Agent Without Distillation, but the Default Exact-Match Reward Is the Worst Choice
arXiv·low signal
Jayadas, Plaat, Serra-Gómez et al. (arXiv 2609.28765) train Qwen3.5-0.8B with GRPO and an interleaved Wikipedia search tool on MuSiQue, varying only the reward shape across three seeds. The best run reached 0.352 average exact match on a seven-benchmark QA suite, against a 0.092 untrained floor. The Search-R1-style exact-match-only reward was the worst of the three at every seed, even on exact match itself.