Enriching the Environment's Feedback Beats Warming Up the Agent for Sparse Long-Horizon RL
Rather than agent-side warm-up via supervised fine-tuning, which is limited by data scarcity and constrained exploration, this work builds Feedback-Enriched Environments that shift from action guidance to observation enrichment in the later stages of both intra-episode exploration and inter-episode evolution. Experiments on SciWorld and BFCL across multiple Qwen3 scales with GRPO, GSPO and DAPO show consistent improvement over standard settings. Analysis reports that FEEs stabilize training by reducing entropy volatility, encourage proactive state-space exploration on hard tasks, internalize environmental guidance into policy weights rather than acting as an inference-time prior, and identify intra-group feedback consistency as the boundary condition for stable optimization.
↳ Follow the thread