Fetching from the wire…
Public story · 2026-02-28 · source-backed
Hybrid on/off-policy RL giving agents non-parametric memory for exploration. 128.6% improvement over GRPO on ScienceWorld, 11.3% on WebShop. Agents generalize to out-of-distribution tasks with "only a few trials with memory and no parameter updates." (arXiv)
Each link below shares sources, entities, or timing with this story.
SkillPyramid benchmarked against ScienceWorld / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (SkillPyramid benchmarked against ScienceWorld); both cover ScienceWorld, WebShop; reported by the same outlet (arxiv.org).
GRPO competes with PPO / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (GRPO competes with PPO); both cover GRPO, Memory; reported by the same outlet (arxiv.org).
Unsloth released GRPO / Shared entity: GRPO / What happened next / Tension
Linked by a graph relationship (Unsloth released GRPO); both cover GRPO; picks up the GRPO thread on 2026-06-20.
Shared entities / Same source domain / Shared topic / What happened next
Both cover GRPO, WebShop; reported by the same outlet (arxiv.org); overlapping topics (agent, grpo, over).
DeepSeek-R1 uses GRPO
Linked by a graph relationship (DeepSeek-R1 uses GRPO).
Microsoft released EMPO2 / Shared entity: Memory / Shared topic / What happened next
Linked by a graph relationship (Microsoft released EMPO2); both cover Memory; overlapping topics (agent, memory).
EMPO2 benchmarked against ScienceWorld / Shared entity: ScienceWorld / Same source domain / What happened next
Linked by a graph relationship (EMPO2 benchmarked against ScienceWorld); both cover ScienceWorld; reported by the same outlet (arxiv.org).
Shared entity: Memory / Same source domain / Shared topic / What happened next / Tension
Both cover Memory; reported by the same outlet (arxiv.org); overlapping topics (agent, memory).