Fetching from the wire…
01
02
03
04
05
06
07
08
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
EMPO2 achieved 128.6% improvement over GRPO baseline on ScienceWorld benchmark.
Source findingMicrosoft released EMPO2 combining on/off-policy reinforcement learning with external memory for agents.
Source findingEMPO2 memory-augmented agent achieves 128.6% improvement over GRPO on ScienceWorld with 7 times better generalization to out-of-distribution tasks
Source findingMicrosoft released EMPO2 combining on/off-policy reinforcement learning with external memory for agents.
Source finding