Fetching from the wire…
Research2026-08-02 · source-backed
Alexi Gladstone, Yilun Du and Heng Ji generate K candidate matches between output and real data at each training step and train only on the best: exploration inside the training loop rather than as RL post-training. 6.2x sample, 4.1x FLOP, 47% better parameter efficiency, with the advantage widening at scale. An Explorative Policy matches Diffusion Policy on robotics with 1 forward pass instead of 100. Single-source blog post, no peer review yet.
Each link below shares sources, entities, or timing with this story.
Shared topic / Tension
Overlapping topics (data, train, training); pushes against this story (but).
Shared entity: FLOP / Earlier coverage
Both cover FLOP; earlier FLOP coverage from 2026-07-20.
Shared entity: ImageNet / Earlier coverage
Both cover ImageNet; earlier ImageNet coverage from 2026-04-29.
Shared topic / Tension
Overlapping topics (best, better); pushes against this story (against).
Overlapping topics (better, data); pushes against this story (but).
Overlapping topics (between, efficiency); pushes against this story (vs).
Overlapping topics (best, data); pushes against this story (but).
Overlapping topics (better, between); pushes against this story (against).