Explorative Modeling Trains on the Best of K Guesses, Claiming 6.2× Sample and 4.1× FLOP Efficiency on ImageNet
Alexi Gladstone, with Yilun Du and Heng Ji, published the approach on July 29 (107 points on HN August 1): at each training step the model generates K candidate matches between its output and the real data, and only the best match is trained on — exploration inside the training loop rather than as a separate RL post-training stage. Reported gains are 6.2× sample efficiency, 4.1× FLOP efficiency and 47% better parameter efficiency on ImageNet, with the advantage widening at scale (7%→36% with more data, 13%→23% with more parameters). Downstream, an Explorative Policy matches Diffusion Policy on robotics with 1 forward pass instead of 100, and an Explorative World Model matches scores with 16–256× fewer forward passes. Single-source blog post, no peer review yet.
↳ Follow the thread