LeVJEPA Matches V-JEPA 2 at 5.6-20.8x Less Pretraining Compute by Dropping the EMA Encoder and Stop-Gradient
arXiv 2608.27395 trains the first video encoder under LeJEPA's collapse-free objective, dispensing with both the architectural asymmetries (EMA target encoder, stop-gradient, capacity-limited predictor) and pixel-space reconstruction that prior self-supervised video methods rely on. The architecture reduces to an encoder plus projector with a single hyperparameter, and because pretraining cost is governed by tokens observed, uniform random token dropping shrinks that count while improving downstream accuracy. At matched epochs on identical data it matches or beats V-JEPA 2 across ViT-S/B/L at 5.6-20.8x less compute, and at matched FLOPs exceeds the strongest video baseline by 7.6 points on ImageNet-1K.
↳ Follow the thread