UniJEPA Unifies Image and Video World Modeling in One Latent Space, Planning Tens of Times Faster Than Generative World Models
JEPA research has fragmented into incompatible recipes — I-JEPA predicts masked image parts, Image World Models predict photometric transformations, and V-JEPA 2 / DINO-World / DINO-WM predict future temporal states — each with separate encoders, predictors, and anti-collapse regularizers. UniJEPA learns photometric and temporal prediction jointly in one shared latent space under a single end-to-end objective (next-embedding loss plus a Gaussian regularizer) that is provably anti-collapse and trains from raw pixels with no EMA, stop-gradient, or pretrained encoder. It matches or surpasses task-specific JEPAs with one loss hyperparameter and plans up to tens of times faster than generative world models at comparable accuracy.
↳ Follow the thread