Sources
DreamX-Phi 1.0 Uses SAM3 Masks and a Frozen V-JEPA Teacher to Stop Robot Arms Losing Their Identity Mid-Prediction
The DreamX Team (arXiv 2608.13489, August 13) released an action-conditioned video world model for robotic manipulation that takes an observed frame, a language instruction and an action sequence and forecasts future observations. Technical contributions include PRoPE-style geometric encoding to preserve arm identity, a lightweight depth branch for scene geometry, SAM3 masks paired with a frozen V-JEPA teacher for object consistency, and distribution-matching distillation for a few-step student. It took first place on Track 1 and second on Track 2 of the WorldArena 2.0 Challenge, with model and code promised publicly.
↳ Follow the thread