PhiZero Replaces Pixel-Space Prediction With a Discrete 'Physical Language' World Model That Reasons Then Renders
PhiZero (arXiv 2607.28624, July 30) proposes physical language — a compact discrete representation of world-state transitions — learned self-supervised from in-the-wild videos, in contrast to physical world models that predict future frames directly in pixel space and leave dynamics implicit inside a high-dimensional visual predictor. It adopts a reason-then-render paradigm: first infer future world evolution as a physical-language sequence, then render those transitions into video. The authors report results across generation and understanding benchmarks plus demonstrations of interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer; it is #4 on Hugging Face trending with 148 upvotes.
Source
↳ Follow the thread