Sources
SolarWM open-sources a 1.43M-clip world-model data engine with frame-aligned camera geometry, captions and provenance
arXiv 2609.02886 (2026-09-02, 112 HF upvotes) is the highest-upvoted world-model release this week and is deliberately framed as infrastructure rather than a model. Its data engine converts 1.43 million canonical clips from 10 datasets into one frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions and provenance, decoupling per-source processing from mixture construction. The paper's argument is that naive data mixing and model-specific implementations are what make long-horizon video world-model results irreproducible, and a backbone-native adaptation framework plus shared camera conditioning is the fix.
Source
↳ Follow the thread