Sources
WorldCrafter gives a video world model a camera-queryable 3D memory so it stops forgetting what it already rendered
Video world models let you explore dynamic environments interactively but drift away from prior observations over long horizons and across viewpoints. WorldCrafter learns an implicit 3D-aware memory that the requested camera viewpoint queries directly, so the target view decides how multi-view evidence gets compressed into the generator's limited token budget, with a memory encoder and pose-conditioned readout trained jointly with the video generator and no explicit depth correspondences. Combined with few-step distillation it supports streaming scene exploration from a single image or text prompt, and it drew 86 upvotes on HuggingFace's September 22 list.
↳ Follow the thread