Latent Spatial Memory Stores 3D Scene State Inside Video World Models' Diffusion Latents
HuggingFace Daily Papers·low signal
A June Hugging Face trending paper proposes keeping 3D scene information in diffusion latent space rather than pixel space, cutting memory use and speeding generation for video world models. It points toward more efficient persistent-memory designs for interactive and long-context video generation.