StateFlow Argues Generative Previsualization Needs a Persistent 3D World State, Not Better One-Shot Video Prompts
arXiv 2608.12314 (submitted 12 Aug 2026, 22 upvotes on HuggingFace today) targets previsualization for film, games, architecture and urban design, where current generative methods try to control scene, action, camera and spatio-temporal dynamics jointly through a single prompt and offer weak controllability and no real iterative editing. The authors' claim is architectural: successive frames are mostly local modifications or recombinations of a shared state, so the missing component is an explicit, persistent working state rather than a better one-shot generator. StateFlow maintains an editable 3D world of scene elements and camera configurations across construct/evolve/access stages — lifting generated 2D content into 3D via prior-guided, conflict-aware dual-view initialization — and calls off-the-shelf video models only when higher fidelity is wanted.
↳ Follow the thread