Seoul World Model: KAIST and NAVER AI Lab Release First Real-City Grounded Video World Model
arXiv / NAVER AI Lab·high signal
SWM generates coherent long-horizon urban video conditioned on NAVER Map street-view retrieval, supporting hundreds of meters of trajectory simulation with text-prompted scenario variations across Seoul, Busan, and Ann Arbor test sets. A 'Virtual Lookahead Sink' mechanism re-grounds generation to future retrieved images, stabilizing long-horizon coherence where prior video world models fail. SWM outperforms all prior video world models on the real-city benchmark and ships with open weights.