Reddit
A 1.57B-Parameter Dreamer 4 World Model Trained From Scratch for Under $150, After the Genie Architecture Failed on Controls
The author's first attempt built on Genie's architecture produced good-looking video where keypresses did essentially nothing, because Genie learns actions unsupervised into only 8 latent codes. Scrapping it for Dreamer 4 and generating every frame with Procgen, so ground-truth actions are known at each step, produced a 1.57B model on 9.6M frames for about $150, with tokenizer PSNR of 40.41 against the 35.7 reported in the Genie paper, end-to-end FVD of 32.19, and coherence out to 144 frames. A commenter made the sharpest critique: report an action-conditioned divergence metric, because PSNR and FVD both stay high when the controls do nothing, which is exactly the failure mode of attempt one.
Source
↳ Follow the thread