AURORA-LM Argues Text Is the Last Holdout Against Continuous Latents, and Ships a 1B Flow-Matching Diffusion Language Model to Prove It
Posted to arXiv August 3, 2026 (2608.02602) and trending on Hugging Face Daily Papers with 61 upvotes, AURORA-LM opens on the observation that "language remains an outlier in generative modeling" — images, video and audio all moved to continuous latent spaces while text stayed discrete-token. The architecture pairs a query-based encoder-decoder with a block-causal Diffusion Transformer trained by flow matching. At 1B parameters and ~1,500 EFLOPs it reports the strongest results among continuous and diffusion-based language models on OpenWebText free generation and XSum summarization, beating a larger publicly released latent-diffusion LM under matched protocols. Code is on GitHub at fyv587/AURORA-LM. Sixty-plus contributors, led out of Nanjing University.
↳ Follow the thread