Reddit
A 2.4-4M Parameter Latent Flow Transformer Generates 128x128 Faces on an RP2350 Microcontroller in ~20 Seconds
An r/MachineLearning post (385 upvotes) documents a tiny image generation model, a latent flow transformer of 2.4 to 4 million parameters quantized to int8, running fully on an RP2350 microcontroller and rendering 128x128 face images to an attached monitor in about 20 seconds at the slowest setting. This was the only post in r/MachineLearning above 70 upvotes for the day. It is a useful floor marker for how far generative image models compress, and the kind of result that never appears in vendor benchmark posts.
↳ Follow the thread