Reddit
r/LocalLLaMA Practitioner Runs Inference Off 12x Gen4 3.2TB SSDs Instead of Buying More GPUs
A satirical r/LocalLLaMA post about a rumored '26T-a3b Le Chaton FAT' (135 upvotes, 43 comments) buries a real hardware configuration: twelve Gen4 3.2TB drives, two per card, used as the weight store for sparse-MoE inference because the poster cannot afford the equivalent 5060Ti count. The joke framing is the delivery mechanism for a genuine adaptation — as frontier open-weight MoEs push past 1T total parameters with tiny active counts, storage bandwidth rather than VRAM capacity becomes the binding constraint for hobbyist rigs. This is the kind of configuration that never appears in vendor documentation.
↳ Follow the thread