OSS
slotstream runs a 104GB model on a 48GB Mac at 12 tok/s by streaming MoE experts off SSD with pread
carloslfu/slotstream hit 220 points on Hacker News overnight and holds 210 stars, MIT, Swift and MLX, created August 28. It runs Qwen3.8-Flash-Next, a 125B MoE weighing 103.8GB across 24 files at 4-bit, by loading only the 3.8GB dense trunk into RAM (about 2 second startup) and reading routed experts with pread into a fixed pool of cache slots shared across all 48 layers, so hot layers borrow slots from cold ones. It avoids mmap because MLX cannot materialize part of a memory-mapped tensor, which would force the full ~100GB. Published throughput scales 3 tok/s at 8GB, 4 at 16GB, 8 at 24GB, 12 at 48GB and up, needing ~110GB free disk.
Source
↳ Follow the thread