AutoArk's Edge0-35B-A3B Streams MoE Experts Off SSD to Run a 35B Model in ~3GiB on a Mac Mini
arXiv 2609.18063 (via Hugging Face)·high signal
Edge0-35B-A3B, built on Qwen3.5-MoE, fires only 4 of 256 experts per token and fetches just those weights from SSD on demand, holding peak active memory near 2.9-3 GiB. A trained prerouter predicts the next step's experts so storage reads overlap compute, adding up to 59% decode throughput; on a 24GB Mac mini M4 Pro it decodes around 20 tok/s and prefills at 113-140 tok/s. A Recover-LoRA distillation pass on a frozen int4 base closes most of the 4-bit quality gap, leaving a reported 3.9 point deficit.