OSS
DeepGrove's Maple-Preview Is a Ternary 20B-A1B MoE That Fits in a 5.31 GB Checkpoint and Hits 218 tok/s on a Mac Mini M4
DeepGrove published Maple-Preview to Hugging Face under MIT — a ternary-weight reasoning MoE with 20B total and 1B active parameters, 24 layers, 256 experts with 8 active, 3:1 SWA-512:GA attention, and a 131,072-token context, all in a 5.31 GB checkpoint. It reports 218 tokens/sec on a Mac mini M4 (5–16x faster than Gemma 4, Qwen3.5 and gpt-oss at comparable quality) with scores on LCBv6, AIME 2026, HMMT 2026 and GPQA-D, and the Show HN thread hit 140 points claiming 120 tok/s on an iPhone. This is a different bet than last week's Swiftlet SSD-streaming approach: ternary weights shrink the model itself rather than paging a large one off disk.
Source
↳ Follow the thread