Fetching from the wire…
Public story · 2026-08-05 · high
Maple-Preview hits 218 tokens per second on a Mac mini M4, 5 to 16 times faster than Gemma 4 and Qwen3.5, DeepGrove says.
Why now: On August 5, Maple-Preview's traction comes from a Show HN thread past 140 points, not from any name recognition DeepGrove has yet.
DeepGrove published Maple-Preview, a 20-billion-parameter mixture-of-experts model that runs with just 1 billion parameters active per token, in a 5.31GB checkpoint, per its Hugging Face listing.
On a Mac mini M4, DeepGrove clocks it at 218 tokens per second. That's fast enough to run a 20B model locally without a GPU server or an API bill. That's 5 to 16 times faster than Gemma 4, Qwen3.5 and gpt-oss at comparable quality, DeepGrove reports. It backs that with scores on four benchmarks: LCBv6, AIME 2026, HMMT 2026 and GPQA-D.
The model has 24 layers and 256 experts, with 8 active on any given pass. Context runs to 131,072 tokens, built from a 3:1 mix of sliding-window and global attention.
The weights are ternary, which shrinks the model itself instead of paging a large one off disk. That's a different bet than the SSD-streaming approach other small-model projects use. There, a large model stays on disk and pages sections into memory as needed.
The Show HN thread backing the release passed 140 points, with one commenter reporting 120 tokens per second on an iPhone. It ships under MIT, so anyone can test the speed and quality claims independently.
Each link below shares sources, entities, or timing with this story.
DiffusionGemma benchmarked against Gemma / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (DiffusionGemma benchmarked against Gemma); both cover Gemma, Hugging Face; reported by the same outlet (huggingface.co).
Hugging Face partners with NVIDIA / Shared entities / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Gemma, Hugging Face; earlier Gemma coverage from 2026-06-11.
Hugging Face partners with NVIDIA / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover AIME, Qwen3; overlapping topics (active, aime).
Gemma built by Google / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Gemma built by Google); both cover Gemma, Hugging Face; overlapping topics (faster, gemma).
Hugging Face criticizes OpenAI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover Gemma, MIT; overlapping topics (comparable, context).
Gemma built by Google / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Gemma built by Google); both cover Gemma, Qwen3; reported by the same outlet (huggingface.co).
Hugging Face partners with NVIDIA / Shared entity: Hugging Face / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Hugging Face; reported by the same outlet (huggingface.co).
Hugging Face criticizes OpenAI / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover Gemma, MIT; earlier Gemma coverage from 2026-04-20.