Hacker News
Neutrino-1 8B Ships Apache-2.0 With Ternary-Packed Weights: 3.88 GB On Disk, 72.1 MMLU, 33.7 tok/s on an M5
Fermion Research released Neutrino-1 8B on July 27, 2026 — a Qwen3-8B derivative retrained with ternary quantization-aware training, storing all 252 transformer linears in a ternary-family format one-eighth the size of fp16, decoded inside the matmul kernels. The 8,190,735,360-parameter model is a 2.56 GB download / 3.88 GB on disk with a 40,960-token context, scoring 72.1 MMLU, 80.2 IFEval-strict, 53.4 GSM8K and 68.9 BFCL v3, and running 763 tok/s on an H100 with speculative decoding versus 33.7 tok/s on an Apple M5 via MLX. Apache-2.0 with no access request; weights are on Hugging Face as FermionResearch/Neutrino-8B alongside 0.6B and 0.6B-Chat variants.
↳ Follow the thread