Fetching from the wire…
Models2026-07-28 · source-backed
A Qwen3-8B derivative retrained with ternary quantization-aware training, storing all 252 transformer linears in a ternary format one-eighth the size of fp16, decoded inside the matmul kernels. 8.19B parameters in a 2.56 GB download with 40,960-token context, scoring 72.1 MMLU, 80.2 IFEval-strict, 53.4 GSM8K, 68.9 BFCL v3, running 763 tok/s on an H100 with speculative decoding versus 33.7 tok/s on an Apple M5 via MLX. No access request. Weights at FermionResearch/Neutrino-8B with 0.6B variants. (Fermion Research)
Each link below shares sources, entities, or timing with this story.
Ollama uses MLX / Shared entities / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover Apache, MLX; earlier Apache coverage from 2026-05-06.
Linked by a graph relationship (Ollama uses MLX); both cover Apache, Qwen3; earlier Apache coverage from 2026-04-02.
Ollama uses MLX / Shared entity: Qwen3 / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover Qwen3; overlapping topics (disk, download).
Ollama uses MLX / Shared entity: Qwen3 / Earlier coverage / Tension
Linked by a graph relationship (Ollama uses MLX); both cover Qwen3; earlier Qwen3 coverage from 2026-04-23.
Linked by a graph relationship (Ollama uses MLX); both cover Qwen3; earlier Qwen3 coverage from 2026-03-22.
headroom benchmarked against GSM8K / Shared entities / Shared topic
Linked by a graph relationship (headroom benchmarked against GSM8K); both cover BFCL, GSM8K; overlapping topics (bfcl, context).
Ollama uses MLX / Shared entity: MLX / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover MLX; earlier MLX coverage from 2026-07-25.
Ollama uses MLX / Shared entity: Qwen3 / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover Qwen3; earlier Qwen3 coverage from 2026-07-13.