Fetching from the wire…
Public story · 2026-08-16 · high
The gain came at the identical 3.3 GB size, a 76% cut from the original 14 GB model, per its Hugging Face model card.
Why now: The comparison was live on Hugging Face as of August 16, with no independent replication published yet.
A quantization method called CADA pushed Gemma 4 E4B's reasoning score from 28.9% to 69.5% at the same size, per its Hugging Face model card. That's a 40.6-point jump at the same 3.32 GiB budget, the range engineers target to fit language models on phones and laptops without a GPU.
Standard 2-bit quantization (IQ2_XXS) squeezes every tensor to the same precision. CADA instead searches for which tensors can tolerate 2-bit compression and which need more bits. It spends the same total byte budget unevenly across the model, per the model card.
The 3.32 GiB result is a 76.32% cut from the 14.02 GiB BF16 source model, the card says. It doesn't say how the reasoning benchmark was scored, or whether anyone besides ByteOtter has run the comparison.
For anyone shipping small models on-device, that's the specific number worth watching for independent replication before trusting a 2-bit checkpoint on reasoning-heavy work.
Each link below shares sources, entities, or timing with this story.
Claude Code uses Hugging Face / Shared entities / Same source domain
Linked by a graph relationship (Claude Code uses Hugging Face); both cover Gemma, Hugging Face; reported by the same outlet (huggingface.co).
Hugging Face partners with NVIDIA / Shared entities / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Gemma, Hugging Face; earlier Gemma coverage from 2026-06-11.
DiffusionGemma benchmarked against Gemma / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (DiffusionGemma benchmarked against Gemma); both cover Gemma, Hugging Face; reported by the same outlet (huggingface.co).
Hugging Face partners with NVIDIA / Shared entity: Hugging Face / Same source domain / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Hugging Face; reported by the same outlet (huggingface.co).
Ollama supports Gemma / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Ollama supports Gemma); both cover E4B, Gemma; reported by the same outlet (huggingface.co).
Claude Code uses Hugging Face / Shared entity: Gemma / Earlier coverage
Linked by a graph relationship (Claude Code uses Hugging Face); both cover Gemma; earlier Gemma coverage from 2026-08-12.
Hugging Face released Safetensors / Shared entity: Hugging Face / Earlier coverage
Linked by a graph relationship (Hugging Face released Safetensors); both cover Hugging Face; earlier Hugging Face coverage from 2026-07-27.
Hugging Face partners with NVIDIA / Shared entity: Hugging Face / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Hugging Face; earlier Hugging Face coverage from 2026-07-27.