Reddit
Per-Tensor Precision Allocation Recovered Gemma 4 E4B's Reasoning From 28.9% to 69.5% at the Exact Same 3.3 GB Budget
A quantization called CADA, posted to r/LocalLLaMA, searches per-tensor mixed precision around an IQ2_XXS byte budget rather than applying uniform 2-bit quantization. On the byte-matched control comparison the model card reports reasoning going from 28.906% (stock IQ2_XXS) to 69.531% — a 40.6 percentage-point gain, or +140.54% relative — at 3.32 GiB, a 76.32% reduction from the 14.02 GiB BF16 source. If the result holds under independent eval, it says the 2-bit quality cliff practitioners have accepted for years is largely an allocation failure, not an information-theoretic floor.
↳ Follow the thread