The 'Fragile' Gated DeltaNet Gates Are the Least Quantization-Sensitive Part, and All 496 Linear Layers Go to NVFP4 W4A4
Community 4-bit quantizations of Qwen3.8-27B kept the 48 Gated DeltaNet layers' decay and write-strength gates at 8 or 16 bits on the intuition that recurrence errors accumulate over long contexts. Minima tests that intuition by pushing NVFP4 W4A4 through all 496 linear layers including GDN, and matches BF16 within seed noise across 4K/32K perplexity, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench and RULER to 64K, with a 5-task average of -0.52, the smallest footprint at 17.5 GiB, and 14-19% faster prefill. The mechanism study finds the gate projections compress roughly 11% GEMM error to about 2% output error and the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens because each write overwrites state along the current key direction.
↳ Follow the thread