APEX MoE Quants: 25+ New Models Added Including New I-Nano Tier for Ultra-Low Memory
r/LocalLLaMA·medium signal
The APEX MoE-aware mixed-precision quantization strategy added 25+ new models since the initial Qwen 3.5 post, plus a new 'I-Nano' tier targeting ultra-low-memory deployment. APEX exploits the sparse activation patterns in Mixture-of-Experts architectures to apply different quantization levels to active vs dormant experts, achieving better quality-per-bit than uniform quantization. The update (104 upvotes, 36 comments) shows the quantization community rapidly expanding coverage beyond the initial proof-of-concept.