Alibaba's Marco-Mini (17.3B/0.86B Active) and Marco-Nano (8B/0.6B Active) — Extreme-Efficiency MoE Models with Sub-1B Active Parameters
r/LocalLLaMA (75↑, 39cmts)·medium signal
Alibaba's AIDC-AI lab quietly released Marco-Mini-Instruct and Marco-Nano-Instruct on HuggingFace, featuring extreme MoE efficiency: Marco-Mini activates only 0.86B of its 17.3B parameters per token, and Marco-Nano activates just 0.6B of 8B. These represent a new floor for active-parameter efficiency in instruction-tuned models, relevant to edge deployment and resource-constrained environments. The models went largely unnoticed for six days before surfacing on r/LocalLLaMA.