Meta Ships Muse Glimmer 30B Under Apache 2.0 — a 29.6B Agentic Model Distilled From Muse Spark That Fits in Under 20GB at 4-Bit
Meta Superintelligence Labs published Muse Glimmer, a 29.6B dense causal transformer (52 layers, 6,656 hidden dim) with a ~1.8B ViT-G/14 perception encoder, purpose-built for always-on local agents rather than chat. It ships Apache 2.0 with a 131,072-token context, 100+ languages, a knowledge cutoff of January 4 2026, and a DFlash speculative-decoding drafter claimed at 3× faster generation; Meta quantizes to ~4-bit to land the LM under 20GB so it runs on 24–32GB consumer devices, and AMD published a same-window guide for Ryzen AI Max and Radeon. Reported SWE-Bench Pro is 51.2%, positioned against Gemma4-31B and Qwen3.6-27B. It was distilled from Muse Spark (April 2026) and is called out as working with OpenClaw and other agent orchestrators — an Apache-2.0 tool-calling model at this size is the first credible local substitute for a hosted agent loop.
Source
↳ Follow the thread