Fetching from the wire…
Public story · 2026-08-31 · high
A LocalLLaMA user ran Muse Glimmer 30B at 3.00bpw and reports no noticeable quality drop for agent work against the 17GB official quant.
Why now: Posted to r/LocalLLaMA on August 31.
A LocalLLaMA user got Muse Glimmer 30B running fully in 12GB of VRAM on an EXL3-SC quant at 3.00bpw H4, with a Q8 KV cache. The model's official K-quant needs 17GB, more than a 12GB card holds, with no noticeable quality drop against it for agent work.
A LocalLLaMA thread on EXL3 quants has the numbers: 100K tokens of context at around 30 tokens a second, all inside that 12GB budget. A 17GB model that doesn't fit on a 12GB card normally means spilling to system RAM or dropping to a smaller model instead. EXL3's 3.00bpw format skips both trade-offs.
The same tester tried Qwen 3.8 27B at a more aggressive SC2.20bpw H3 quant and called it usable. For coding, they still went back to Unsloth's UD_Q4_K_XL, a GGUF quant.
The post doesn't name the hardware beyond the 12GB VRAM ceiling. It has no formal eval score, just one person's read against their daily-driver quant.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover GGUF, GPU, LocalLLaMA, Qwen; reported by the same outlet (reddit.com).
Alibaba released Qwen / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover GPU, LocalLLaMA, Muse Glimmer, Qwen; reported by the same outlet (reddit.com).
Unsloth partners with NVIDIA / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Unsloth partners with NVIDIA); both cover GPU, LocalLLaMA, Qwen; reported by the same outlet (reddit.com).
Ollama supports Qwen / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama supports Qwen); both cover LocalLLaMA, Qwen, Unsloth; reported by the same outlet (reddit.com).
Unsloth supports Blackwell / Shared entities / Same source domain / Earlier coverage / Downstream implication
Linked by a graph relationship (Unsloth supports Blackwell); both cover GPU, LocalLLaMA, Qwen; reported by the same outlet (reddit.com).
Alibaba released Qwen / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover LocalLLaMA, Qwen; reported by the same outlet (reddit.com).
Unsloth supports DeepSeek / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Unsloth supports DeepSeek); both cover GGUF, LocalLLaMA, Muse Glimmer; reported by the same outlet (reddit.com).
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover LocalLLaMA, Qwen, Unsloth; overlapping topics (around, cache).