Reddit
EXL3 at 3.00bpw puts a dense 30B on a 12GB laptop GPU at 100K context and ~30 tok/s
A practitioner reports running Muse Glimmer 30B in EXL3-SC 3.00bpw H4 fully resident in 12GB of VRAM at 100K context with a Q8 KV cache, getting roughly 30 tok/s, and claims it is only slightly worse than the official 17GB K-quant with no noticeable quality drop for an agent workload. The same person tried Qwen 3.8 27B at SC2.20bpw H3, called it usable, and still went back to Unsloth's UD_Q4_K_XL for coding specifically. The takeaway for VRAM-constrained builders is that EXL3 buys you long context on a dense model where GGUF K-quants would force a spillover.
Source
↳ Follow the thread