Tools
Unsloth v0.1.804-beta runs Qwen3.8-Flash-Next on 75 GB RAM and GLM-5.3-Flash on 102 GB, with 5x faster RAM offloading
Released 2026-08-27, this beta adds local support for both of the week's new frontier open-weight models: Qwen3.8-Flash-Next (125B) at 75 GB RAM and GLM-5.3-Flash at 102 GB combined RAM plus VRAM, with GGUFs published for both. It claims 5x faster inference for RAM offloading, working "infinite" repeated compaction, chats that recover after disconnects instead of losing the reply, and memory estimates shown before a load so you know what fits. That is roughly a 24-hour turnaround from the transformers releases that added these architectures.
Source
↳ Follow the thread