Fetching from the wire…
Public story · 2026-08-03 · high
Unsloth's number has 809 upvotes on Reddit, but Alibaba hasn't shared a benchmark, license, or parameter count.
Why now: Alibaba still hasn't published a benchmark, license, or parameter count for the 27B, leaving Unsloth's number as the only spec in public view.
The unreleased Qwen3.8-27B will run on 17GB of RAM or VRAM, per Daniel Han of Unsloth.
That puts a 27-billion-parameter model within reach of a single 24GB consumer GPU. Hobbyists already run smaller open models on that same card instead of a multi-GPU rig.
Han's claim ran in a post on r/LocalLLaMA that pulled 809 upvotes and 143 comments. Unsloth's own X account confirmed the same 17GB figure. Alibaba, meanwhile, hasn't published a benchmark table, a license, or the model's activated-parameter count.
A tooling vendor is telling a local-inference community what hardware they'll need for a model the lab hasn't shipped, priced, or licensed. Unsloth builds fine-tuning and quantization tools for that exact crowd, so getting ahead of Alibaba's own announcement lines up with its business. That doesn't make the 17GB number wrong.
If the shipped model needs more than 17GB, Unsloth's number is the spec people hold it to, not anything Alibaba wrote.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
The Hugging Face page is marked "Upcoming release" with no model card, license, architecture details, context length or benchmarks, after Alibaba promised both Qwen3.8-Max and the 27B weights for the week of August 10. A ModelScope countdown pointed at August 15. Unsloth signa...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.