Fetching from the wire…
Public story · 2026-09-02 · high
The beta also adds Apple Silicon fine-tuning for both models and Codex-style local edits via apply_patch.
Why now: Follow-up fixes came after the initial beta post, which is why the numbers attached to it have already moved.
Unsloth turned on multi-token prediction by default for two open MoE models, Qwen3.8-Flash-Next and GLM-5.3-Flash, in the 0.1.805-beta release. The library claims up to 2x faster generation as a result.
For anyone running open models locally, generation speed decides whether the setup is usable day to day. A default speedup this size changes which models make sense to run on a laptop rather than in the cloud. So far, only Unsloth's own benchmark backs that 2x claim.
The same release adds fine-tuning for both MoE models on text or image datasets, through Apple's MLX framework. That lets Mac users train Qwen3.8-Flash-Next and GLM-5.3-Flash locally.
Follow-up fixes brought more changes on top of the beta. Long Qwen chats on Mac run up to 30x faster, and MLX models use their full context size instead of a truncated one. GLM-5.3 MLX fine-tunes export to GGUF for use outside MLX. Unsloth also added support for Codex's apply_patch tool, letting locally run models make edits the same way Codex does.
Each link below shares sources, entities, or timing with this story.
Released August 27 with GGUFs for both, claiming 5x faster inference for RAM offloading, working repeated compaction, chats that recover after disconnects instead of losing the reply, and memory estimates shown before a load (GitHub). That's roughly a 24-hour turnaround from t...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.