Tools
Unsloth 0.1.805-beta doubles Qwen3.8-Flash-Next and GLM-5.3-Flash inference with MTP on by default
Released 2026-09-02, Unsloth v0.1.805-beta enables multi-token prediction by default for Qwen3.8-Flash-Next and GLM-5.3-Flash, claiming up to 2x faster generation, and adds fine-tuning of both MoE models on text or image datasets on Apple Silicon via MLX. Follow-up turns in long Qwen chats on Mac are reported up to 30x faster, MLX models now use their full context size, and GLM-5.3 MLX fine-tunes export to GGUF. Local models can also now use Codex's `apply_patch` tool for code edits.
Source
↳ Follow the thread