Tools
Ollama speeds up Qwen 3.8 prompt processing on Apple Silicon by up to 19%
ollama/ollama PR #18550, merged 2026-09-23, uses MLX's gated-delta kernel for long scans and folds dense MLP global scales into SwiGLU. On an M5 Max, prompt throughput rose from 715.2 to 848.5 tokens/s at 2k tokens (+18.7%), 695.6 to 828.1 at 8k (+19.1%), and 704.6 to 802.6 at 16k (+13.9%). A companion PR (#18601) stops the macOS app from hanging while it asks System Events whether ChatGPT or Codex is running.
Source
↳ Follow the thread