Ollama 0.19 Adds MLX Backend — 93% Faster Decode on Apple Silicon for Local LLMs
Ollama Blog·medium signal
Ollama released version 0.19 in preview with MLX (Apple's ML framework) as the backend for Apple Silicon Macs, delivering 57% faster prefill and 93% faster decode compared to v0.18. On M5 chips, the update leverages GPU Neural Accelerators for both time-to-first-token and generation speed, achieving 1,851 tokens/s prefill and 134 tokens/s decode with int4 quantization. Currently supports Qwen3.5-35B-A3B with more models coming; requires 32GB+ unified memory.