Fetching from the wire…
Tools2026-08-27 · source-backed
Released August 26, it brings MLX support for Qwen3.8 Flash Next, adds structured output to mlxrunner, and stops Metal GPU timeouts when loading models from slow storage (GitHub). Structured output on the MLX path is the practical unlock: Apple Silicon local inference can now be driven by JSON-schema-constrained tool calls the same way the llama.cpp path already could.
Each link below shares sources, entities, or timing with this story.
Ollama uses MLX / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Ollama uses MLX); both cover Apple Silicon, GitHub, MLX, Qwen3; reported by the same outlet (github.com).
Apple released MLX / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Apple released MLX); both cover MLX, Qwen3; overlapping topics (inference, local).
Ollama uses MLX / Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover GitHub; reported by the same outlet (github.com).
Ollama uses MLX / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover Apple Silicon, MLX; reported by the same outlet (github.com).
Ollama uses MLX / Shared entity: GitHub / Same source domain / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover GitHub; reported by the same outlet (github.com).
Linked by a graph relationship (Ollama uses MLX); both cover GitHub; reported by the same outlet (github.com).
Ollama uses MLX / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover JSON, MLX; reported by the same outlet (github.com).
Ollama uses MLX / Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama uses MLX); both cover GitHub; reported by the same outlet (github.com).