Fetching from the wire…
Models2026-09-25 · source-backed
v0.40.0-rc0, published September 25, runs every architecture the MLX runner supports on MLX by default on Apple Silicon, with qwen3.8 as the worked example, and the maintainers say more models get tested and enabled during the RC. (release) Mac users get Apple's native ML stack without touching a flag. Anyone with published local-model benchmarks on M-series hardware should re-run them against this build, because the old numbers no longer describe the default path.
Each link below shares sources, entities, or timing with this story.
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
Edge0-AI/Edge0, created September 8, went from 269 to 583 stars in two days. It packages SSD expert offload, Recover-LoRA and prerouter routing prediction into an MLX-backed framework: edge0-35b is a 4-bit 40-layer 256-expert model built on Qwen3.5-MoE 35B-A3B needing about 2....
Jared Palmer published Kev September 17 on Qwen3.5, Apache-2.0 with training code and frozen eval suites. One request carries yes/no, multiple-choice and rating questions that share input text but can't read each other, and the API matches TypeSafe's System One so their Python...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
Released September 3, PR #2206 confines storage names against directory traversal, confines symlink targets to configured storage directories, requires API keys for non-loopback server bindings, authenticates the Ollama routes, and defaults the REST server to loopback port 808...
laya-mlx, created September 19 at 549 stars, is a native Apple Silicon runtime skipping PyTorch and Transformers entirely (GitHub). It reports 13.4ms median end-to-end for short English questions and 7.4ms on the multilingual checkpoint, 146.8 and 395.0 q/s batched on an M3 Ma...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.