Fetching from the wire…
Infra2026-09-16 · source-backed
Released September 14, it graduates MLX safetensors ollama create out of experimental, while GGUF creation now requires llama.cpp tooling for conversion and quantization. Runaway repeat-token detection now needs 100 repeated tokens before firing, cutting false positives on OCR-style output, and typical_p is deprecated for new models while existing GGUF models keep it.
Each link below shares sources, entities, or timing with this story.
The July 6 release delivers nearly 90% faster Gemma 4 token generation through multi-token prediction with automatic draft-length tuning, on by default, output-preserving, no config (Ollama). It also adds MLX-engine support for more model families and flash attention for older...
Google DeepMind shipped Quantization-Aware Training checkpoints for every Gemma 4 size, and the headline number is genuinely useful: the smallest model goes from 11.4GB to 1.1GB. That's 0.84GB if you go text-only. Up to ~72% lower VRAM and 2x faster inference on mobile NPUs, w...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Lily is Apache 2.0 inside pplx-garden, a small Metal inference server built for Qwen3.6-35B-A3B converted to MLX affine 4-bit, exposing a minimal OpenAI-compatible chat API with greedy decoding. The README explicitly rules out dense and smaller Qwen checkpoints, BF16, GGUF, AW...
Claims up to 2x faster generation, and adds fine-tuning of both MoE models on text or image datasets on Apple Silicon via MLX (release). Follow-up turns in long Qwen chats on Mac are reported up to 30x faster, MLX models now use full context size, and GLM-5.3 MLX fine-tunes ex...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.