Fetching from the wire…
Infra2026-09-09 · source-backed
auto now means "pick a probably good mode for your system"; its old behavior of lazily loading tensors over 4 GiB becomes mode large, and loading all becomes all. The old default effectively re-enabled mmap on iGPUs, which lose substantial performance that way. A day later, --mmap, --mlock and --direct-io were removed from the arg parser entirely, so scripts passing the old flags now fail to parse. GitHub
Each link below shares sources, entities, or timing with this story.
Block released it August 27 with a security section dominated by defaults that previously failed open: fail closed on malformed tool visibility, permission denies now take precedence, fail closed on invalid default GCP credentials and invalid Codex ACP mode, honor plugin enabl...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
The abliteration tool gained 215 stars to reach 30,103, but the stronger signal is downstream: the HF trending endpoint returns DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU and Momoking/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4, both naming the too...
PR 27742 adds gated delta net layers, 512-expert MoE with top-10 selection, hyper-connections, query-key sparse attention and the per-layer n-gram embeddings as a mmap table that can sit in RAM or on disk (GitHub). Reported 55 tok/s on 4x3090 with the Q4 GGUF, and one commente...
The August 26 release adds Decode Context Parallel for Kimi-K3, fused FlashKDA decode and prefill kernels, combined all-gathers with a claimed 1.5-3x kernel-level speedup, an adaptive speculative token budget worth about 60% better DSpark TTFT, and optional shared-expert shard...
Replicated David Ng's RYS method — duplicating layers 12-14 routes hidden states through the reasoning circuit twice, boosting BBH Logical Deduction from 0.22 to 0.76 and GSM8K from 0.48 to 0.64. Zero training, no weight changes. Same technique on Qwen2.5-Coder-32B improved re...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.