Fetching from the wire…
Public story · 2026-09-26 · high
Token generation doesn't speed up at all, so the gain only helps CPU boxes running long prompts, not chatbots answering one token at a time.
Why now: The PR merged September 26.
A new llama.cpp CPU matmul path speeds prompt processing 3 to 7 times on x86 chips, merged into the project September 26. Prefill, feeding a model its context before it generates anything, ate most of the cost of CPU inference. This patch cuts that cost for long-context RAG pipelines and batch summarization jobs running on CPU-only boxes.
PR #27851 unpacks k-quant weights into 256x256 int8 tiles. A 16x16 VNNI microkernel runs over those tiles, replacing a vec_dot path that unpacked the same quantized weights repeatedly.
On an AMD 9950X3D running 8 threads at 8192x8192, q3_K matmul rose from 0.73 to 5.16 TFLOPS, about 7 times faster. q4_K reached 4.97 TFLOPS, up from 1.06, close to a 4.7x gain. Both roughly double the throughput of the existing repack path, and both show lower RMSE, so the speedup doesn't cost accuracy.
Token generation doesn't benefit: the new path engages only above 64 rows, and pure GEMV below that runs at 80% of stock speed.
Each link below shares sources, entities, or timing with this story.
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
Dario Amodei published "We Must Pace the Frontier" on September 12. Altman and Musk agreed within hours. Hassabis called the direction correct. By Monday morning the market had priced it. Nasdaq 100 futures fell 1.5% and S&P 500 futures 0.8%. Nvidia dropped 3%, AMD 5.7%, and A...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
Luu's September 1 post drew 852 points and over 1,000 comments, walking predictions from February 2024 through November 2025 after removing unfalsifiable and tautological ones. Misses include "AI has peaked" (Feb 2024), Meta "dying" (Nov 2024) against revenue going $135B to $2...
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.