Fetching from the wire…
Infra2026-09-25 · source-backed
PR #27952 adds an int8 coopmat1 MMQ shader, so Vulkan no longer dequantizes to fp16 before touching matrix cores. Coverage spans q4_0 through q6_k, mxfp4, nvfp4 and iq4_nl. On a Strix Halo Radeon 8060S, q4_0 MUL_MAT improves 1.29x over master and beats ROCm. RDNA4 is roughly neutral except for MoE prompt processing, and slower quants are disabled there. The same day, llama.cpp merged llama_batch_ext (#24669), superseding a batch-API proposal open since #11875, so bindings and servers built on llama_batch should watch for the migration.
Each link below shares sources, entities, or timing with this story.
Strix Halo and Strix Point default to Vulkan instead of ROCm for up to 23% faster prompt processing and 8% faster generation, and AMD iGPUs without ROCm move to Vulkan instead of CPU on Linux. On Apple Silicon, gated-delta models train up to 25% faster and quantized MLX KV cac...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
The project wires a CUDA application through ZLUDA to a cuBLAS/cuSPARSE/cuFFT shim to rocBLAS/hipBLASLt/rocSPARSE/HIP, using ZLUDA v6-preview.69, AMD HIP SDK 6.4 and LibTorch 2.3.0 against CUDA 11.8, validated by training a 2.2M-parameter PPO network. Only the RX 9060 XT (gfx1...
A llama.cpp fork by fewtarius aimed at AMD APUs, iGPUs and handhelds adds a persistent SSD-backed KV cache with hot/warm/cold tiering. On an Ayaneo Flip KB running Qwen3.6-35B over a 15,700-token prompt, cold TTFT was 143.1 seconds versus 0.99 seconds warm, a 144.5x speedup. A...
The v1.11.0 release reads DeepSeek V4.1 Flash's vendor checkpoint natively with no conversion, 510 GB on disk, fp8 dense with 32x32 ue8m0 tiles and fp4 experts byte-identical to the mxfp4 its Kimi K3 engine already reads. A cold turn went from 78.7s to 25.1s through batched ex...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.