Fetching from the wire…
Infra2026-09-26 · source-backed
Merges on September 24 and 25 (#493, #494, #500) add an attention-only adapter, because V2 bypassed kvcached's hooks by calling a module-level init_kv_cache(). They also add packed four-dimensional K/V storage and hybrid Mamba page binding for V1. V2 still rejects hybrid/Mamba and cross-layer sharing, and ROCm stays on split K/V. 1,504 stars and no tagged release since v0.1.5 in April, so these are main-branch-only.
Each link below shares sources, entities, or timing with this story.
The project wires a CUDA application through ZLUDA to a cuBLAS/cuSPARSE/cuFFT shim to rocBLAS/hipBLASLt/rocSPARSE/HIP, using ZLUDA v6-preview.69, AMD HIP SDK 6.4 and LibTorch 2.3.0 against CUDA 11.8, validated by training a 2.2M-parameter PPO network. Only the RX 9060 XT (gfx1...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
Strix Halo and Strix Point default to Vulkan instead of ROCm for up to 23% faster prompt processing and 8% faster generation, and AMD iGPUs without ROCm move to Vulkan instead of CPU on Linux. On Apple Silicon, gated-delta models train up to 25% faster and quantized MLX KV cac...
DeepSeek posted a community notice: once V4.1 Flash launches around September 10 Beijing time, and until a V4.1 Pro exists, every V4 Pro request routes to V4.1 Flash and bills at Flash unit pricing. The stated reason is that Flash has surpassed Pro on performance, cost, speed...
One number predicts whether your agent finishes the task, and it isn't the benchmark score. Shubhra Mittal's paper (arXiv 2609.01660) analyzed 10,664 trajectories across nine models spanning 1.2B to 671B parameters and found task success follows P(n) = p^n, where p is a single...
A single-commit repo documents 304B parameters in 156.67 GB, 7.9–8.5K tok/s prefill, 830 tok/s at a 64-stream burst without OOM. The correctness fix is the reusable part: MI300X uses AMD's FNUZ FP8 variant rather than OCP standard, requiring a cache-writer overlay selecting fl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.