Fetching from the wire…
Infra2026-09-25 · source-backed
PR #58566 moves a contiguous copy in rocm_unquantized_gemm_impl into the only two branches that use it. Shapes passing the broad skinny-GEMM gate but matching neither branch were paying for a copy that got discarded. On Kimi-K3 FP4 across 8x MI355X, per-step copies dropped from 73 to 4, saving 0.339ms per decode step, almost all kernel-launch overhead. The author notes both earlier hypotheses from reading source were wrong, and only the profile found it. Good reminder for the "let the agent read the code and reason about perf" workflow.
Each link below shares sources, entities, or timing with this story.
A single-commit repo documents 304B parameters in 156.67 GB, 7.9–8.5K tok/s prefill, 830 tok/s at a 64-stream burst without OOM. The correctness fix is the reusable part: MI300X uses AMD's FNUZ FP8 variant rather than OCP standard, requiring a cache-writer overlay selecting fl...
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that'...
The project wires a CUDA application through ZLUDA to a cuBLAS/cuSPARSE/cuFFT shim to rocBLAS/hipBLASLt/rocSPARSE/HIP, using ZLUDA v6-preview.69, AMD HIP SDK 6.4 and LibTorch 2.3.0 against CUDA 11.8, validated by training a 2.2M-parameter PPO network. Only the RX 9060 XT (gfx1...
Strix Halo and Strix Point default to Vulkan instead of ROCm for up to 23% faster prompt processing and 8% faster generation, and AMD iGPUs without ROCm move to Vulkan instead of CPU on Linux. On Apple Silicon, gated-delta models train up to 25% faster and quantized MLX KV cac...
The repo appeared on trending with +135 stars and a repositioned pitch, pivoting from the general local-code-execution tool it launched as in 2023. It's now aimed directly at Claude Code and Codex but on the open-weight side. Single-source on the repositioning, so check the re...
Reuters, via Tech Startups, reports capital released against deployment milestones with Anthropic deploying up to two gigawatts of Instinct MI450 starting 2027. Same structure as Nvidia/OpenAI: compute vendor capital flowing to the lab that commits to buy the silicon. A two-gi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.