Fetching from the wire…
Vibe Coding2026-09-14 · source-backed
The top r/LocalLLaMA post of the day itemizes it: 4x AMD V620 at $1,400, 256GB DDR4 RDIMM 2666 at $610, a Huananzhi D12D board at $410, an EPYC 7452 at $170, a 1600W PSU at $220. Measured 1.3k tok/s prefill and 70 tok/s code generation at 128k+ context on Qwen3.8-next-flash AutoRound W4A16 with MTP-2 on a vLLM fork, drawing 700-900W during prefill. The builder returned a Lenovo P620 first over hardware lock-in, which is its own data point about who's buying workstations now.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
AMD argues agents require so much CPU-heavy orchestration alongside GPU inference that the ratio is moving from 1:8 toward 1:1. If AMD is right, that means entirely new racks of CPU servers in every AI data center. Major implications for AMD's EPYC roadmap and Strix Halo APUs,...
Supports both an integrated MTP head and a separate -md draft file, with the poster reporting 45 to 90 tok/s on a 5090 with 128GB system RAM and working configs down to a 12GB 4070. A separate 2x3090 plus DDR4 build reports 37-41 t/s decode with UD-Q4_K_XL plus expert cache pl...
A user moved from UD-Q3_K_XL at 140k context to UD-IQ3_XXS and cleared 200k on a 16GB eGPU over Thunderbolt 4, with KV cache at q5_1 and llama.cpp built with DGGML_CUDA_FA_ALL_QUANTS=ON (r/LocalLLaMA). Prompt processing fell from 700-800 tok/s to 400. A commenter on an RTX 508...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.