A $3,010 build gets 128GB of VRAM and 256GB of RAM, at 700-900W under prefill
r/LocalLLaMA (817 upvotes, 231 comments)·medium signal
The top r/LocalLLaMA post of the day (817 upvotes, 231 comments) itemizes a home inference server: 4x AMD V620 at $1,400, 256GB DDR4 RDIMM 2666 at $610, a Huananzhi D12D board at $410, an EPYC 7452 at $170, and a 1600W ASRock PSU at $220. The builder measured 1.3k tokens/sec prefill and 70 tok/s code generation at 128k+ context on Qwen3.8-next-flash AutoRound W4A16 with MTP-2 on a vLLM fork, and reports 700-900W during prefill, 500-600W during decode. He also returned a Lenovo P620 workstation first over proprietary hardware lock-in.