Three-Year Local Cluster Retrospective: 30M Tokens Since January, 'Cloud Is Cheaper, Hands Down,' and a Near House Fire From Daisy-Chained PSUs
An r/LocalLLaMA 'Showoff Saturday' post (205 upvotes / 80 comments) traces one builder from two 3090s in September 2023 to a 4x RTX 6000 Pro Max-Q (300W each) plus 4x power-limited 3090s cluster, and is unusually honest about the economics: 30M tokens generated since January 2026, and 'I don't expect to break even ever.' The failure catalogue is the useful part — bad PCIe cables, low-quality PSUs causing non-reproducible throughput drops, burnt-out risers from incorrect multi-PSU wiring, and an add2psu board that burned out under load after daisy-chaining three 1300W consumer supplies. Stated motivation is privacy (keeping private keys and business data off cloud APIs) and freedom from rate limits, not cost.
↳ Follow the thread