Reddit
DeepSeek V4 Flash 0731 at IQ2_M Runs ≈3.5 tok/s on Dual RTX 3060s and 96GB RAM — and LM Studio Refused to Load It
An r/LocalLLaMA member (66 upvotes, 43 comments) posted a real-hardware datapoint for the 304B V4-Flash-0731 at IQ2_M quantization: roughly 3.5 tokens/sec on two RTX 3060s plus 96GB system RAM, with the caveat it was not a controlled benchmark. The operationally useful detail is that LM Studio refused to distribute weights onto the second GPU while Unsloth Studio did — a loader difference, not a hardware limit, that would otherwise read as an out-of-memory failure. This is the kind of dual-consumer-GPU floor number that never appears in official model cards.
↳ Follow the thread