Reddit
A single-request 600 tok/s run of Qwen3.6 35B-A3B on one RTX Pro 6000 using Ninfer
An r/LocalLLaMA user posted video of Ninfer serving Qwen3.6 35B-A3B at 600 tok/s on a single request on one RTX Pro 6000, describing it as a deliberate brute-force tier: a weaker model used for read-and-find and code tasks where spending 20x the tokens still finishes faster than a smarter local model. The framing is the useful part for builders sizing local agent fleets, since it treats throughput and intelligence as a tradeable pair rather than picking the best model available. Single-source and self-reported, with no benchmark harness or reproducible config published.
↳ Follow the thread