Reddit
A Legal-Document Rig Runs Qwen3.8-27B in BF16 at 262K Context Out of a Lunchbox, 1,715 tok/s Prefill and 45 tok/s Generation
An r/LocalLLaMA user detailed a portable build for 200K+ token legal OCR and summarization: FormD T1 case, Minisforum BD770i SE, Ryzen 7745HX, 96GB DDR5 SODIMM and a 96GB RTX Pro 6000 Blackwell, hitting 77GB VRAM at full 262K context plus MMPROJ. Measured throughput is 1,715 tok/s prefill (175K tokens in 102 seconds) and 45 tok/s generation with MTP, the slow half being the price of BF16. The claim worth testing is that BF16 beats even UD Q8_K_XL when legal precision matters, and that the rig outperforms Gemini Pro and ChatGPT 5.6 Sol on this task, though not Opus, partly because sampler settings are under his control.
Source
↳ Follow the thread