Tim Dettmers' lab opens a release week: Qwen 3.6 35B at 1.5 bits per weight hitting 450 tokens/sec
timdettmers.com (via Hacker News, 168 points)·high signal
Tim Dettmers posted on 2026-09-21 that his lab is putting out two open-source projects and four papers starting the next day. The inference framework targets Mac/Metal and CUDA and claims Qwen 3.6 35B at 1.5 bits per weight running 450 tokens/sec, Qwen 3.8 Flash Next 125B on a single 24GB GPU, and DeepSeek V4.1 550B on AMD Strix, NVIDIA DGX or a 128GB MacBook. Also shipping: an agent harness for long-running autonomous research sessions, an information retrieval system he claims beats frontier-lab deep research, and CliffCompaction, an auto-compaction technique cutting costs roughly 50% for million-token sessions.