Tim Dettmers' dlab ran an open-source week of 2 frameworks and 4 papers for frontier-class AI on your own hardware, plus a bitsandbytes2 beta
Tim Dettmers·medium signal
Dettmers' CMU lab released work on compression, context compaction, agent harnesses, autonomous research and test-time scaling, all aimed at local hardware. He also opened a private beta of bitsandbytes2, built around runtime dynamic compression of Mixture-of-Experts, and says it runs a quantized Qwen model at 450 tokens/s at 1.5 bits per weight. The post drew 180 HN points and LavX News covered it separately. The 450 tok/s figure is unreplicated beta data.