Syzygy Research's Mach-1 Additive Runs a 35B Model Without Ever Multiplying by a Weight — 1.7 Bits/Weight, 95% of Qwen 3.6 35B, 7GB
Syzygy Research announced Mach-1 Additive, a 35B-parameter model that performs inference without a single weight multiplication, at 1.7 bits per weight. They report recovering 95% of full-precision Qwen 3.6 35B performance across 12 agentic and reasoning benchmarks while being 10x smaller — 7GB total, running up to 120 tokens/sec on consumer laptops — and claim the conversion needs under 15 GPU hours of retraining. r/LocalLLaMA surfaced it at 480 upvotes / 137 comments with the framing 'why nobody is talking about this,' which is the honest read: the claims are impressive and currently rest on the lab's own announcement. Syzygy says it will release models up to 3 trillion parameters compressed with the same algorithm over the coming weeks — that is the claim to watch, because 1.7-bit additive inference at trillion scale would reset what 'runs locally' means.
↳ Follow the thread