Fetching from the wire…
Public story · 2026-08-05 · high
Syzygy says it keeps 95% of full-precision performance, and a 3-trillion-parameter version is coming in weeks.
Why now: Syzygy's numbers hit r/LocalLLaMA before any outside lab weighed in, and the August 5 coverage is still just the lab's own claim.
Syzygy Research compressed a 35-billion-parameter model to 1.7 bits per weight while holding onto 95% of full-precision performance, per its own announcement on X.
The result, called Mach-1 Additive, is a tenth the size of the original at 7GB total. Syzygy says it runs at up to 120 tokens per second on a consumer laptop, without a data center GPU.
Those benchmark numbers come from 12 agentic and reasoning tests run against Qwen 3.6 35B, the full-precision model Syzygy started from. Converting an existing model into this format takes under 15 GPU hours of retraining, per Syzygy, far cheaper than training one from scratch. That conversion cost matters for any lab sitting on a model it doesn't want to retrain from zero.
r/LocalLLaMA picked it up at 480 upvotes under the title "why nobody is talking about this."
None of this is independently verified yet. As of August 5, no outside lab has replicated the 95% figure, checked the GPU-hours claim, or benchmarked the 120 tok/s number.
Each link below shares sources, entities, or timing with this story.
Anthropic criticizes Qwen / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic criticizes Qwen); both cover LocalLLaMA, Qwen; overlapping topics (agentic, benchmark, model, weight).
Ollama supports Qwen / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Ollama supports Qwen); both cover GPU, LocalLLaMA, Qwen; earlier GPU coverage from 2026-04-23.
Anthropic criticizes Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic criticizes Qwen); both cover LocalLLaMA, Qwen; overlapping topics (agentic, benchmark, model).
Alibaba released Qwen / Shared entity: LocalLLaMA / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover LocalLLaMA; overlapping topics (benchmark, claim, consumer, weight).
Qwen benchmarked against Claude / Shared entity: Qwen / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover Qwen; overlapping topics (model, weight).
Alibaba released Qwen / Shared entity: Qwen / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover Qwen; overlapping topics (agentic, model).
Alibaba released Qwen / Shared entity: Qwen / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover Qwen; earlier Qwen coverage from 2026-06-28.
Anthropic criticizes Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic criticizes Qwen); both cover LocalLLaMA, Qwen; overlapping topics (model, weight).