Fetching from the wire…
Public story · 2026-08-05 · high
Syzygy says it keeps 95% of full-precision performance, and a 3-trillion-parameter version is coming in weeks.
Why now: Syzygy's numbers hit r/LocalLLaMA before any outside lab weighed in, and the August 5 coverage is still just the lab's own claim.
Syzygy Research compressed a 35-billion-parameter model to 1.7 bits per weight while holding onto 95% of full-precision performance, per its own announcement on X.
The result, called Mach-1 Additive, is a tenth the size of the original at 7GB total. Syzygy says it runs at up to 120 tokens per second on a consumer laptop, without a data center GPU.
Those benchmark numbers come from 12 agentic and reasoning tests run against Qwen 3.6 35B, the full-precision model Syzygy started from. Converting an existing model into this format takes under 15 GPU hours of retraining, per Syzygy, far cheaper than training one from scratch. That conversion cost matters for any lab sitting on a model it doesn't want to retrain from zero.
r/LocalLLaMA picked it up at 480 upvotes under the title "why nobody is talking about this."
None of this is independently verified yet. As of August 5, no outside lab has replicated the 95% figure, checked the GPU-hours claim, or benchmarked the 120 tok/s number.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.