Fetching from the wire…
Models2026-07-28 · source-backed
A Qwen3-8B derivative retrained with ternary quantization-aware training, storing all 252 transformer linears in a ternary format one-eighth the size of fp16, decoded inside the matmul kernels. 8.19B parameters in a 2.56 GB download with 40,960-token context, scoring 72.1 MMLU, 80.2 IFEval-strict, 53.4 GSM8K, 68.9 BFCL v3, running 763 tok/s on an H100 with speculative decoding versus 33.7 tok/s on an Apple M5 via MLX. No access request. Weights at FermionResearch/Neutrino-8B with 0.6B variants. (Fermion Research)
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
Google released open-source Multi-Token Prediction (MTP) drafters for the Gemma 4 model family. The concept: pair a heavy target model (Gemma 4 31B) with a lightweight drafter that predicts several future tokens in parallel. The target model verifies the predictions in a singl...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Z.ai's 744B MoE is being ranked as July's strongest open-source model at a reported 91.2% GPQA Diamond and 62.1% SWE-bench Pro, at a fraction of frontier API pricing, and it's listed in Ollama's supported-model line. That's the distinction that matters this week: Kimi K3's wei...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.