Fetching from the wire…
Models2026-09-19 · source-backed
Published to GitHub September 18, one day after the Science paper, with checkpoints on Hugging Face and a Zenodo archive. The vision-language model reads contrast-enhanced CT across 18 abdominal organs and scored an average AUC of 0.913 on 146 clinical findings across about 40,000 real examinations, outperforming 23 of 26 radiologists on average in a study. Training used more than 420,000 abdominal CT examinations and over 15 million anatomy-focused image-text pairs. SCMP Weights plus training framework under a permissive license is the unusual part, not the accuracy.
Each link below shares sources, entities, or timing with this story.
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Per a roundup, ZAYA1-8B is an Apache-2.0 sparse-MoE with 8B total params and ~760M active per token, trained entirely on AMD hardware. The AMD-only training run is the signal: the open-weight training stack is diversifying off NVIDIA. Verify the numbers against the official mo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.