Fetching from the wire…
Models2026-09-16 · source-backed
GSQ-RCO GGUFs for Qwen3.8-Flash-Next cut the model from roughly 80-95GB to 68-76GB, with IQ3_XXS matching the base model exactly on AIME25 and within 0.51 on GPQA-Diamond. The unusual variant is Q2_0, which trades 0.09 points of task average for 3.4x prompt throughput and 1.9x lower end-to-end latency against IQ2_XS by keeping decode cost flat. One commenter on dual 3090s reports it's currently slower for them because tensor parallel and MTP support are missing.
Each link below shares sources, entities, or timing with this story.
Two methods: GSQ (Gumbel-Softmax Quantization), post-training scalar quantization that jointly learns grid assignments and scales at 2-3 bits, and RCO (Riemannian Constrained Optimization), assigning a quant type per tensor under a strict size budget by gradient descent on the...
Created August 24, it holds a trendingScore of 3,967 against second-place GLM-5.3-Flash at 1,376 (Hugging Face). The near-1:1 like-to-download ratio means almost everyone bookmarking it hasn't pulled weights, and the unsloth GGUF conversion at 4,354 downloads is absorbing comp...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Edge0-AI/Edge0, created September 8, went from 269 to 583 stars in two days. It packages SSD expert offload, Recover-LoRA and prerouter routing prediction into an MLX-backed framework: edge0-35b is a 4-bit 40-layer 256-expert model built on Qwen3.5-MoE 35B-A3B needing about 2....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.