Fetching from the wire…
Models2026-09-07 · source-backed
A practitioner running 2x Strix Halo 128GB over USB-C 4 with llama.cpp RPC compared both models at Q8_K_XL on real coding work (r/LocalLLaMA). A task Qwen finished in 25 minutes on medium took DeepSeek 12. The poster attributes the gap to fewer hallucinated detours. The sharpest data point is a failure: Qwen3.8-Flash-Next on xhigh failed to complete in about three hours the task it finished in 25 minutes on medium, so raising reasoning effort on that model is counterproductive for agentic coding.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
A 232-upvote r/LocalLLaMA thread builds an open-weights argument on Ramp's mid-August corporate spending data: Fable 5, the most capable and expensive model in the lineup, accounts for 11% of those companies' Anthropic spend. (r/LocalLLaMA) The thread's read is that Qwen, GLM...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.