Fetching from the wire…
Models2026-09-07 · source-backed
Eleven days and about 167 hours on an RTX 5090 covering weight comparison, 13 benchmarks, KL divergence and HarmBench's 400 classic behaviors (abliterlitics.dev). Attack success rates: orcarouter 82.2%, apostate 78.7%, huihui 75.6%, ultra_heretic 70.5%, coder3101 70.0%, blackfrost 68.5%, obliteratus 63.9%, trohrbaugh 57.5%, against a 4.5% base. Capability cost tracks separately: apostate has the lowest KL divergence at 0.0439 while obliteratus sits at 1.5427. The author says orcarouter was the only card where every published claim checked out against the weights.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The letter to Senators Tim Scott and Elizabeth Warren, dated June 10 and surfacing publicly this week, frames it as model distillation run against Claude at scale (Anthropic). A related claim pegs it at 28.8 million fraudulent exchanges, though that figure is single-sourced an...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Alibaba post-trained its flagship in place on September 1, keeping the 2.4T-parameter base and 1M context. All eight published coding benchmarks improved: TerminalBench 3.0 from 11.3 to 29.0, DeepSWE 1.1 from 56.6 to 69.3, QwenSWEbench V2 from 55.1 to 70.0, JobBench from 53.4...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.