Fetching from the wire…
Models2026-09-25 · source-backed
UkisAI released Swift1.5 27B (-58.5% thinking tokens, +0.35% score), Swift Flash Next (-63.4% tokens, 1.8x faster, -0.2% at xhigh) and an experimental Swift Bonsai 2 (-39.8%), trained by penalizing overthinking patterns then restoring accuracy with GSPO RL and on-policy distillation. (r/LocalLLaMA) A separate community Aider eval (2 runs, Q8_0, llama.cpp 0.5.0) found bottlecapai's ThinkingCap-Qwen3.8-27B matched vanilla exactly at 27.1% first-try and 77.6% retry pass, using 7,436 median tokens against 12,547 and 777s per case against 1,481. Swift scored 30.8%/75.7% at 750s. One daily user reported Swift falling into loops more often than vanilla Qwen, which is the failure mode to watch for.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
UkisAI post-trained Qwen 3.8 27B by identifying tokens tied to overthinking and penalizing those specifically instead of capping reasoning length, then repaired accuracy with on-policy distillation (r/LocalLLaMA). Reported 58.3% fewer thinking tokens, 1.95x speedup, under 1% a...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.