Fetching from the wire…
Public story · 2026-09-06 · high
One of three templates tested stayed flat at 94% regardless of reasoning effort while the top scorer paid over 4 hours to reach 99%.
Why now: The sweep posted to r/LocalLLaMA on 2026-09-06 tested three templates across two reasoning efforts each, isolating the template as the one variable.
A 100-task sweep found the same model swing 8 points on SWE-bench Verified, depending only on which chat template served it. Anyone deploying this model through an agent harness picks a template blind to which behavior it produces. That blind choice costs hours of wall time.
The 100-task template sweep ran Qwen3.8 Flash Next through mini-SWE-agent 2.4.6 on an RTX PRO 6000 workstation. It used sglang, NVFP4 weights, and the full 262K context window.
Three templates, two reasoning efforts each. Stock scored 91% at medium reasoning and 99% at xhigh. Fixed came in lower both times, 87% to 98%. Sharp didn't move: 94% at either effort.
Stock's 99% costs more: +143.5% more output tokens and 4 hours 31 minutes of wall time, against 1 hour 47 minutes at medium. Sharp's climb to 94% costs +28.3% more tokens, with no reasoning-effort dial to turn.
Stock at xhigh beats Sharp by 5 points if a task can wait 4 hours 31 minutes. Sharp's flat 94% beats stock's 91% at medium when it can't, without a reasoning-effort tax.
The run doesn't say why Sharp resists reasoning effort while the other two respond to it. That leaves an open question for anyone building on this model.
Each link below shares sources, entities, or timing with this story.
Kwaipilot's release is 35B total / 3B activated, built on Qwen3.6-35B-A3B, Apache 2.0, 262,144-token context, posting 69.40% on SWE-bench Verified, 63.00% multilingual, 45.96% SWE-bench Pro, 41.02% Terminal-Bench 2.1. An r/LocalLLaMA user reports it one-shotting a five-level T...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Somebody finally measured the thing everyone complains about, and the numbers are worse than the vibes. A Level1Techs writeup that hit 384 points and 144 comments on Hacker News captured full-vocabulary logits and computed KL divergence in FP64 to trace exactly where local inf...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.