Fetching from the wire…
Public story · 2026-09-10 · high
Independent benchmarks put the open model two points behind GPT-5.6 Sol and two ahead of Claude Opus 5 on terminal coding tasks.
Why now: AWS published the walkthrough dated September 10, 2026.
AWS published a step by step guide for running Qwen3.8's 2.4T-parameter model on SageMaker HyperPod, using vLLM and NVFP4 quantization to serve it through an OpenAI-compatible endpoint with reasoning support. Teams weighing self-hosting against an API now have a number to work with: Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, two points behind GPT-5.6 Sol's 88.8 and two points ahead of Claude Opus 5's 84.6, per independent benchmark compilations.
AWS's HyperPod walkthrough covers cluster provisioning end to end. The model has 95B active parameters in a mixture-of-experts setup, so this isn't something you point at a single GPU and walk away from.
The model is Apache 2.0 licensed. That license means any team can pull the weights and run this exact recipe on its own cluster, no vendor approval needed.
For teams with a compliance reason to keep weights local, the math used to mean eating a real capability tax to get there. Data residency rules, audit requirements, or just not wanting a model to disappear behind a price change are the kinds of reasons that used to cost you frontier performance. This walkthrough is evidence that tax has shrunk to nearly nothing for at least one workload.
What the post doesn't cover is cost. Running NVFP4 quantization on a cluster big enough to serve a 2.4T model isn't free. AWS's guide gets the deployment working; it says nothing about per-token cost against the closed alternatives. Anyone doing this math needs to run their own GPU bill before assuming self-hosting wins on price and not just on control.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The full family, Sol, Terra, and Luna, is generally available on Amazon Bedrock with IAM and VPC controls (LLM Boss, AWS). Sol targets coding, biology, and cybersecurity agentic work. Terra runs everyday tasks at about half GPT-5.5's cost, and Luna optimizes for speed. The thr...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.