Fetching from the wire…
Public story · 2026-09-20 · high
The 600B-parameter model costs about an eighth as much per task, but StepFun's own benchmarks show a 24-point gap on terminal coding tasks.
Why now: StepFun published Step 5 Preview's scores and pricing as of September 20, with full weights due October 15.
StepFun released Step 5 Preview, a 600 billion parameter sparse mixture-of-experts model with 27 billion active parameters and a 1 million token context window. It's live through StepFun's products and API, with full weights due October 15.
Step 5 scores 44 on the Artificial Analysis Intelligence Index, at about one-eighth Claude Opus 5's per-task cost, according to StepFun's release page. For teams running high volume inference, that ratio changes which model gets the default slot.
StepFun published its own gaps alongside the price. Step 5 scores 67.7% on DeepSWE v1.1 against GPT-6 Astra's 74.1%. On StepCodeBench it scores 49.0% against Opus 5's 63.9%. Those are real deficits, but ones a lot of workloads can absorb.
Terminal-Bench v4 shows the widest gap: 33.3% for Step 5 against Astra's 57.9%. A 24-point deficit on terminal work, running commands, reading output, recovering from a failed one, sits right where a coding agent burns through budget on retries. An eighth of the price stops looking cheap once the agent needs three or four extra turns to finish a task Astra completes in one.
Step 5 fits work that doesn't loop through a shell: chat, summarization, single-shot generation where the 1 million token context matters more than agentic reliability. For terminal-heavy coding, the benchmark gap is StepFun's own number to explain.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
The walkthrough covers cluster provisioning, NVFP4 quantization and an OpenAI-compatible endpoint with reasoning support for the 95B-active MoE. Independent benchmark compilations put Qwen3.8-Max at 86.6 on Terminal-Bench 2.1 between GPT-5.6 Sol at 88.8 and Claude Opus 5 at 84...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.