Fetching from the wire…
Public story · 2026-09-11 · high
SWE-2 drops to 27.3% on Terminal-Bench 4, about half of Fable 5.1's 55.8% and GPT-6 Astra's 57.9%.
Why now: Cognition posted the SWE-2 benchmarks as teams weigh a cheaper model against frontier reliability on long, multi-step work.
Cognition scored its new coding model, SWE-2, at 50.0% on FrontierCode 1.1 Main, a point behind Fable 5.1's 50.9%. It runs at 64% lower cost than Fable 5.1, per Cognition, changing the math for teams pointing agents at a large backlog of routine tickets.
SWE-2 is trained with reinforcement learning on Kimi K3, a 2.8-trillion-parameter base model, with three effort levels trained in a single run. On Terminal-Bench 2.1 it scores 92.8%, close to frontier-model territory on shorter, tool-use tasks.
The story flips on Terminal-Bench 4, built for longer, multi-step work. SWE-2 scores 27.3% there, less than half of Fable 5.1's 55.8% and GPT-6 Astra's 57.9%. Cognition's post doesn't say why the longer-horizon score falls off so hard.
I'd route routine tickets to SWE-2 and keep long, multi-step work on a frontier model. The FrontierCode near-tie and the Terminal-Bench 4 collapse both point at a model built for short, well-defined tasks over long ones.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.