Fetching from the wire…
Public story · 2026-07-19 · high
It's also third on DeepSWE, the first open-weights model to reach frontier level there, per AINews.
Why now: AINews surfaced these throughput and benchmark numbers in its July 19 roundup on Latent Space, ahead of any independent throughput testing.
K3's Kimi Delta Attention architecture claims up to 6x cheaper throughput at a 1 million token context window, per AINews's roundup on Latent Space.
That number matters more than where K3 lands on any leaderboard. Long-context agent loops call a model over and over inside the same task. A throughput multiple compounds across a run in a way a single benchmark score doesn't. A 6x cut at 1M context is the difference between a long-context agent workflow that's affordable and one that blows a budget.
AINews also has K3 sitting third on DeepSWE, and calls it the first open-weights model to reach frontier level on that benchmark. Artificial Analysis separately scores K3 at 64% on DeepSWE and 84% on Terminal-Bench v2, per the same rundown.
AINews doesn't say what baseline the 6x figure is measured against, or what K3 costs per token in dollar terms. That leaves the throughput claim directional until someone runs it under real load.
K3 is a bet that throughput economics beat leaderboard position for agent-heavy workloads. If the 6x figure holds up outside AINews's report, it undercuts frontier models priced by the token for anything that runs long.
Each link below shares sources, entities, or timing with this story.
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
Latent Space's AINews breakdown puts Opus 5 at 159 on the Epoch Capabilities Index against Fable 5's 161, but at SWE-ECI 161 they're at parity on software engineering specifically. Artificial Analysis separately reported the ~150 Elo gap at 20% lower cost per task. Users flagg...
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardo...
AINews published the hard placement numbers: K3's Coding Agent Index of 57 matches GPT-5.6 Terra and GPT-5.5, and it ranks #3 among open-weight models on DeepSWE. An open-weight model sitting above Opus 4.8 on the aggregate index is the first quantified read on how close the g...
AINews resumed publishing after its post-Kimi-K3 blackout with a "not much happened today" edition (Latent Space). That's a real signal after GPT-5.6 Sol, Grok 4.5, Meta Muse, Kimi K3, and the Qwen3.8-Max preview all landed inside a fortnight. Matthew Berman published another...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.