Fetching from the wire…
Public story · 2026-09-10 · high
The model card also shows DeepSeek's own harness beating Claude Code and Codex on the same weights.
Why now: The card posted at 02:17 UTC on September 10, hours before any independent test of the scaffold numbers could run.
DeepSeek posted a model card for V4.1-Flash at 02:17 UTC on September 10, releasing the weights under an MIT license.
The model activates 8B parameters on prefill and 16B on decode out of a 552B backbone. Its persistent KV cache footprint drops to about an eighth of V4-Flash's, which matters for anyone running the full 1M-token context.
The drop comes from a Causal Encoder-Decoder design. It projects the decoder's global KV cache from the encoder's final hidden states, instead of building that cache layer by layer. DeepSeek pairs it with a second technique it calls SWA Bounded Replay.
On DeepSeek's own numbers, V4.1-Flash scores 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE v1.1, and 3471 on Codeforces.
The card also runs a scaffold comparison on the same weights. DeepSeek's minimal in-house harness scores 90.6, ahead of Claude Code at 88.0 and Codex at 84.1. DeepSeek built that harness and ran the grading itself, against two scaffolds it doesn't control. That's a vendor picking the ruler it gets measured with, not an independent test.
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA post at 1,330 upvotes reports the first run of full K3, Moonshot's 2.8T open-weight MoE, on a 16x NVIDIA GB10 cluster with dspark speculative decoding: 20+ tok/s average, 38 peak, 750 prefill. That's roughly $64K of hardware for frontier-adjacent tokens at your...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
A Chinese lab shipped a runtime that manages two American coding agents as subagents, and it went from repo creation to 145,439 stars in four days. deepseek-ai/deepseek-harness published dsh-v0.1.0-rc.7 at 12:01 UTC today, its first tagged release since the repo appeared on Au...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Z.ai's GLM-5 (744B/40B MoE, MIT license, 205K context) is free on NVIDIA NIM at 40 req/min with no credit card. Benchmarks: 77.8% SWE-bench Verified (highest open-source), 56.2 Terminal-Bench 2.0 (approaching Opus 4.5's 59.3). Trained entirely on 100,000 Huawei Ascend chips. Y...
Cline released @cline/sdk on May 13, an open-source TypeScript agent runtime that powers their CLI, VS Code, and JetBrains extensions. Running claude-opus-4.7, Cline CLI scores 74.2% on Terminal-Bench 2.0. Claude Code on the same model: 69.4%. Same model. Different harness. Al...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.