Fetching from the wire…
Models2026-09-13 · source-backed
Latent Space's September 12 roundup collects reaction to the causal encoder-decoder release: 763B total parameters, 8B active at prefill and 16B at decode, 1M context, KV cache cut to about 890 bytes per token, roughly one eighth of V4 Pro. Sebastian Raschka argues the architectural break warrants calling it v5. Fraser Price reports 300+ TPS on four RTX Pros at full precision with under 32GB peak system RAM using SSD offloading. TeortaxesTex supplies the counterweight, questioning whether DeepSeek ships internal research artifacts instead of products, and the model is verbose at ~89k tokens per task, 25-62% more than competitors. That verbosity is the number to hold against the KV-cache win before you price a workload.
Each link below shares sources, entities, or timing with this story.
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
DeepSeek posted a community notice: once V4.1 Flash launches around September 10 Beijing time, and until a V4.1 Pro exists, every V4 Pro request routes to V4.1 Flash and bills at Flash unit pricing. The stated reason is that Flash has surpassed Pro on performance, cost, speed...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
DeepSeek dropped V4 in mid-June as an open-weight model with a 1-million-token context window, priced at $1.74 per million input tokens, posting near-parity with GPT-5.4 on math and Q&A benchmarks (MindStudio). That's the headline number. The architecture underneath is more in...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
The model card went up at 02:17 UTC on September 10, MIT-licensed, multimodal MoE with a 1M-token context. The Causal Encoder-Decoder architecture projects the decoder's global KV cache from final encoder hidden states rather than per-layer, and SWA Bounded Replay cuts persist...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.