Fetching from the wire…
Models2026-09-14 · source-backed
The API changelog carries it verbatim: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." Every deepseek-v4-pro request was scheduled to be answered by V4.1-Flash at Flash rates from 04:00 UTC today. V4 Flash and V4 Flash Vision Exp are still retired with their names aliased to V4.1 Flash, so anyone who rewrote model strings for the cutover should work out which of the two changes applied to them.
Each link below shares sources, entities, or timing with this story.
DeepSeek posted a community notice: once V4.1 Flash launches around September 10 Beijing time, and until a V4.1 Pro exists, every V4 Pro request routes to V4.1 Flash and bills at Flash unit pricing. The stated reason is that Flash has surpassed Pro on performance, cost, speed...
384 tokens maximum per image after automatic resizing to roughly 800x800, up to 600 images per request, 32 MiB per image inline or 64 MiB via the Files API, 48 MiB request body, 8192 px per side dropping to 4096 px once a request carries 15+ images. Images accepted only in use...
DeepSeek's docs confirm deepseek-chat and deepseek-reasoner fully retire after July 24, 15:59 UTC. Migrate to explicit V4-Pro (1.6T total / 49B active) or V4-Flash (284B / 13B) names or face broken calls. If you have DeepSeek in production, that's a hard ops deadline, not a su...
The changelog lists Terminal Bench 2.1 at 82.7, NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon verified 70.3, DSBench-FullStack 68.7 and DSBench-Hard 59.6: figures DeepSeek says far exceed V4-Pro-Preview. Native Responses API support and specific Codex adaptation. Only...
Latent Space's September 12 roundup collects reaction to the causal encoder-decoder release: 763B total parameters, 8B active at prefill and 16B at decode, 1M context, KV cache cut to about 890 bytes per token, roughly one eighth of V4 Pro. Sebastian Raschka argues the archite...
The model card went up at 02:17 UTC on September 10, MIT-licensed, multimodal MoE with a 1M-token context. The Causal Encoder-Decoder architecture projects the decoder's global KV cache from final encoder hidden states rather than per-layer, and SWA Bounded Replay cuts persist...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.