Voices
Raschka on DeepSeek v4.1-Flash: "should have been called DeepSeek v5"
Latent Space's Sep 12 roundup collects practitioner reaction to the causal encoder-decoder release: 763B total parameters with only 8B active at prefill and 16B at decode, a 1M context, and a KV cache cut to roughly 890 bytes per token, about one eighth of V4 Pro. Sebastian Raschka argues the architectural break warrants a major version number, and Fraser Price reports 300+ TPS on four RTX Pros at full precision with under 32GB peak system RAM using SSD offloading. TeortaxesTex supplies the counterweight, questioning whether DeepSeek ships internal research artifacts rather than products, and the model is verbose at ~89k tokens per task, 25-62% more than competitors.
Source
↳ Follow the thread