DeepSeek Ships v4.1 Flash and Will Auto-Route All V4-Pro API Traffic to It on September 14
DeepSeek launched deepseek-v4.1-flash around September 10, 2026 with off-peak pricing of $0.003 per million cached input tokens, $0.15 cache-miss input and $0.60 output, doubling during peak hours (01:00-04:00 and 06:00-10:00 UTC weekdays). Reported benchmarks put it at 88.1 on CyberGym versus 84.5 for GPT-5.6 Sol and GLM 5.3, 74.2 on DeepSWE v1.1 versus Claude Opus 5's 74.0, and 54.8 on Automation-Bench versus 50.3, while trailing badly on knowledge (HLE 36.8 versus 56.3). The detail builders should act on is the forced migration: from September 14 at 12:00 Beijing time, V4-Pro API requests get silently routed to V4.1 Flash at Flash pricing, which the top HN comments object to because validated production workflows get swapped models without consent.
↳ Follow the thread