Markets
CROSS-CATEGORY: Four Independent Moves in 48 Hours All Attack the Same Line Item, the Per-Token Inference Bill
Between September 8 and 9, Inception shipped Mercury 2.5 at $0.04/$0.15 per million tokens on an 80% launch discount, DeepSeek posted a V4.1 Flash price schedule effective September 10 at $0.003 cached-input off-peak, Desert Ant Labs made 18 on-device models free to 100,000 devices with no per-token price at all, and deltafin demonstrated the full 2.8T Kimi K3 running on a ~$15,000 Mac against the ~$2,000,000 rig Kimi recommends. Four different architectures, one target. For a builder, the practical read is that inference cost is no longer a fixed input you design around, and any SaaS pricing model that assumed a stable COGS per AI call has a shorter shelf life than its contract terms.
Source
↳ Follow the thread