Reddit
Inception Labs' diffusion LLM Mercury 2.5 hits 1,107 tok/s at $0.20/$0.75 per million tokens
Mercury 2.5 shipped September 8 running at 1,107 tokens per second on NVIDIA GPUs with a 260K context window, which Inception Labs claims is a 40% intelligence gain over Mercury 2 and comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite and Claude Haiku 4.5. List pricing is $0.20 per million input and $0.75 per million output, currently discounted 80% to $0.04/$0.15 at launch. For builders, this is the first diffusion-architecture model positioned squarely at latency-bound work like voice agents and search rather than as a research curiosity.
↳ Follow the thread