Research
NVIDIA Puts Groq 3 LPX Into Full Production, Claiming 3,400 Output Tokens Per Second and 4x Agent Responsiveness
Announced at Hot Chips on 2026-08-24, Groq 3 LPX is an interactive inference accelerator extending the Vera Rubin NVL72 platform, aimed specifically at the token-generation phase that determines how responsive an agent loop feels. NVIDIA cites 3,400 output tokens per second running Gemma 4 31B at 100,000-token context and claims 4x faster agent responsiveness than the nearest alternative platform, with agentic coding tasks completing in minutes rather than hours. Nebius is the first AI cloud to adopt it, bringing it to Nebius Token Factory through its existing API, with Groq itself among the earliest planned adopters; no pricing was disclosed.
↳ Follow the thread