Fetching from the wire…
Public story · 2026-08-25 · high
The interactive inference accelerator claims 3,400 output tokens a second, and Nebius is the first cloud to run it.
Why now: Nvidia disclosed the production milestone on August 24 at Hot Chips, its venue for staking inference-speed claims.
Nvidia put Groq 3 LPX into full production on August 24, claiming 3,400 output tokens per second, according to Nvidia's announcement at Hot Chips.
LPX targets the token-generation phase, the step that decides how responsive an agent feels. That phase is what stalls while a loop waits on a multi-step tool call. Nvidia calls it an interactive inference accelerator and says it runs 4x faster on agent responsiveness than the nearest alternative.
The 3,400-tokens-per-second figure comes from running Gemma 4 31B at a 100,000-token context. LPX extends the Vera Rubin NVL72 platform.
Nebius is the first cloud provider to put it into production. Nvidia hasn't disclosed pricing, so what that speed costs at scale is still unknown.
That figure is Nvidia's own benchmark, run on Nvidia's own comparison set. It stays a marketing number until an independent test reproduces the 4x agent-responsiveness claim against a workload nobody at Nvidia picked. Watch for Nebius or a third party to publish that number once LPX is live outside Nvidia's demos.
Each link below shares sources, entities, or timing with this story.
NVIDIA released Vera Rubin NVL72 / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA released Vera Rubin NVL72); both cover August, NVIDIA, Vera Rubin NVL72; earlier August coverage from 2026-08-24.
Vera Rubin NVL72 competes with Blackwell / Shared entity: NVIDIA / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Vera Rubin NVL72 competes with Blackwell); both cover NVIDIA; reported by the same outlet (nvidianews.nvidia.com).
Ollama supports Gemma / Shared entities / Earlier coverage
Linked by a graph relationship (Ollama supports Gemma); both cover August, Gemma, NVIDIA; earlier August coverage from 2026-08-12.
NVIDIA released Vera Rubin NVL72 / Shared entities / Shared topic / Tension
Linked by a graph relationship (NVIDIA released Vera Rubin NVL72); both cover Nvidia, Vera Rubin NVL72; overlapping topics (agent, token).
NVIDIA released Vera Rubin NVL72 / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA released Vera Rubin NVL72); both cover August, Nvidia; overlapping topics (agent, august).
DFlash uses Gemma / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (DFlash uses Gemma); both cover August, NVIDIA; overlapping topics (agent, token).
NVIDIA partners with AWS / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (NVIDIA partners with AWS); both cover Nebius, Vera Rubin NVL72; reported by the same outlet (nvidianews.nvidia.com).
NVIDIA released Vera Rubin NVL72 / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA released Vera Rubin NVL72); both cover August, NVIDIA; earlier August coverage from 2026-08-20.