Dispatch
NVIDIA claims Vera Rubin NVL72 delivers 30x throughput per megawatt on real recorded agentic coding sessions
NVIDIA published results on 2026-08-24 claiming Vera Rubin NVL72 hits 30x higher throughput per megawatt and 35x lower cost per million tokens than GB300 NVL72, measured on the SemiAnalysis AgentX benchmark using recorded real-world agentic coding sessions with actual context growth, tool calls, and sub-agent spawning preserved. Models tested included DeepSeek V4 Pro and Qwen3.5. NVIDIA says results are still pending SemiAnalysis review and gives no GA date, so treat the multiples as vendor-reported, but the benchmark choice is the notable part: inference economics are now being marketed on agent sessions rather than single-turn tokens.
↳ Follow the thread