Fetching from the wire…
Public story · 2026-08-26 · high
The chip claims up to 1.9x Nvidia's efficiency at half the power draw, and OpenAI's senior data center executive left the company on August 25.
Why now: OpenAI published these benchmark numbers on August 26, one day after a TechCrunch report detailed the August 25 departure of its senior data center executive.
OpenAI published first benchmark numbers for Jalapeño, its Broadcom-built inference chip, claiming up to 1.9x the throughput per kilowatt of Nvidia's GB300 racks. The gap matters because Jalapeño runs on 700 watts against GB300's 1,400. Nvidia's flagship racks draw twice the power for numbers OpenAI says its chip already beats on throughput and latency.
Tom's Hardware's benchmark report measured Jalapeño on the SemiAnalysis InferenceX suite. OpenAI's numbers there show 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia's GB200 and GB300 rack systems. Jalapeño's single-token output per megawatt also edges past the multi-token figures Nvidia and CoreWeave published for Vera Rubin in July.
Per the companion OpenAI post, Jalapeño went from design to silicon in about nine months. Broadcom handled physical implementation and networking, and Celestica built the boards and racks. It targets prefill and inter-chip communication, the two bottlenecks OpenAI names as dominant in serving. Nine months is a product cycle, not a chip program. If OpenAI can repeat it, that changes who can plausibly build custom inference silicon.
The chip is still at engineering-sample stage while Rubin already ships to customers. Total cost per token between the two comes out roughly even. OpenAI hasn't run larger models like DeepSeek V4 Pro or Kimi K3 on Jalapeño at all. The workload mix behind the headline ratios is narrower than it looks.
OpenAI's senior data center executive left the company on August 25, according to a TechCrunch report on the departure. OpenAI said the move followed a reorganization of its infrastructure group "to support the scale and pace of our work." Turning nine months of engineering samples into racks that customers can order needs that organization intact.
Each link below shares sources, entities, or timing with this story.
Perplexity partners with NVIDIA / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Perplexity partners with NVIDIA); both cover Ben Thompson, July, Kimi K3, LocalLLaMA; reported by the same outlet (stratechery.com).
NVIDIA invested in OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA invested in OpenAI); both cover Apple, July, OpenAI; reported by the same outlet (openai.com).
Linked by a graph relationship (NVIDIA invested in OpenAI); both cover Apple, July, OpenAI; reported by the same outlet (techcrunch.com).
NVIDIA uses Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Ben Thompson, July, OpenAI; reported by the same outlet (stratechery.com).
NVIDIA partners with AWS / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA partners with AWS); both cover Apple, NVIDIA, OpenAI; reported by the same outlet (techcrunch.com).
NVIDIA uses TSMC / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA uses TSMC); both cover Broadcom, Nvidia, September; overlapping topics (broadcom, nvidia).
Anthropic uses NVIDIA / Shared entities / Earlier coverage
Linked by a graph relationship (Anthropic uses NVIDIA); both cover Apple, LocalLLaMA, NVIDIA, OpenAI; earlier Apple coverage from 2026-08-10.
NVIDIA invested in OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA invested in OpenAI); both cover Apple, OpenAI; reported by the same outlet (techcrunch.com).