Reddit
OpenAI Published Jalapeño's First Benchmarks: 1.5x to 1.9x More Throughput Per Kilowatt Than Nvidia's GB300, at Half the Power
OpenAI posted first results for Jalapeño, its Broadcom co-developed inference ASIC, claiming 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 rack systems on the SemiAnalysis InferenceX suite, from a 700W part against Nvidia's 1,400W flagship. Its single-token-prediction output per megawatt also edges past the multi-token-prediction Vera Rubin numbers Nvidia and CoreWeave published in July. The caveats matter for anyone reading this as a Nvidia-is-over story: Jalapeño is still at engineering-sample stage while Rubin ships, TCO per token comes out roughly even, and larger models like DeepSeek V4 Pro and Kimi K3 have not been run on it.
↳ Follow the thread