
The neocloud company is using debt to fund Nvidia GPU infrastructure for major customers.
OpenAI’s first in-house inference chip, Jalapeño, reportedly outperforms Nvidia’s Blackwell and Rubin systems in key inference benchmarks, according to results discussed at Hot Chips and analyzed by SemiAnalysis.

OpenAI showed benchmarks for Jalapeño, its first in-house inference chip, at the Hot Chips conference. The chip is built for inference only, meaning it runs AI models rather than training them. The Decoder reports that Jalapeño is positioned as a general-purpose LLM inference accelerator rather than a chip tuned only for OpenAI’s own models.

According to SemiAnalysis testing cited by The Decoder, Jalapeño beats Nvidia’s Blackwell and even Rubin in throughput and energy efficiency. OpenAI claims the chip delivers 1.5x to 1.9x more AI work per watt at peak throughput across three tested models. The company also claims 1.7x to 3.6x lower end-to-end latency than the best commercially available systems, and 2.1x to 4.1x higher performance for interactive workloads.
The tests used SemiAnalysis’s public InferenceX benchmark, with OpenAI providing the numbers and SemiAnalysis verifying some runs on-site in the lab. The tested models were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T; Jalapeño reportedly reached about 1,400 tokens per second per user on GPT-OSS and more than 700 tokens per second on a single concurrent Deepseek R1 request. The Decoder notes important limits: Nvidia and AMD have published results on larger models not yet tested on Jalapeño, and Rubin systems are already shipping while Jalapeño reportedly remains at the engineering-sample stage.
OpenAI developed Jalapeño with Broadcom, and the company says it used its own AI models during development. The reported speed of development and the benchmark results raise questions about how durable Nvidia’s software and hardware advantage remains. OpenAI CFO Sarah Friar frames the chip as part of a broader compute strategy that complements partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and others rather than replacing them.

The neocloud company is using debt to fund Nvidia GPU infrastructure for major customers.

Z.ai’s new open-source model pairs low task cost with Chinese-chip inference.
More than 100 organizations joined OpenAI’s call to strengthen cyber defenses against AI-enabled threats.

The reported $12.9B deal strengthens Nvidia’s open AI push as closed labs move toward custom chips.