
OpenAI details performance benchmarks for custom Jalapeño inference processor
OpenAI has shared initial benchmark results for its custom Jalapeño inference chip at the Hot Chips conference. The hardware demonstrates significant efficiency gains over current state-of-the-art systems.
Published by Jin · 2 min read · 26 AUG 2026
- 700 watts
- 550 watts or below
- 9 months
- Gen 2 is deep in development, and Gen 3 is taking shape
| Metric | Jalapeño | Comparison System (GB200) |
|---|---|---|
| Package TDP | 700 W | 1,200 W |
| Peak mixed TPS / kW | 85,448 | 44,960 |
| End-to-end latency | 1.03 s | 1.80 s |
| Min TBT | 0.69 ms (1,459 tok/s/user) | 1.87 ms (535 tok/s/user) |
At the Hot Chips conference, OpenAI shared a more detailed look at its custom inference chip, known as Jalapeño, alongside the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, the hardware registered both higher token output per user and greater throughput per kilowatt compared to currently available state-of-the-art inference processors.
Richard Ho, OpenAI’s head of hardware, noted that the results reflect a substantial performance advance. The architecture is designed to serve higher volumes of AI workloads per unit of power while simultaneously maintaining low latency for end users.
Competitive landscape and deployment timeline
The preliminary comparisons were made against Nvidia Blackwell systems. However, the competitive landscape is expected to shift by the time Jalapeño reaches full deployment. Ho estimated that the chip will be deployed in very small volumes by the end of 2026, with more significant scaling anticipated through 2027.
First announced in October, Jalapeño was developed in close collaboration with Broadcom, utilizing OpenAI’s own models to assist in the design process. The organization intends for Jalapeño to serve as a multigenerational platform, synchronizing the development of AI products, models, silicon, and memory.
Addressing inference bottlenecks
This full-stack approach allowed the engineering team to target specific phases of the inference pipeline that traditionally introduce friction. Jalapeño is specifically optimized to minimize delays during the prefill and communication phases, which frequently act as system bottlenecks.
By reducing data movement and communication delays, the system allows model states, including the KV cache, to remain localized. This ensures that compute, memory, and networking resources are activated efficiently for each distinct phase of the inference process.
Source — Original announcement ↗
Worth a read?
Comments · 0