OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
By Pradhyuman,
OpenAI shared benchmark results for its Jalapeño chip at the Hot Chips conference. In tests on the InferenceX benchmark by SemiAnalysis, the chip registered more tokens per user and more throughput per kilowatt than an Nvidia Blackwell system. Richard Ho, the head of hardware at OpenAI, stated that the chip serves more work per unit of power and returns responses faster. OpenAI developed the chip with Broadcom, and used its own AI models to assist in the design process.
Design and deployment timeline
The chip is designed to limit data movement and communication delays. In a company blog post, OpenAI wrote that the system keeps the model state local during processing. Ho estimated that OpenAI will deploy the chip in small volumes at the end of 2026, with a larger deployment planned for 2027.
Maintained by Pradhyuman.
Filed under: OpenAI, Chips & Infrastructure, Models & Research