OpenAI’s Jalapeño Chip Shows Strong AI Inference Performance in Early Benchmarks

OpenAI reveals Jalapeño chip benchmarks showing improved AI inference speed, efficiency, and scalability compared with current systems.

Aug 25, 2026 - 15:01
 1
OpenAI’s Jalapeño Chip Shows Strong AI Inference Performance in Early Benchmarks
IMAGE CREDITS: OPENAI

OpenAI has shared the first benchmark results for Jalapeño, its custom AI inference chip designed to improve the speed and efficiency of running artificial intelligence models at scale.

The company presented new details about the system at the Hot Chips conference, showing benchmark results from SemiAnalysis’ InferenceX test. OpenAI said Jalapeño demonstrated higher tokens per user and greater throughput per kilowatt compared with currently available inference processors.

Jalapeño targets faster and more efficient AI inference

OpenAI’s hardware team said the chip is designed to handle large-scale AI workloads while reducing power requirements and response delays. Richard Ho, OpenAI’s head of hardware, said the results showed significant performance improvements compared with current state-of-the-art inference systems.

The company’s comparison was made against an Nvidia Blackwell-based system. However, OpenAI expects Jalapeño to enter limited deployment at the end of 2026, with broader deployment planned for 2027, meaning competing hardware platforms may advance before wider adoption.

OpenAI first announced Jalapeño in October and developed the chip in collaboration with Broadcom. The company said its own AI models were also used during the development process.

Custom hardware designed around AI model needs

OpenAI is developing Jalapeño as a multigenerational hardware platform where AI models, software, memory systems, and chips can be designed together.

The company said this approach allows it to address specific challenges in AI inference, particularly delays that occur during the prefill and communication stages of processing.

OpenAI said Jalapeño is designed to reduce unnecessary data movement and communication delays by keeping important model information, including the KV cache used during response generation, closer to the computing resources handling each inference task.

OpenAI expands focus on AI infrastructure

The Jalapeño project reflects OpenAI’s broader effort to build infrastructure optimised for running AI systems efficiently. Instead of relying solely on general-purpose processors, the company is developing specialised hardware to improve performance for its AI workloads.

OpenAI said the system is designed to balance faster response times with the ability to serve large numbers of users, addressing two major requirements for modern AI applications.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav’s current bio says she reports on technology-focused developments “in India”, but the same profile publishes stories about U.S. NHTSA investigations, Hugging Face, global AI startups and other international topics.