Monday, September 14, 2026
AINews

OpenAI Publishes First Benchmarks for Its In-House Jalapeño Inference Chip

OpenAI has published the first measured results of Jalapeño, its first custom inference chip, saying it can serve more AI work per unit of power while also cutting response latency. The company said the chip delivers both higher throughput and lower latency with a single architecture, where existing hardware systems often have to trade one against the other.

Key Takeaways

  • OpenAI published the first measured results of Jalapeño, its first custom inference chip
  • Across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, it delivered 1.5 to 1.9 times more work per watt at peak throughput
  • End-to-end latency was 1.7 to 3.6 times lower than comparison systems
  • OpenAI plans to deploy the chip in its compute infrastructure before the end of 2026

Dual gains in power efficiency and latency

OpenAI said Jalapeño offers higher throughput and lower latency at once, where existing systems often force a choice between the two. For customers, it said, this can mean faster responses, more responsive agents and more reliable access as demand grows.

Jalapeño’s performance spans both OpenAI and third-party models. In testing across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, the chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems. For highly interactive workloads, performance was 2.1 to 4.1 times higher.

On Kimi, the largest public model tested, Jalapeño delivered about 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. OpenAI said its advantage widened further on frontier OpenAI models in internal testing, suggesting the architecture becomes more valuable as workloads grow.

Multi-generation roadmap and deployment plan

OpenAI described Jalapeño as evidence of a full-stack advantage, saying it can design models, products, serving software, chips, memory, networking and systems together, using real-workload observations to improve every layer. The chip supports ultra-fast, fast and batched inference modes, pushing each mode’s efficiency to a higher level.

OpenAI plans to begin deploying Jalapeño within its compute infrastructure before the end of 2026. It called the chip the first generation of a multi-generation roadmap, with Gen 2 deep in development and Gen 3 taking shape. OpenAI said it will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference.

Ahead of deployment, OpenAI continues production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models.

Source: OpenAI

Cover image generated by AI for illustrative purposes only.