OpenAI Publishes First Benchmarks for Its Jalapeño AI Chip
OpenAI's first custom inference chip, Jalapeño, delivered up to 1.9 times more work per watt than comparison systems in independently run benchmarks. Image: OpenAI
Cloud & Infrastructure

OpenAI Publishes First Benchmarks for Its Jalapeño AI Chip

OpenAI’s first custom inference chip, Jalapeño, delivered up to 1.9 times more work per watt and far lower latency than rival systems, with AI accelerating its own design.

By Olivia Grant • 4 mins read Edited by Maria Konash Published: Updated:

Key Notes

  • OpenAI published first benchmark results for Jalapeño, its first custom inference chip, showing 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency than comparison systems across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, tested on SemiAnalysis's InferenceX benchmark. SemiAnalysis itself ran the tests in OpenAI's labs and reported Jalapeño beating Nvidia's Blackwell (GB200/GB300) on performance-per-watt across nearly all scenarios.
  • OpenAI plans to begin deploying Jalapeño in its own data centers by year-end, calling it the first of a multigenerational roadmap (Gen 2 in development, Gen 3 taking shape). .

OpenAI published its first detailed performance results for Jalapeño, its first custom silicon chip, built with Broadcom and designed specifically for inference, the work of running already-trained AI models rather than training new ones. The company presented the results at the Hot Chips conference and in a blog post, positioning the chip as evidence of a genuine full-stack hardware advantage rather than a marketing exercise.

Tested on InferenceX, a public benchmark from the analysis firm SemiAnalysis that measures the complete process of serving an AI request, Jalapeño was run against leading commercial inference systems across three open-weight models: GPT-OSS 120B, DeepSeek R1 670B and the trillion-parameter Kimi K2.5.

Across all three, OpenAI reported Jalapeño delivering 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems, with gains reaching 2.1 to 4.1 times for highly interactive, agent-style workloads. On Kimi, the largest model tested, the chip showed roughly 1.5 times higher performance per watt and 3.4 times lower latency.

Notably, these were not purely self-reported claims taken on faith. SemiAnalysis, which built and maintains the InferenceX benchmark, said it ran the tests itself inside OpenAI’s labs alongside OpenAI engineers, and in its own independent writeup reported that Jalapeño beat Nvidia’s Blackwell-generation GB200 and GB300 systems on performance per watt across almost all tested scenarios, without being tuned for any single point on the performance curve.

That independent confirmation matters, since it distinguishes this from a purely internal marketing claim, though the tests were still conducted on OpenAI’s own hardware and with OpenAI’s cooperation.

Jalapeño is rated at 700 watts but ran at or below 550 watts sustained on the tested workloads. Architecturally, OpenAI designed the chip to minimize data movement and communication delays between compute, memory and networking across the distinct phases of language-model inference, aiming for a balanced accelerator that performs well at both the compute-heavy prompt-processing stage and the memory-bound response-generation stage, rather than trading one off against the other.

AI Accelerating Chip’s Own Creation

OpenAI says its models helped design Jalapeño, exploring implementations and shortening design and verification loops enough to go from initial design to manufacturing tapeout in roughly nine to sixteen months, an unusually fast cycle for custom silicon.

The company also designed the chip to be a clear, predictable target for AI to program, and used its Codex coding agent to bring three open-weight models that were not originally part of Jalapeño’s production plan to strong performance within two months; for selected attention and expert-routing code blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-written versions. That points toward a compounding loop in which AI increasingly designs and optimizes the hardware it will eventually run on.

What Remains Unproven

Meaningful caveats apply before reading too much into the results. Jalapeño is not yet generally available, and OpenAI says it is still completing production qualification, maturing its software stack and preparing to operate the chip at scale. The comparison systems and specific configurations were chosen by OpenAI, first-generation chips are historically uncompetitive against established rivals, and a favorable benchmark result does not guarantee the same advantage holds across every real-world workload or model OpenAI eventually runs on it.

OpenAI has been explicit that it is not abandoning Nvidia, saying it will continue deploying accelerators widely from Nvidia and other partners for both training and inference. OpenAI plans to begin deploying Jalapeño within its own infrastructure by the end of 2026, describing it as the first generation of a multigenerational platform, with a second generation already deep in development and a third taking shape.

The near-term payoff, if the results hold at scale, is straightforward: cheaper, faster inference that could let OpenAI serve growing demand without cost scaling in lockstep, an increasingly important lever as the company works to make its enormous compute spending sustainable.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cloud & Infrastructure, News