
Prefer to listen instead? Here’s the podcast version of this article.
Artificial intelligence may look like software, but the latest AI race is getting increasingly physical.
At Hot Chips 2026, OpenAI pulled back the curtain on Jalapeño, its first custom AI inference accelerator, offering a much deeper look at a chip that could become an important part of the infrastructure behind ChatGPT, Codex, AI agents, and future OpenAI services.
The OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026 story matters for more than semiconductor enthusiasts. It is a window into where the economics of artificial intelligence are heading: away from a world where every major AI workload automatically lands on a general-purpose GPU and toward one where hyperscalers and AI labs increasingly design specialized silicon around their own models, workloads, networking requirements, latency targets, and power constraints.
And OpenAI is not exactly dipping a toe in the water.
Jalapeño is intended to become the first generation of a multi-generation OpenAI compute platform.
Jalapeño is an application-specific integrated circuit, or ASIC, designed primarily for large-language-model inference.
That distinction is important.
Training is the computationally enormous process used to build or update a model. Inference is what happens after training, when millions of users ask that model questions, generate code, run AI agents, summarize documents, create content, or interact with AI-powered applications.
OpenAI formally introduced Jalapeño on June 24, 2026, through its announcement, [OpenAI and Broadcom unveil LLM-optimiezed inference chip] OpenAI said the accelerator was designed from scratch around modern LLM inference rather than adapting an architecture originally intended for a broader range of computing jobs. Broadcom is contributing silicon implementation and networking technology, while Celestica is involved in board, rack, and system integration.
That makes Jalapeño less of an attempt to create “another GPU” and more of an attempt to answer a very OpenAI-specific question:
What would the ideal computer look like if you could design the hardware, software, model-serving system, networking, memory architecture, and AI products together?
For more context on why specialized accelerators are becoming so important, Quantilus recently explored the same industry shift in [A New Era of AI Hardware Built for Training and Inference]
Hot Chips 2026 ran from August 23–25 at Stanford University, with OpenAI presenting during the conference’s final AI session on August 25. Richard Ho, Ravi Narayanaswami, and Chris Leary delivered a wonderfully understated presentation titled “You Can Just Build Things … Chips.” The official Hot Chips program confirms OpenAI shared the AI 2 session with Google and SambaNova.
What followed was considerably more technical than the original June announcement.
Detailed coverage from [ServeTheHome’s OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026 analysis] describes a full inference platform rather than merely a silicon die: accelerator hardware, memory, networking, racks, software, kernels, and a programming model designed to operate together.
That full-stack approach may ultimately prove more important than any individual specification.
According to the architecture details presented at Hot Chips, a Jalapeño accelerator delivers approximately 13.4 PFLOP/s of MXFP4 matrix compute, paired with 216 GiB of HBM4 memory and roughly 15.4 TB/s of HBM bandwidth.
The package has a rated thermal design power of 700 watts, although OpenAI says sustained power during the workloads it tested remained at or below approximately 550 watts.
That is particularly interesting because AI infrastructure is increasingly power-constrained.
More GPUs do not magically create more electricity. Data-center operators face fixed utility connections, cooling capacity, power-distribution infrastructure, construction timelines, transformers, backup systems, and local grid limitations. As a result, how much useful AI work can be produced from every kilowatt is becoming nearly as important as raw chip performance.
Now for the spicy part. Yes, apparently the chip name has made pepper jokes legally mandatory.
OpenAI benchmarked Jalapeño using InferenceX, a public inference benchmarking framework developed by SemiAnalysis.
The tests included OpenAI’s GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. According to OpenAI’s reported results, Jalapeño reached the performance-per-watt and latency Pareto frontier across those workloads.
On GPT-OSS 120B, OpenAI reported roughly 1.9 times higher peak throughput per kilowatt and approximately 1.7 times lower end-to-end latency at matched operating points against the comparison system.
For DeepSeek R1, reported improvements reached roughly 1.7 times higher peak performance per watt and 3.6 times lower end-to-end latency in one comparison.
On the enormous Kimi K2.5 workload, OpenAI reported approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency.
Independent technical analysis is available in [SemiAnalysis and InferenceX’s deep dive into OpenAI Jalapeno] SemiAnalysis says it benchmarked the system with OpenAI engineers and describes Jalapeño as an unusually competitive first-generation custom accelerator.
There is, however, an important caveat.
SemiAnalysis argues that comparing Jalapeño primarily with Nvidia’s Blackwell-generation GB200 and GB300 does not tell the complete competitive story because newer Nvidia Rubin-generation hardware is the more forward-looking comparison.
That nuance matters. Benchmark headlines saying “OpenAI beats Nvidia” are irresistible click candy, but AI accelerator performance varies dramatically by model, batching strategy, precision, software stack, networking configuration, latency target, and system scale.
The sensible conclusion is not that Nvidia has suddenly been dethroned.
It is that OpenAI has produced a first-generation ASIC credible enough to enter the conversation.
That alone is significant.
Perhaps the most fascinating part of the Jalapeño story is how quickly OpenAI says it was developed.
OpenAI says the core project moved from initial design to manufacturing tape-out in roughly nine months, with OpenAI models helping engineers during design and optimization.
There is some useful nuance around that figure: SemiAnalysis describes a longer period of roughly 16 months when counting from early team formation to tapeout, while OpenAI’s nine-month claim refers more specifically to the design-to-tapeout phase.
At Hot Chips, OpenAI also described using AI to optimize hardware kernels. ServeTheHome reported that AI-optimized attention and mixture-of-experts kernels achieved roughly 1.5x to 1.8x higher performance than existing expert-written implementations in the examples shown.
This creates a fascinating feedback loop.
OpenAI builds AI models.
Those AI models help engineers design and program better chips.
Those chips run future AI models more efficiently.
The improved infrastructure makes larger and more capable AI services economically viable.
Those services can then assist with the next generation of infrastructure.
Software eating the world was apparently just the appetizer.
No.
OpenAI has been unusually explicit on this point.
The company says it expects to continue deploying accelerators from Nvidia and other partners for both training and inference, even as Jalapeño enters its infrastructure fleet.
That makes sense.
AI compute demand is expanding so rapidly that OpenAI’s problem is unlikely to be choosing one chip supplier and abandoning everyone else. The larger challenge is securing enough economically viable compute from multiple sources.
Jalapeño therefore looks less like an Nvidia replacement and more like a strategic new layer in OpenAI’s supply portfolio.
It gives OpenAI another source of compute, greater control over inference economics, tighter hardware-software optimization, and potentially more negotiating leverage throughout the semiconductor ecosystem.
Inference is where an AI business repeatedly pays the computational bill.
A frontier model may require an enormous one-time training investment, but serving that model to hundreds of millions of requests creates an ongoing infrastructure cost.
Every reduction in energy per token, latency per request, server count, or networking overhead can therefore multiply across massive volumes.
For OpenAI, successful custom silicon could eventually mean lower operating costs, greater capacity, more predictable infrastructure supply, and improved performance for ChatGPT, Codex, API users, and AI agents.
For enterprises, those infrastructure gains could eventually surface as faster products, higher rate limits, more capable agentic workflows, or lower costs.
Jalapeño also highlights an increasingly important environmental and infrastructure issue.
AI’s energy requirements are growing rapidly, and large data-center deployments can create pressure on electricity grids, water systems, land use, and local communities.
That means performance per watt is not merely an engineering benchmark or cost-saving trick.
It is becoming part of responsible AI infrastructure design.
If a system can perform substantially more useful AI inference within the same power envelope, operators may be able to serve more users without requiring an equivalent increase in electricity consumption.
Of course, efficiency has a catch: cheaper and faster compute can encourage dramatically more AI usage. Better efficiency per request does not automatically guarantee lower total energy consumption.
This is why sustainability reporting should increasingly include both efficiency metrics and absolute consumption.
Policymakers, enterprises, and infrastructure providers will also need better transparency around electricity sources, grid impact, water consumption, carbon intensity, semiconductor supply chains, and lifecycle emissions.
The hardware may be getting smarter. Governance still needs to keep up.
The bigger competitive story extends far beyond OpenAI and Nvidia.
Google has spent years building TPUs. Amazon has developed Trainium and Inferentia. Microsoft has its Maia accelerator program. Meta is developing custom AI silicon. Meanwhile, Nvidia, AMD, Cerebras, SambaNova, and other semiconductor companies continue pushing specialized architectures of their own.
Hot Chips 2026 effectively showcased this fragmentation in real time, with presentations from Meta, Microsoft, Google, Nvidia, Cerebras, SambaNova, AMD and OpenAI across the program.
The future of AI computing therefore probably will not be one universal processor.
It is increasingly likely to be a heterogeneous world where organizations select hardware according to the workload: training, high-throughput inference, interactive inference, long-context reasoning, agentic workloads, multimodal applications, or specialized enterprise deployment.
That competition should ultimately benefit AI users—provided increasingly powerful infrastructure is deployed responsibly and access does not become concentrated among only a handful of companies.
OpenAI’s Jalapeño custom AI ASIC is more than a new chip—it is a sign that the AI industry is entering a new phase where software, models, networking, memory, power efficiency, and silicon are being designed together as one integrated system.
At Hot Chips 2026, OpenAI showed that it is no longer relying solely on third-party accelerators to shape the performance and economics of its AI services. By building specialized hardware around real inference workloads, OpenAI is aiming to reduce latency, improve performance per watt, increase infrastructure efficiency, and support the growing demands of AI agents and large-scale model serving.
That does not mean Nvidia or other accelerator vendors are suddenly irrelevant. Far from it. The more likely future is a diverse AI infrastructure ecosystem in which custom ASICs, GPUs, TPUs, and other specialized processors coexist, each optimized for different workloads and deployment needs.
What makes Jalapeño especially important is the precedent it sets. If OpenAI can successfully combine custom silicon with its models, software stack, networking architecture, and data-center strategy, it could gain far greater control over the cost and performance of running AI at scale. That could influence everything from API pricing and response times to the capabilities of future autonomous agents.
The competitive race will also extend beyond raw speed. Energy efficiency, sustainability, supply-chain resilience, responsible infrastructure expansion, and transparency around resource consumption are becoming critical measures of success.
Ultimately, the OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026 story is not simply about whether one processor can outperform another. It is about who can build the most efficient, scalable, and intelligent computing platform for the next generation of AI.
And if Jalapeño is only OpenAI’s first step into custom silicon, the real impact may become clearer with the generations that follow.
WEBINAR