Custom AI silicon is edging closer to the mainstream, and the latest evidence arrived at a major semiconductor conference where OpenAI revealed the first real performance numbers for its in-house inference chip. Co-developed with Broadcom and known as Jalapeño, the accelerator posted efficiency and latency figures that outpace Nvidia's current Blackwell-generation hardware on a widely watched benchmark. The results are a milestone for OpenAI's strategy of building its own compute stack, and they underscore how aggressively the largest AI labs are moving to reduce their dependence on a single supplier.
What Jalapeño is and how it was built
Jalapeño is OpenAI's first custom-designed chip, an application-specific integrated circuit, or ASIC, built specifically to run large language model inference rather than general-purpose workloads. The company unveiled the project with Broadcom on June 24, 2026, describing it as an "Intelligence Processor" purpose-built for serving LLMs at scale, and reporting that the design was completed in an unusually short period of roughly nine months. Broadcom contributes chip design and its Tomahawk networking fabric for the systems around the processor, while OpenAI brings the model-side optimization and the serving software that runs on top of the silicon.
The design philosophy sets Jalapeño apart from general-purpose accelerators such as Nvidia's H100 or GB200, which are engineered to handle both training and inference across many different workloads. Because Jalapeño does one thing, OpenAI was able to co-design compute, memory, networking, and serving software as a single integrated system, adapting each layer to the demands of real LLM traffic. OpenAI has also emphasized that AI was used to design the chip itself, summarizing the approach as using AI to design hardware and designing that hardware so AI can be programmed onto it efficiently.
The partnership structure is notable in its own right. OpenAI is not trying to manufacture or market the chip as a standalone product; instead Broadcom handles silicon implementation and networking while OpenAI supplies the model-level expertise and the full-stack software that extracts maximum value from the hardware. This split mirrors a broader strategy among cloud and AI companies, which are increasingly teaming up with established ASIC designers rather than building the specialized, expensive engineering teams required to bring a processor to life from scratch. Because Jalapeño serves OpenAI's own workloads rather than being sold to outside customers, the chip's success is measured in internal cost and latency, not external market share.
First benchmarks: efficiency and latency gains over Nvidia
The headline numbers emerged at the Hot Chips 2026 conference on August 25, 2026, when OpenAI disclosed the first batch of measured results. Using the InferenceX benchmark published by the analytics firm SemiAnalysis and running it in OpenAI's own labs, Jalapeño delivered roughly 1.5 to 1.9 times more AI work per watt at peak throughput than Nvidia's Blackwell-class systems, which include the GB200 and GB300. The same tests showed 1.7 to 3.6 times lower end-to-end latency, a meaningful advantage for interactive assistants where response speed directly shapes the user experience. The figures were measured across several large models, including OpenAI's open-weight GPT-OSS, DeepSeek R1, and the Kimi K2.5 1T model.
Power is central to the efficiency story. Jalapeño is rated at 700 watts, while the Nvidia flagship system it was compared against is specified at roughly 1,400 watts, and OpenAI reported that sustained power on the workloads tested remained at or below 550 watts. Getting comparable throughput per watt at roughly half the power budget is precisely the kind of result that matters at data-center scale, where energy is becoming one of the biggest constraints on AI expansion. OpenAI has said these early results point to an inference cost reduction of roughly 50 percent versus running the same workloads on general-purpose Nvidia hardware, a claim it bases on internal measurements.
The results come with important caveats. Jalapeño was not tested against Nvidia's newer Vera Rubin platform, which began shipping recently and also relies on HBM4 memory, so the comparison is against the previous generation rather than the latest available silicon. In addition, Jalapeño is an inference-only accelerator, which means it does not displace Nvidia for model training, the compute-heavy process of updating model weights. OpenAI continues to rely on Nvidia GPUs such as its H100 and B200 clusters for training work, a relationship the company has said it is not walking away from. Analysts broadly agreed the benchmark is a genuine technical signal, even while noting that the fairer head-to-head test, and the one that matters most going forward, will come against Rubin.
What the chip means for the AI infrastructure race
OpenAI's decision to bring inference in-house reflects a wider shift among the largest AI labs, which are increasingly investing in custom silicon to control cost, availability, and performance at scale. Nvidia still dominates the AI chip market, having reported $96.2 billion in revenue for its second quarter of fiscal 2026, up about 106 percent year over year, but the economics of serving billions of inference requests are pushing even its biggest customers to explore alternatives. For OpenAI, owning the full stack from models through serving software and now chips creates a path to lower unit costs and faster iteration on features such as reasoning models that demand large amounts of inference compute.
Deployment will be gradual rather than immediate. OpenAI has said Jalapeño will begin deploying in very small volumes toward the end of 2026, with more significant production expansion expected in 2027. This phased ramp reflects both the cautious validation of a brand-new architecture and the practical realities of manufacturing and integration. In the meantime, the company frames Jalapeño as the first step in a multigenerational platform, signaling that first-party silicon will be a permanent and growing part of its infrastructure strategy rather than an experiment.
For the broader industry, the development is a reminder that the AI compute market is becoming more contestable. Broadcom is emerging as a major alternative supplier to Nvidia by partnering with leading labs on custom accelerators, and the performance data from OpenAI gives these challengers a concrete, if narrow, data point to point to. The caveats around Vera Rubin and training workloads, however, mean Nvidia's role is far from diminished, and the real contest is only beginning. The next wave of benchmarks, and the race between Rubin and the second generation of first-party chips, will determine just how much of the fast-growing inference market the challengers can capture.
Key Takeaways
- First custom chip: OpenAI and Broadcom unveiled Jalapeño, an inference-only ASIC purpose-built for large language model serving, with the design completed in about nine months and computing, memory, and networking co-designed as one system.
- Benchmark results: On the InferenceX benchmark, Jalapeño delivered 1.5–1.9x more AI work per watt at peak throughput and 1.7–3.6x lower end-to-end latency than Nvidia's Blackwell-class GB200/GB300 systems, tested on models including GPT-OSS, DeepSeek R1, and Kimi K2.5 1T.
- Power advantage: Jalapeño is rated at 700W versus roughly 1,400W for the comparable Nvidia flagship, with measured sustained power at or below 550W on tested workloads and a claimed inference cost reduction of about 50 percent.
- Key caveats: The chip was not tested against Nvidia's newer Vera Rubin platform, and as an inference-only accelerator it does not displace Nvidia for model training, where OpenAI still relies on its H100 and B200 clusters.
- Gradual ramp: OpenAI plans to deploy Jalapeño in very small volumes by late 2026, with more significant production expansion in 2027, positioning it as the first step in a multigenerational first-party silicon roadmap.