OpenAI has presented detailed benchmarks for its inference chip Jalapeño for the first time at the Hot Chips conference in Stanford. The chip, developed in collaboration with Broadcom, delivers more computing power per watt than Nvidia’s current systems Blackwell and Rubin, confirms industry analyst SemiAnalysis. This could make ChatGPT faster in the medium term and make OpenAI less dependent on Nvidia’s scarce hardware.
Broadcom manufactures the custom chip after 16 months of development
OpenAI and Broadcom had publicly announced the partnership in June 2026, but details about performance remained vague until now. Jalapeño is a so-called ASIC – an application-specific chip that, unlike an Nvidia graphics card, is built exclusively for running pre-trained language models (inference), making it cheaper and more energy-efficient. Development began in mid-2024, and the finished design went into production in November 2025 – a cycle of about 16 months.
OpenAI President Greg Brockman stated to CNBC that their own AI models accelerated programming and optimization during the nine-month design phase. This step is part of a series of announcements by major AI providers aiming to make their data centers less dependent on individual suppliers. Anthropic is also building its own chip design team for this purpose, and the Chinese lab DeepSeek is developing its own inference chip. According to OpenAI, it continues to operate Nvidia and AMD chips in parallel and does not fully replace them with Jalapeño.
Benchmark values significantly exceed Nvidia’s current systems
OpenAI hardware chief Richard Ho described the results to TechCrunch as “a very, very significant performance leap over the current state of the art.” In tests conducted on-site by SemiAnalysis, Jalapeño achieved more than 700 tokens per second and user with a single concurrent request on the open model DeepSeek R1 with 670 billion parameters. The value was similarly high on the Kimi-K2.5 model, while the next-best tested chip managed only about 100 tokens per second. How hard such speed figures are to compare shows in Nvidia’s own inference chip: for the Groq 3 LPX, the company reports 3,400 tokens per second – spread across at least 64 chips.
On the comparison metric of token yield per kilowatt, Jalapeño achieved 54 to 104 times that of comparable systems at the same response speed, depending on the tested model, according to SemiAnalysis. The raw values come from OpenAI itself and are therefore not independently verified, even though SemiAnalysis engineers accompanied the tests in the lab.
A fairer comparison is Nvidia’s upcoming Rubin generation, which has not yet shipped: the two systems are roughly equal on cost per generated token, but Jalapeño achieves this without the acceleration technique of speculative decoding, which Rubin already uses. First chip generations are usually not competitive, commented SemiAnalysis founder Dylan Patel, but OpenAI is already beating Blackwell and even Rubin here.
Software maturity seen as the next test
The published values come exclusively from a relatively simple test scenario with 8,000 input and 1,000 output tokens. No numbers are yet available for the more meaningful production benchmark AgentX, which models the multi-step processes of real production systems. The tested models also do not belong to the largest available variants from their respective makers.
The current results also come from the first chip stepping. An improved successor with roughly 25 percent higher power efficiency is already in fabrication. Jalapeño also still lacks the competition’s acceleration technology, which analysts say leaves room for further gains.
SemiAnalysis views the pace of software development positively: according to Patel, the Jalapeño team stood up a complex parallelization technique within eight days – a fraction of the time comparable changes to established Nvidia software usually take. If the chip succeeds, the analyst argues, it would signal that Nvidia’s CUDA software ecosystem could lose ground as a competitive moat.
What matters next is whether the efficiency gains hold up in a second chip generation with speculative decoding, and whether future independent tests on the stricter AgentX benchmark show similar results. The next real test comes at the end of 2026, when the first Jalapeño systems are due to answer actual customer traffic in small volumes.


