AI-Economy

OpenAI surpasses Nvidia's Blackwell and Rubin with Jalapeño chip

3 min read

TL;DR Too Long; Didn’t read

OpenAI's new inference chip Jalapeño achieves up to 104 times more token output per kilowatt in tests by SemiAnalysis than comparable Nvidia systems. The values presented at the Hot Chips conference originally come from OpenAI itself, so they are not fully independently verified. The chip, developed in collaboration with Broadcom, is expected to run in small quantities starting in late 2026, with larger quantities following in 2027.

A microchip in the shape of a jalapeño pepper lies in front of a server rack with the OpenAI logo, in the background a receding chip with the Nvidia logo. Image generated with GPT Image 2

Key takeaways

  • OpenAI showcased the chip in detail for the first time at the Hot Chips conference in Stanford.
  • Broadcom manufactures the chip – development took about 16 months according to SemiAnalysis.
  • On the DeepSeek R1 model, Jalapeño achieved over 700 tokens per second per user.
  • Nvidia remains, according to OpenAI, a central partner for training and part of the inference.
  • OpenAI's hardware chief Richard Ho speaks of a significant advancement over the previous state of the art.
  • Lack of tests with the practical benchmark AgentX leaves conclusions about real production load open.

OpenAI has presented detailed benchmarks for its inference chip Jalapeño for the first time at the Hot Chips conference in Stanford. The chip, developed in collaboration with Broadcom, delivers more computing power per watt than Nvidia’s current systems Blackwell and Rubin, confirms industry analyst SemiAnalysis. This could make ChatGPT faster in the medium term and make OpenAI less dependent on Nvidia’s scarce hardware.

Broadcom manufactures the custom chip after 16 months of development

OpenAI and Broadcom had publicly announced the partnership in June 2026, but details about performance remained vague until now. Jalapeño is a so-called ASIC – an application-specific chip that, unlike an Nvidia graphics card, is built exclusively for running pre-trained language models (inference), making it cheaper and more energy-efficient. Development began in mid-2024, and the finished design went into production in November 2025 – a cycle of about 16 months.

OpenAI President Greg Brockman stated to CNBC that their own AI models accelerated programming and optimization during the nine-month design phase. This step is part of a series of announcements by major AI providers aiming to make their data centers less dependent on individual suppliers. Anthropic is also building its own chip design team for this purpose, and the Chinese lab DeepSeek is developing its own inference chip. According to OpenAI, it continues to operate Nvidia and AMD chips in parallel and does not fully replace them with Jalapeño.

Benchmark values significantly exceed Nvidia’s current systems

OpenAI hardware chief Richard Ho described the results to TechCrunch as “a very, very significant performance leap over the current state of the art.” In tests conducted on-site by SemiAnalysis, Jalapeño achieved more than 700 tokens per second and user with a single concurrent request on the open model DeepSeek R1 with 670 billion parameters. The value was similarly high on the Kimi-K2.5 model, while the next-best tested chip managed only about 100 tokens per second. How hard such speed figures are to compare shows in Nvidia’s own inference chip: for the Groq 3 LPX, the company reports 3,400 tokens per second – spread across at least 64 chips.

On the comparison metric of token yield per kilowatt, Jalapeño achieved 54 to 104 times that of comparable systems at the same response speed, depending on the tested model, according to SemiAnalysis. The raw values come from OpenAI itself and are therefore not independently verified, even though SemiAnalysis engineers accompanied the tests in the lab.

A fairer comparison is Nvidia’s upcoming Rubin generation, which has not yet shipped: the two systems are roughly equal on cost per generated token, but Jalapeño achieves this without the acceleration technique of speculative decoding, which Rubin already uses. First chip generations are usually not competitive, commented SemiAnalysis founder Dylan Patel, but OpenAI is already beating Blackwell and even Rubin here.

Software maturity seen as the next test

The published values come exclusively from a relatively simple test scenario with 8,000 input and 1,000 output tokens. No numbers are yet available for the more meaningful production benchmark AgentX, which models the multi-step processes of real production systems. The tested models also do not belong to the largest available variants from their respective makers.

The current results also come from the first chip stepping. An improved successor with roughly 25 percent higher power efficiency is already in fabrication. Jalapeño also still lacks the competition’s acceleration technology, which analysts say leaves room for further gains.

SemiAnalysis views the pace of software development positively: according to Patel, the Jalapeño team stood up a complex parallelization technique within eight days – a fraction of the time comparable changes to established Nvidia software usually take. If the chip succeeds, the analyst argues, it would signal that Nvidia’s CUDA software ecosystem could lose ground as a competitive moat.

What matters next is whether the efficiency gains hold up in a second chip generation with speculative decoding, and whether future independent tests on the stricter AgentX benchmark show similar results. The next real test comes at the end of 2026, when the first Jalapeño systems are due to answer actual customer traffic in small volumes.

Frequently asked questions

When will the Jalapeño chip be operational?

First systems are expected to start running in small quantities by the end of 2026, with broader deployment planned for 2027.

Can companies rent or buy the chip themselves?

No, so far Jalapeño is exclusively intended for OpenAI's own inference infrastructure behind ChatGPT and the API; a standalone cloud offering is not known.

What distinguishes an ASIC like Jalapeño from an Nvidia GPU?

An ASIC is tailored to a specific task, in this case executing pre-trained models, and is therefore cheaper and more efficient than a flexibly deployable graphics card.

Does the chip mean the end of OpenAI's dependence on Nvidia?

No, Nvidia remains, according to the company, a central hardware partner for training and part of the inference; Jalapeño complements the capacity.

How reliable are the published benchmark figures?

The raw data comes from OpenAI itself, and SemiAnalysis accompanied the tests in the lab without fully independently re-measuring them.

Sources (5)
  1. OpenAI and Broadcom unveil LLM-optimized inference chip
  2. OpenAI and Broadcom reveal Jalapeno, first AI chip in partnership
  3. OpenAI Jalapeño: Better Than Nvidia Blackwell
  4. OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show
  5. OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog