AI-Economy

Nvidia launches Groq chip – speed record under criticism

4 min read

TL;DR Too Long; Didn’t read

Nvidia has transitioned the inference chip Groq 3 LPX into mass production and reports 3,400 tokens per second for an open 31 billion parameter model. The chip comes from the 20 billion dollar acquisition of Groq and is intended to allow AI agents to respond faster, with Nebius as the first customer. Experts consider the fourfold speed advantage over Cerebras to be misleading, as Nvidia requires 64 chips instead of one to two.

A microchip with a Groq logo sticker sits on an Nvidia logo plaque, while a hare with a lightning-bolt symbol races past a slower tortoise wearing a Cerebras sticker Image generated with GPT Image 2

Key takeaways

  • Groq 3 LPX has been in mass production since August 24, 2026, manufactured by Samsung.
  • The chip achieves 3,400 tokens per second at a 100,000-token context window in the agent benchmark.
  • Nvidia acquired Groq's technology in December 2025 for 20 billion dollars as a licensing and personnel deal.
  • Cloud provider Nebius integrates the chip as the first customer into its Token Factory platform.
  • For the comparison with Cerebras, Nvidia needs at least 64 chips, while Cerebras only needs one to two.
  • Senator Warren questioned the Groq deal in March 2026 as a possible circumvention of antitrust laws.

Nvidia has transitioned the specialized inference chip Groq 3 LPX into mass production and positions it as a response to the growing demands of AI agents. In the benchmark, the chip achieves 3,400 tokens per second with an open 31-billion-parameter model with a 100,000-token context window – according to Nvidia, four times faster than the next competing platform.

Series chip comes from controversial licensing deal

The chip is based on technology from a deal that Nvidia, according to its own press release, transitioned into mass production on August 24, 2026. The basis is an agreement closed in December 2025 worth $20 billion – Nvidia’s largest acquisition to date. Legally, it is not a classic takeover: Nvidia acquires a non-exclusive license for Groq’s chip technology and hires a large part of the management team, including the former CEO, while Groq remains formally independent. Company head Jensen Huang stated that the Vera Rubin portfolio is thus expanding with a platform specifically tailored to agent-based AI workloads.

The Groq 3 LPX complements the Vera Rubin NVL72 platform with a dedicated component for the so-called decode phase of inference – the section that determines how quickly a system generates individual response tokens. For this, the chip manufactured by Samsung incorporates 500 megabytes of cache directly on the die to relieve the memory bandwidth, which often becomes a bottleneck in other inference accelerators. A price for the chip itself has not been disclosed; it is not sold individually but provided exclusively through cloud providers as computing infrastructure.

Nebius becomes first customer, SpaceX relies on Vera Rubin

The first customer is the AI cloud provider Nebius, in which Nvidia already holds a 9.3 percent stake. Nebius integrates the Groq 3 LPX into its production platform Token Factory. Chief Technology Officer Danila Shtan justified the move by stating that the generation phase of inference determines the perceived response speed of an AI system – exactly where the new chip comes into play. Nebius is the first AI cloud to productively use the component, enabling faster response times for multi-step agent applications.

According to a report by SiliconANGLE, another major customer for the Vera Rubin platform is the aerospace company SpaceX. It intends to build its next AI architecture on Nvidia’s Vera CPUs, which will handle orchestration, tool execution, and simulation tasks both in terrestrial data centers and on satellites. The combination of Nebius as a cloud customer and SpaceX as an infrastructure partner demonstrates how Nvidia is leveraging the Groq acquisition across various application fields – from publicly accessible AI cloud infrastructure to internal enterprise agent infrastructure. A start date for customers in Germany or the EU is not yet known; access is only available indirectly through international cloud platforms.

Experts consider comparison with Cerebras to be embellished

Nvidia promotes the Groq 3 LPX with a speed advantage that, according to an analysis by The Register, is only partially meaningful. In the benchmark, the chip achieved 3,400 tokens per second compared to 882 tokens per second for a system from Cerebras – according to Nvidia’s calculations, a fourfold lead. However, the analysis points out that for the Nvidia value, at least 64 Groq 3 LPX chips are needed, while the same model size fits into one to two accelerators at Cerebras. Moreover, the tested model with 31 billion parameters represents a particularly favorable case for Nvidia’s architecture, not a typical production deployment. Cerebras has also already announced successor systems with double the performance, so the comparison may quickly become outdated.

The acquisition is also under political scrutiny: U.S. Senators Elizabeth Warren and Richard Blumenthal questioned Nvidia in an open letter dated March 20, 2026, whether the deal is a so-called reverse-acquihire construction aimed at circumventing antitrust pre-review under the Hart-Scott-Rodino Act. Nvidia’s roughly 90 percent market share in AI accelerators exacerbates these concerns, according to the senators; a final decision from the Department of Justice or the FTC is still pending.

It will be crucial whether the performance advantage holds up in production environments with larger models and mixed workloads, rather than just in the particularly favorably chosen test scenario. It also remains open how U.S. antitrust authorities will respond to the reverse-acquihire construction – a decision likely to have signaling effects for similar deals by other tech companies beyond the Groq case.

Frequently asked questions

What is the price of the Groq 3 LPX?

Nvidia has not published a price. The chip is not sold individually but provided through AI cloud providers like Nebius as computing infrastructure.

Is the chip available in Germany or the EU?

Not directly. Access is only available indirectly through international cloud platforms that integrate the chip into their infrastructure; a European launch date is not yet known.

What distinguishes the Groq 3 LPX from regular Nvidia GPUs?

It is exclusively specialized for the token generation phase of inference and uses large cache directly on the chip, instead of covering training and inference together like GPUs.

What is Groq's stance on the deal?

According to the agreement, Groq remains formally an independent company and described the agreement as a non-exclusive licensing agreement, while executives transitioned to Nvidia.

When will US authorities review the deal?

Senators urged the Department of Justice and the FTC to respond by April 3, 2026; a public decision has not yet been made.

Sources (5)
  1. NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
  2. What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble
  3. Nvidia buying AI chip startup Groq's assets for about $20 billion in its largest deal on record
  4. Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
  5. Warren, Blumenthal Question Whether NVIDIA's $20 Billion Groq Deal is Attempt to Avoid Antitrust Laws

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog