Nvidia has transitioned the specialized inference chip Groq 3 LPX into mass production and positions it as a response to the growing demands of AI agents. In the benchmark, the chip achieves 3,400 tokens per second with an open 31-billion-parameter model with a 100,000-token context window – according to Nvidia, four times faster than the next competing platform.
Series chip comes from controversial licensing deal
The chip is based on technology from a deal that Nvidia, according to its own press release, transitioned into mass production on August 24, 2026. The basis is an agreement closed in December 2025 worth $20 billion – Nvidia’s largest acquisition to date. Legally, it is not a classic takeover: Nvidia acquires a non-exclusive license for Groq’s chip technology and hires a large part of the management team, including the former CEO, while Groq remains formally independent. Company head Jensen Huang stated that the Vera Rubin portfolio is thus expanding with a platform specifically tailored to agent-based AI workloads.
The Groq 3 LPX complements the Vera Rubin NVL72 platform with a dedicated component for the so-called decode phase of inference – the section that determines how quickly a system generates individual response tokens. For this, the chip manufactured by Samsung incorporates 500 megabytes of cache directly on the die to relieve the memory bandwidth, which often becomes a bottleneck in other inference accelerators. A price for the chip itself has not been disclosed; it is not sold individually but provided exclusively through cloud providers as computing infrastructure.
Nebius becomes first customer, SpaceX relies on Vera Rubin
The first customer is the AI cloud provider Nebius, in which Nvidia already holds a 9.3 percent stake. Nebius integrates the Groq 3 LPX into its production platform Token Factory. Chief Technology Officer Danila Shtan justified the move by stating that the generation phase of inference determines the perceived response speed of an AI system – exactly where the new chip comes into play. Nebius is the first AI cloud to productively use the component, enabling faster response times for multi-step agent applications.
According to a report by SiliconANGLE, another major customer for the Vera Rubin platform is the aerospace company SpaceX. It intends to build its next AI architecture on Nvidia’s Vera CPUs, which will handle orchestration, tool execution, and simulation tasks both in terrestrial data centers and on satellites. The combination of Nebius as a cloud customer and SpaceX as an infrastructure partner demonstrates how Nvidia is leveraging the Groq acquisition across various application fields – from publicly accessible AI cloud infrastructure to internal enterprise agent infrastructure. A start date for customers in Germany or the EU is not yet known; access is only available indirectly through international cloud platforms.
Experts consider comparison with Cerebras to be embellished
Nvidia promotes the Groq 3 LPX with a speed advantage that, according to an analysis by The Register, is only partially meaningful. In the benchmark, the chip achieved 3,400 tokens per second compared to 882 tokens per second for a system from Cerebras – according to Nvidia’s calculations, a fourfold lead. However, the analysis points out that for the Nvidia value, at least 64 Groq 3 LPX chips are needed, while the same model size fits into one to two accelerators at Cerebras. Moreover, the tested model with 31 billion parameters represents a particularly favorable case for Nvidia’s architecture, not a typical production deployment. Cerebras has also already announced successor systems with double the performance, so the comparison may quickly become outdated.
The acquisition is also under political scrutiny: U.S. Senators Elizabeth Warren and Richard Blumenthal questioned Nvidia in an open letter dated March 20, 2026, whether the deal is a so-called reverse-acquihire construction aimed at circumventing antitrust pre-review under the Hart-Scott-Rodino Act. Nvidia’s roughly 90 percent market share in AI accelerators exacerbates these concerns, according to the senators; a final decision from the Department of Justice or the FTC is still pending.
It will be crucial whether the performance advantage holds up in production environments with larger models and mixed workloads, rather than just in the particularly favorably chosen test scenario. It also remains open how U.S. antitrust authorities will respond to the reverse-acquihire construction – a decision likely to have signaling effects for similar deals by other tech companies beyond the Groq case.


