NVIDIA (NASDAQ: NVDA) announced on Monday that its Groq 3 LPX rack-scale system has entered full production, marking the official commercialization of its low-latency AI inference technology following the company's roughly $20 billion acquisition of Groq-related assets last year. The first systems will be deployed at AI cloud computing provider NEBIUS (NASDAQ: NBIS), with plans to go live later this year.
NVIDIA Senior Director Dion Harris told media that Groq 3 LPX will be deployed alongside NVIDIA's Vera central processing units and Rubin graphics processing units at Nebius data centers. The rapid progression of Groq products to full production reflects how low-latency AI inference is becoming a key battleground for NVIDIA as AI applications shift from model training toward actual deployment. This is particularly evident in AI agents and AI coding applications, where models need to generate content continuously at faster speeds to reduce user wait times.
NVIDIA acquired AI chip startup Groq's assets for approximately $20 billion in December last year, marking the largest acquisition in the company's history. A distinctive feature of the Groq chip architecture is the integration of 500MB of high-speed SRAM directly on the chip, which reduces data transfer bottlenecks caused by traditional memory access and improves response speed during AI model inference. Unlike NVIDIA's primary GPUs, which are manufactured by TSMC (NYSE: TSM), Groq chips are produced by Samsung.
NVIDIA currently integrates 256 Groq 3 chips into a single LPX rack. According to benchmark data from Artificial Analysis cited by NVIDIA, the Groq 3 LPX system can achieve a performance of approximately 3,400 tokens generated per second. For generative AI, tokens represent the basic units that models process and generate text. Higher token generation speed per second means AI can respond to user requests faster, which is especially important for latency-sensitive applications like AI coding and real-time agents.
Harris noted that for cloud computing companies offering AI inference services, lower latency enables them to charge premium prices to customers who demand faster response times. However, Groq chips are not designed to replace NVIDIA's traditional GPUs. GPUs remain the core of current AI computing infrastructure, capable of both AI model training and inference tasks, while offering greater versatility to accommodate different models and technical architectures. Groq's low-latency chips focus more specifically on particular stages of AI inference, especially the "decoding" phase when models generate content.
Harris stated: "This is not about replacing GPUs, but rather using the most appropriate processor in terms of price and performance for different parts of the workload." This indicates NVIDIA is attempting to establish a more segmented AI computing architecture: Vera CPU and Rubin GPU continue to handle broader AI computing tasks, while Groq chips are optimized for latency-sensitive inference workloads.
As the AI industry transitions from large-scale model training to a phase of rapidly growing inference demand, competition in the low-latency inference market is intensifying. AMD (NASDAQ: AMD) announced earlier this year that it would integrate its rack-scale AI systems with Cerebras (NYSE: CBRS) chips, also targeting low-latency AI inference. The importance of this sector is growing further with the increasing adoption of AI coding applications. Reports indicate that OpenAI's newly announced Ultrafast mode promises 750 tokens per second, with underlying computing power provided by Cerebras.
In comparison, benchmark data cited by NVIDIA shows Groq 3 LPX can achieve 3,400 tokens per second. However, different systems may have varying test conditions and application scenarios, so these figures cannot be simply viewed as direct performance comparisons. Meanwhile, NVIDIA is accelerating shipments of its Vera Rubin system, which entered production earlier this year. NVIDIA CEO Jensen Huang estimated in March, when unveiling Vera Rubin and Groq 3 LPX, that cumulative sales from the current Blackwell chips through the new-generation Vera Rubin system could reach $1 trillion by 2027.
Huang also revealed that for data center space dedicated to AI coding applications, he plans to allocate approximately one-quarter of it to Groq chips, with the remainder deployed with Vera Rubin systems. This allocation ratio further demonstrates NVIDIA's emphasis on the low-latency inference market. With the rapid development of AI agents, AI coding, and real-time generative AI applications, inference speed is becoming a key competitive metric for cloud computing companies and AI developers.
The full production launch of Groq 3 LPX also signals that NVIDIA is expanding from a GPU-centric AI chip supplier to offering specialized computing architectures for different AI workloads. The market will next focus on NVIDIA's latest earnings results. The company is scheduled to report financials this Wednesday, and beyond Blackwell and Vera Rubin demand, AI inference business and Groq commercialization progress may become new focal points for investors.