Highlights

  • NVIDIA Groq 3 LPX interactive AI inference accelerator enters full production
  • Benchmarked at 3,400 output tokens per second on Gemma 4 31B at 100K token context
  • 4x faster than the nearest alternative platform for agentic AI inference
  • Nebius is the first cloud provider to deploy Groq 3 LPX
  • Part of the NVIDIA Vera Rubin NVL72 platform designed for ultrafast token generation

NVIDIA has announced at Hot Chips 2026 that its Groq 3 LPX interactive AI inference accelerator is now in full production. This marks a significant milestone in the company's efforts to dominate the AI inference market, which is becoming increasingly important as AI agents and real-time applications demand faster response times.

The announcement comes at a time when the AI industry is shifting focus from training large models to deploying them efficiently at scale. Inference speed has become a critical competitive advantage, and NVIDIA's latest hardware is designed to set new standards.

NVIDIA Groq 3 LPX: A New Era for AI Inference

The Groq 3 LPX is NVIDIA's latest generation inference accelerator, specifically designed for agentic AI systems that require ultrafast token generation. Unlike traditional inference chips that prioritize throughput for batch processing, the Groq 3 LPX is optimized for interactive scenarios where response latency matters most.

Key specifications include:

Breaking Records: 3,400 Tokens Per Second

The Groq 3 LPX has achieved a remarkable benchmark of 3,400 output tokens per second on Gemma 4 31B, an open-source agentic model, at a 100K token context. This is 4 times faster than the nearest alternative platform.

To put this in perspective:

These performance gains are particularly significant for agentic AI applications, where AI systems need to make multiple decisions and generate multiple responses in rapid succession.

The Vera Rubin NVL72 Platform

The Groq 3 LPX is a key component of NVIDIA's Vera Rubin NVL72 platform, which represents the company's next-generation infrastructure for AI inference. The platform is designed to support the growing demand for real-time AI applications across industries.

The Vera Rubin platform builds on NVIDIA's previous DGX and HGX architectures but is specifically optimized for inference workloads rather than training. This strategic shift reflects the industry's evolution as more companies move from AI development to AI deployment.

Nebius: First Cloud Deployment

Nebius has been announced as the first cloud provider to deploy the Groq 3 LPX. This partnership gives Nebius customers early access to the fastest inference hardware available, positioning the company as a leader in high-performance AI cloud services.

The deployment is expected to benefit:

The HBM vs DDR5 Memory Challenge

At the same Hot Chips 2026 event, Micron issued a warning about the widening silicon penalty between HBM (High Bandwidth Memory) and DDR5. As each new generation of AI chips demands more memory bandwidth, the cost differential between HBM and standard DDR5 memory continues to grow.

This has significant implications for AI hardware economics:

What This Means for AI Development

The availability of the Groq 3 LPX in full production has several implications for the AI industry:

Frequently Asked Questions (FAQ)

What is the NVIDIA Groq 3 LPX?

The Groq 3 LPX is NVIDIA's latest AI inference accelerator designed for agentic AI systems. It achieves 3,400 output tokens per second on Gemma 4 31B at 100K token context, making it 4x faster than alternatives.

What is the Vera Rubin NVL72 platform?

The Vera Rubin NVL72 is NVIDIA's next-generation AI inference platform, optimized for real-time token generation and agentic AI workloads. The Groq 3 LPX is a key component of this platform.

Why is inference speed important for AI?

Inference speed determines how quickly AI models can generate responses. Faster inference enables real-time conversational AI, agentic workflows, and interactive applications that require instant feedback.

What is the HBM vs DDR5 memory challenge?

HBM (High Bandwidth Memory) provides the bandwidth needed for AI workloads but costs significantly more than standard DDR5. The cost gap is widening with each generation, raising concerns about AI hardware economics.