Coding Horizon

NVIDIA Groq 3 LPX Makes AI Agents Feel Alive Now

Every figure, name and comparison this video puts on screen, chased to a primary source. Numbers that could not be sourced are not rendered.

The acquisition

NVIDIA agreed to acquire Groq’s intellectual property and key engineering staff for about 20 billion dollars, in a definitive agreement dated 24 December 2025. It is the largest transaction in the company’s history, ahead of the roughly 7 billion dollar Mellanox purchase in 2019.

The structure matters and the video’s phrasing simplifies it: NVIDIA licensed Groq’s LPU intellectual property and hired members of its engineering team, including founder Jonathan Ross and president Sunny Madra. It did not buy Groq as a company, and Groq continues to operate as an inference cloud.

What Groq 3 LPX is

NVIDIA’s own description: “the interactive AI inference accelerator for NVIDIA Vera Rubin”, a rack mount accelerator that extends the Vera Rubin platform. Rubin GPUs handle large scale context processing while LPX accelerates the latency sensitive decode stage. Announced in full production on 24 August 2026.

The benchmark number

“Median speed across samples with 100K input context length was 3,431 tokens/second”, measured on Gemma 4 31B in Artificial Analysis benchmarking. The comparison baseline NVIDIA gives in the same post is 870 output tokens per second for the fastest public endpoint.

Same post, other figures:

The newsroom release rounds the headline figure to “record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B”.

Hardware specifications

Per LPU accelerator:

Per rack:

The technical blog describes the same rack as “256 LP30 local processing units (LPUs)” with “128 GB of total SRAM-based memory collectively in those chips” and “96 C2C links per chip running at 112 Gbps each”, scheduling computation and communication overlap on 320 byte vectors.

The power and revenue claims

NVIDIA states “35x higher throughput per megawatt (MW) for trillion-parameter models” and “10x more revenue per watt” for Vera Rubin paired with LPX. The video says “10 times more revenue opportunity”, which is a looser wording of the same claim.

NVIDIA also claims “4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform”.

All four are vendor figures, published by NVIDIA about its own product, and the product page marks its performance table “Projected performance subject to change”. They are presented on screen as NVIDIA’s claims rather than as independent results.

Early adopters

Nebius is named as the first AI cloud to adopt Groq 3 LPX. Groq itself plans to be among the earliest adopters of the platform built on the IP it sold.

Caveats