Nvidia presents Groq 3 LPX rack, cites 3,431 tokens/s third-party benchmark
This digest was compiled by AI from multiple sources — links to the originals are below.

Nvidia presented the Groq 3 LPX inference rack at Hot Chips 2026, publishing a third-party benchmark of 3,431 output tokens per second on a 100K-context Gemma 4 31B workload. The rack is already in production, built on the LP30 chip from Nvidia's $20 billion Groq acquisition. The figure is roughly four times the next-fastest public endpoint, though the comparison used a private pre-release endpoint under single-request conditions.
Key Facts
- Artificial Analysis measured the Groq 3 LPX rack at 3,431 output tokens per second on a 100K-context Gemma 4 31B reasoning workload.
- The next-fastest public endpoint achieved 870 tokens per second, making the LPX rack roughly four times faster.
- Nvidia's on-stage demo showed 10,996 tokens per second on the same model, which Nvidia VP Igor Arsovski flagged as self-reported.
- Each LP30 chip carries approximately 500MB of on-die SRAM and no HBM, giving a full 256-chip LPX rack 128GB of memory and 40 PB/s of aggregate bandwidth.
- The LPX rack delivers 315 PFLOPS of FP8 compute with 350 ns chip-to-chip latency in a Vera Rubin-compatible, MGX liquid-cooled design.
Benchmark Methodology
Artificial Analysis ran the comparison on a private, pre-release Gemma 4 31B endpoint served through Google Cloud, taking the median of 50 sequential client requests at a concurrency of one. The public providers measured against ran shared production serverless endpoints, so the single-request condition produces the highest per-user token rate the hardware can post. Nvidia's on-stage demo showed a higher figure of 10,996 tokens per second on the same 31B model, which Igor Arsovski, Nvidia's VP of hardware, flagged as self-reported. Arsovski told the audience the aim was third-party verified independent benchmarks that customers can trust.
LP30 Architecture
Each LP30 carries roughly 500MB of on-die SRAM and no HBM, so a full LPX rack of 256 chips holds 128GB of memory delivering 40 PB/s of aggregate bandwidth. The rack provides 315 PFLOPS of FP8 compute with 350 ns of chip-to-chip latency in a Vera Rubin-compatible, MGX liquid-cooled design that scales past 1,000 LPUs. Keeping model weights resident in SRAM rather than streaming them from HBM removes the memory-access latency that dominates single-token decode. The design drops caches, branch prediction, and out-of-order execution in favor of a fully deterministic pipeline that the compiler schedules at clock-cycle granularity. The architecture descends directly from the Tensor Streaming Processor that Groq, founded by ex-Google TPU engineer Jonathan Ross, described in a 2020 ISCA paper titled Think Fast.
Production Status
Igor Arsovski, now Nvidia's VP of hardware, said the rack is already in production, built on the LP30 chip Nvidia obtained through its $20 billion Groq deal in December 2025. The same deal pushed the Rubin CPX it replaced off Nvidia's roadmap. Gemma 4 31B is a dense model small enough to sit inside a single LPX rack, and the picture at trillion-parameter mixture-of-experts scale, where memory capacity becomes the main constraint, went unaddressed.
6 sources
Nvidia presents Groq 3 LPX rack, cites 3,431 tokens/s third-party benchmark



