mimile
Back to feed

Nvidia's Groq 3 LPX AI inference chip enters full production

AI digest

This digest was compiled by AI from multiple sources — links to the originals are below.

Nvidia's Groq 3 LPX AI inference chip enters full production

Nvidia Corp. announced at Hot Chips 2026 that its Groq 3 LPX AI inference accelerator has entered full production. The chip is designed to boost token generation speeds for the Vera Rubin platform, with Nebius Group N.V. as its first customer. The announcement comes as Nvidia seeks to maintain its lead in AI compute.

Key Facts

  • Nvidia's Groq 3 LPX AI inference accelerator is now in full production, the company announced at Hot Chips 2026.
  • Nebius Group N.V. is the first customer to commit to using the Groq 3 LPX chip.
  • In Artificial Analysis benchmarks, Groq 3 LPX achieved a record 3,400 tokens per second running the Gemma 4 31B model with a 100,000-token context window.
  • A full rack-scale deployment can harness up to 256 LP30 accelerators linked by high-bandwidth interconnects.
  • The chip was built using technology licensed from Groq Inc.

Production and Platform Integration

Nvidia announced full production of the Groq 3 LPX at Hot Chips 2026, following mass production of Vera CPUs and Vera Rubin servers. The Groq 3 LPX is a purpose-built extension to Nvidia's flagship Vera Rubin data center platform, designed to offload latency-sensitive decode workloads from GPUs. A full rack-scale deployment can harness up to 256 LP30 accelerators linked by Nvidia's ultra-high-bandwidth chip interconnects. The chip was built using technology licensed from Groq Inc., a smaller chipmaker.

Performance and Customer Adoption

In Artificial Analysis benchmarks, Groq 3 LPX output a record-breaking 3,400 tokens per second running the open-source Gemma 4 31B agentic model with a 100,000-token context window. Nvidia claims this makes the chip four times more responsive for latency-sensitive workloads compared to rival platforms, enabling multistep agentic tasks to complete in minutes instead of hours. Neocloud provider Nebius Group N.V. has signed on as the first customer to commit to using the new chip.

3 sources

Nvidia's Groq 3 LPX AI inference chip enters full production