500MB on-chip SRAM · 150 TB/s bandwidth · 256-LPU LPX rack · FP8 decode-only · 35× throughput vs Blackwell NVL72
Purpose-built Language Processing Unit optimized exclusively for autoregressive token generation (decode phase).
Liquid-cooled rack-scale inference system packaging 256 LPUs into a single deterministic compute fabric.
NVIDIA's recommended disaggregated serving: GPUs run prefill + attention, LPUs offload FFN layers for decode.
NVIDIA (Groq LPU
licensed technology)
Groq 3 LPU / LP30
Language Processing Unit
(inference accelerator)
500 MB per chip
150 TB/s per chip
FP8 (INT8/BF16 unverified)
256 LPUs · 128 GB SRAM
40 PB/s aggregate
No (decode-only)
~150 tokens/W
35× vs Blackwell NVL72
Flat decode latency
at any batch size
GTC 2026 · $20B Groq deal
Early access 2026
GA 2H 2026
Available
Quote
NVIDIA introduced the Groq 3 LPU (Language Processing Unit) at GTC 2026, following a $20 billion non-exclusive licensing deal with AI chip startup Groq — the largest NVIDIA has ever paid for technology and personnel. Unlike a GPU, the LPU is not a general-purpose accelerator: it does one thing — autoregressive token generation (the decode phase of LLM inference) — and it does it by replacing HBM with a massive on-chip SRAM pool.
Each Groq 3 LPU packs 500 MB of SRAM at 150 TB/s bandwidth — roughly 45× the bandwidth of an H100's HBM3 (3.35 TB/s) and 7× faster than a Rubin GPU's HBM4. Because decode is fundamentally memory-bandwidth-bound (every token reads the full weight matrix and KV cache), that SRAM wall-breaking bandwidth translates directly into tokens. NVIDIA cites ~150 tokens/watt for 70B FP8 models and 35× throughput improvement versus a Blackwell NVL72 for a 1-trillion-parameter model when paired with Vera Rubin NVL72.
The LPX rack packages 256 LPUs (LP30 chips) with 128 GB of aggregate SRAM and 40 PB/s of aggregate bandwidth, holding a full 70B FP8 model plus KV cache in a single rack with deterministic, flat decode latency. Deployment targets dense LLM serving (7B–70B) using a ~25% LPU / 75% GPU hybrid — GPUs handle prefill and attention math while LPUs offload feed-forward network (FFN) layers to their high-bandwidth SRAM.
Need NVIDIA Groq 3 LPU / LPX?
Contact QS Compute for availability, configuration, and volume pricing.
Request Quote