Intel Gaudi 3 Architecture · 64 TPC + 8 MME · 600W TDP · PCIe Gen 5 x16 · 24× 200GbE · Open Ethernet-native AI
Gaudi 3 HL-338
Intel Gaudi 3
5nm
128 GB HBM2e
3.7 TB/s
96 MB · 19.2 TB/s
8 MME + 64 TPC
64 (VLIW SIMD)
1,835 TFLOPS
917 TFLOPS
Gen 5.0 x16
128 GB/s Bidirectional
600W (Air Cooling)
Full-Height Dual-Slot
10.5" PCIe Card
FP8, BF16, FP16, TF32, FP32
24× 200GbE RoCE v2
RDMA
900 GB/s per card
512 GB (4 × 128 GB)
1,024 GB (8 × 128 GB)
Dedicated HW Decode/Pre-process
PCIe · Mezzanine (HL-325L)
UBB (HLB-325)
The Intel Gaudi 3 AI Accelerator (HL-338) is Intel's third-generation deep learning processor, delivering 1,835 TFLOPS of FP8 AI compute in a standard PCIe Gen 5.0 add-in card form factor. Built on a 5nm process with 128 GB of HBM2e memory at 3.7 TB/s bandwidth, Gaudi 3 is purpose-built for large-scale AI training and inference with an open, Ethernet-native architecture.
Unlike proprietary GPU interconnects, Gaudi 3 uses standard 200GbE RoCE v2 RDMA with 24 integrated ports — eliminating the need for InfiniBand or NVLink. Each accelerator connects directly to Ethernet switches, enabling cost-effective scale-out to thousands of nodes. For scale-up, 4-card quads deliver 900 GB/s of inter-card bandwidth with 512 GB of pooled HBM2e memory, and 8-card configurations double that to 1,024 GB.
Gaudi 3's 64 TPC (Tensor Processor Cores) and 8 MME (Matrix Multiplication Engines) are natively designed for deep learning workloads including large language models, diffusion models, and multi-modal AI. The dedicated media processor offloads image/video decode and pre-processing. Available in PCIe, mezzanine (HL-325L), and UBB (HLB-325) form factors, Gaudi 3 integrates with Dell PowerEdge XE7740 and other OEM server platforms. QS Compute can source Intel Gaudi 3 accelerators for enterprise AI infrastructure deployments.
Need Intel Gaudi 3 Accelerators?
Enterprise hardware — configured, tested, deployed. Contact us for pricing and availability.
Request Quote