Specifications

Product

QuantaGrid D75E-4U rackmount server

Architecture

NVIDIA MGX modular design on an Intel Xeon platform

Processors

2x Intel Xeon 6, up to 350 W TDP each

GPU Support (NVIDIA)

Up to 8x double-width 600 W PCIe GPUs, including RTX PRO 6000, H200 NVL, L40S, L4, A2 and A16

Accelerator Support (Intel)

Up to 8x Intel Gaudi 3 AI accelerators in PCIe form factor at 600 W

Flexible GPU Count

Configurable with 1, 2, 4 or 8 GPUs to match existing rack infrastructure

PCIe Slots (no switch)

4x double-width FHFL PCIe 5.0 x16 for GPU plus 3x single-width FHFL for networking

PCIe Slots (with switch)

8x double-width FHFL PCIe 5.0 x16 for GPU, 4x single-width FHFL plus 1x FHHL and 1x HHHL for networking

Slot Power

All PCIe 5.0 expansion slots designed to supply up to 150 W

Memory

32x DDR5 RDIMM up to 6,400 MHz plus 16x MRDIMM up to 8,000 MHz

Storage SKU 1

12x hot-swappable E1.S SSDs

Storage SKU 2

24x hot-swappable E1.S SSDs

Power Supply

3+1 high-efficiency redundant hot-plug Titanium 2700 W or 3200 W AC

Cooling

10 dual-rotor fans with a remote heatsink solution for GPU thermal efficiency

Serviceability

Tool-less and hot-pluggable design

Management

Orqestra (QCT System Manager) monitoring up to 5,000 devices per module

Form Factor

4U, 438 x 176 x 800 mm

BMC

Integrated ASPEED AST2600 video with 1 GbE dedicated management

Security

Optional TPM 2.0 SPI module

Environment

5 to 35 C operating; 50 to 85% relative humidity

Overview

The QuantaGrid D75E-4U is Quanta Cloud Technology's volume AI inference and HPC node. It is built to the NVIDIA MGX modular reference architecture but sits on dual Intel Xeon 6 processors, which gives it an unusual property for a 4U GPU platform: the same chassis accepts either eight NVIDIA PCIe GPUs at up to 600 W or eight Intel Gaudi 3 accelerators at 600 W. Buyers can standardise on one chassis, power and cooling design and still keep two accelerator supply chains open.

On the NVIDIA side the supported list is deliberately the PCIe rather than SXM tier — RTX PRO 6000, H200 NVL, L40S, L4, A2 and A16 — with QCT noting that H200 NVL and RTX PRO 6000 are particularly well suited to low-power air-cooled enterprise racks. Configurations run from one to eight GPUs, so a chassis can be populated to whatever density the existing power and thermal budget allows rather than forcing a full eight-way build on day one.

Beyond the accelerators, the D75E is a current-generation server: 32 DDR5 RDIMM slots up to 6,400 MHz plus 16 MRDIMM slots up to 8,000 MHz, either 12 or 24 hot-swappable E1.S SSDs, 3+1 redundant Titanium power supplies at 2700 W or 3200 W, and a remote-heatsink layout that reduces GPU pre-heating and lets fans run slower. Management runs through QCT's Orqestra system manager, which monitors up to 5,000 devices per module.

This is the node shape that most enterprise AI inference actually ships in: air cooled, standard 4U, PCIe accelerators, and a chassis that does not require a datacentre redesign. QS Compute quotes the D75E alongside the SXM-class 8-GPU systems and the E1.S storage tiers that populate it.

Key Benefits

NVIDIA MGX plus Intel Gaudi 3 support in one chassis keeps two accelerator supply chains open on a single qualified platform. 1, 2, 4 or 8 GPU configurations let you match the existing rack power budget instead of rebuilding it. Air-cooled 600 W PCIe GPUs avoid a liquid infrastructure project. Up to 24 E1.S SSDs provide the local capacity that inference caching and dataset staging need. 3+1 Titanium PSUs at up to 3200 W give headroom for fully populated eight-GPU builds. Remote heatsink and 10 dual-rotor fans lower GPU pre-heating and fan power.

Applications

Enterprise AI inference in air-cooled racks; LLM serving with PCIe H200 NVL or RTX PRO 6000; HPC and simulation clusters; mixed-silicon accelerator evaluation and production; video analytics and computer-vision inference at scale; private-cloud AI model hosting; research clusters needing high E1.S local storage density.

Request a Quote — QUANTA QCT QUANTAGRID D75E-4U — NVIDIA MGX 8-GPU AI SERVER WITH XEON 6

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Supermicro SYS-421GE-TNRT GPU Server GIGABYTE G593-SD0 8-GPU Server Inventec Atlas Edge AI Server HPE Compute XD690 8x B300 Server