Specifications

Platform

Lenovo ThinkSystem SC777 V4 Neptune server tray

Architecture

NVIDIA GB200 NVL4 — Grace CPUs and Blackwell GPUs in one tightly coupled node

CPUs

2x NVIDIA Grace, 72 Arm Neoverse V2 cores each

GPUs

4x NVIDIA B200 (GB200-class modules)

LPDDR5x Memory

960 GB ECC, 480 GB per CPU

HBM3e Memory

744 GB total, 186 GB per GPU

CPU-to-CPU Interconnect

300 GB/s bidirectional NVLink-C2C

CPU-to-GPU Interconnect

450 GB/s bidirectional NVLink-C2C with unified coherent memory

GPU-to-GPU Interconnect

600 GB/s bidirectional NVLink

Internal Storage

Up to 10x E3.S SSDs

Networking

NVIDIA NDR or XDR via CPU PCIe slots or GPU-direct modules, up to 800 Gbps per GPU with GPUDirect

Enclosure

ThinkSystem N1380 Neptune enclosure, 13U, with 8x SC777 V4 servers installed vertically

Rack Density

Up to 96 GPUs per standard 19-inch rack width

Cooling

100% direct water-cooled design eliminating airflow constraints

Inlet Water Temperature

Supported up to 45 C, allowing some data centres to run without chilled water

Measured GPU HPL

149 TF/s against a 160 TF/s peak, and 343 TF/s with FP64 emulation

Measured GPU GEMM

6,729 TF/s at FP4 and 3,001 TF/s at FP8

Measured Interconnect

588 GB/s NCCL All-Reduce GPU-to-GPU against a 600 GB/s theoretical peak

Performance Verdict

Operating at or within 5% of expected levels across all major subsystems on pre-production systems

Workloads

Traditional HPC, LLM training and inference, computer vision, generative AI, graph analytics, AI for science

Overview

The Lenovo ThinkSystem SC777 V4 is a GB200 NVL4 node: two NVIDIA Grace CPUs and four B200 GPUs on a single tray, joined by NVLink-C2C. Each Grace package carries 72 Arm Neoverse V2 cores, and the tray holds 960 GB of ECC LPDDR5x for the CPUs plus 744 GB of HBM3e for the GPUs. Critically, the CPU-to-GPU link runs at 450 GB/s bidirectional with unified coherent memory, which means the GPUs and CPUs share one address space and the software stack stops paying for explicit data migration between them.

The practical consequence of NVL4 is that a single tray behaves like one large compute domain rather than four discrete accelerators plus two hosts. GPU-to-GPU NVLink runs at 600 GB/s bidirectional, and the CPU-to-CPU fabric at 300 GB/s. Lenovo's measured NCCL All-Reduce result of 588 GB/s against a 600 GB/s theoretical peak — 98% of peak — is the clearest evidence that the fabric is not the limiting factor for distributed training.

Lenovo's published early performance numbers put the tray at 149 TF/s on GPU HPL against a 160 TF/s peak, 343 TF/s with FP64 emulation, 6,729 TF/s on FP4 GEMM and 3,001 TF/s on FP8 GEMM, with GPU STREAM Triad at 7,508 GB/s against an 8,000 GB/s peak. The vendor's summary is that the system operates at or within 5% of expected levels across every major subsystem. Those figures come from pre-production laboratory systems, so they are directional rather than final.

The deployment story is the Neptune enclosure. Eight SC777 V4 trays install vertically inside a 13U N1380 Neptune chassis, and Lenovo specifies up to 96 GPUs in a standard 19-inch rack width. The design is 100% direct water cooled with no airflow path at all, and Lenovo rates it for inlet water up to 45 C — which allows facilities in temperate climates to run without mechanical chillers. For buyers comparing GB200 NVL4 against NVL72 rack-scale systems, the SC777 V4 is the option that keeps a familiar server-tray operational model while accepting a direct liquid loop.

Key Benefits

NVLink-C2C coherent memory at 450 GB/s removes explicit CPU-GPU data movement and raises GPU utilisation. 4 GPUs plus 2 Grace CPUs per tray gives a self-contained training and inference domain. 96 GPUs per 19-inch rack delivers density without a bespoke rack standard. 100% direct water cooling with 45 C inlet tolerance can eliminate mechanical chillers entirely. Measured performance within 5% of expected across all subsystems derisks a first deployment. Up to 800 Gbps per GPU with GPUDirect keeps large-model training fed.

Applications

Large language model training and fine-tuning; AI inference at scale; traditional HPC — CFD, molecular dynamics, quantum chemistry, FEA, weather modelling; AI-for-science surrogate modelling and drug discovery; computer vision and generative AI training; graph analytics; mixed HPC and AI production clusters.

Request a Quote — LENOVO THINKSYSTEM SC777 V4 — GB200 NVL4 NEPTUNE LIQUID-COOLED AI NODE

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA GB200 NVL72 NVIDIA DGX B200 Supermicro SYS-821GE-TNHR AMD Threadripper Halo Station