Specifications
Platform
Lenovo ThinkSystem SC777 V4 Neptune server tray
Architecture
NVIDIA GB200 NVL4 — Grace CPUs and Blackwell GPUs in one tightly coupled node
CPUs
2x NVIDIA Grace, 72 Arm Neoverse V2 cores each
GPUs
4x NVIDIA B200 (GB200-class modules)
LPDDR5x Memory
960 GB ECC, 480 GB per CPU
HBM3e Memory
744 GB total, 186 GB per GPU
CPU-to-CPU Interconnect
300 GB/s bidirectional NVLink-C2C
CPU-to-GPU Interconnect
450 GB/s bidirectional NVLink-C2C with unified coherent memory
GPU-to-GPU Interconnect
600 GB/s bidirectional NVLink
Internal Storage
Up to 10x E3.S SSDs
Networking
NVIDIA NDR or XDR via CPU PCIe slots or GPU-direct modules, up to 800 Gbps per GPU with GPUDirect
Enclosure
ThinkSystem N1380 Neptune enclosure, 13U, with 8x SC777 V4 servers installed vertically
Rack Density
Up to 96 GPUs per standard 19-inch rack width
Cooling
100% direct water-cooled design eliminating airflow constraints
Inlet Water Temperature
Supported up to 45 C, allowing some data centres to run without chilled water
Measured GPU HPL
149 TF/s against a 160 TF/s peak, and 343 TF/s with FP64 emulation
Measured GPU GEMM
6,729 TF/s at FP4 and 3,001 TF/s at FP8
Measured Interconnect
588 GB/s NCCL All-Reduce GPU-to-GPU against a 600 GB/s theoretical peak
Performance Verdict
Operating at or within 5% of expected levels across all major subsystems on pre-production systems
Workloads
Traditional HPC, LLM training and inference, computer vision, generative AI, graph analytics, AI for science
Overview
The Lenovo ThinkSystem SC777 V4 is a GB200 NVL4 node: two NVIDIA Grace CPUs and four B200 GPUs on a single tray, joined by NVLink-C2C. Each Grace package carries 72 Arm Neoverse V2 cores, and the tray holds 960 GB of ECC LPDDR5x for the CPUs plus 744 GB of HBM3e for the GPUs. Critically, the CPU-to-GPU link runs at 450 GB/s bidirectional with unified coherent memory, which means the GPUs and CPUs share one address space and the software stack stops paying for explicit data migration between them.
The practical consequence of NVL4 is that a single tray behaves like one large compute domain rather than four discrete accelerators plus two hosts. GPU-to-GPU NVLink runs at 600 GB/s bidirectional, and the CPU-to-CPU fabric at 300 GB/s. Lenovo's measured NCCL All-Reduce result of 588 GB/s against a 600 GB/s theoretical peak — 98% of peak — is the clearest evidence that the fabric is not the limiting factor for distributed training.
Lenovo's published early performance numbers put the tray at 149 TF/s on GPU HPL against a 160 TF/s peak, 343 TF/s with FP64 emulation, 6,729 TF/s on FP4 GEMM and 3,001 TF/s on FP8 GEMM, with GPU STREAM Triad at 7,508 GB/s against an 8,000 GB/s peak. The vendor's summary is that the system operates at or within 5% of expected levels across every major subsystem. Those figures come from pre-production laboratory systems, so they are directional rather than final.
The deployment story is the Neptune enclosure. Eight SC777 V4 trays install vertically inside a 13U N1380 Neptune chassis, and Lenovo specifies up to 96 GPUs in a standard 19-inch rack width. The design is 100% direct water cooled with no airflow path at all, and Lenovo rates it for inlet water up to 45 C — which allows facilities in temperate climates to run without mechanical chillers. For buyers comparing GB200 NVL4 against NVL72 rack-scale systems, the SC777 V4 is the option that keeps a familiar server-tray operational model while accepting a direct liquid loop.
Key Benefits
NVLink-C2C coherent memory at 450 GB/s removes explicit CPU-GPU data movement and raises GPU utilisation. 4 GPUs plus 2 Grace CPUs per tray gives a self-contained training and inference domain. 96 GPUs per 19-inch rack delivers density without a bespoke rack standard. 100% direct water cooling with 45 C inlet tolerance can eliminate mechanical chillers entirely. Measured performance within 5% of expected across all subsystems derisks a first deployment. Up to 800 Gbps per GPU with GPUDirect keeps large-model training fed.
Applications
Large language model training and fine-tuning; AI inference at scale; traditional HPC — CFD, molecular dynamics, quantum chemistry, FEA, weather modelling; AI-for-science surrogate modelling and drug discovery; computer vision and generative AI training; graph analytics; mixed HPC and AI production clusters.
Request a Quote — LENOVO THINKSYSTEM SC777 V4 — GB200 NVL4 NEPTUNE LIQUID-COOLED AI NODE
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →