Specifications

MPN

900-2G133-0080-000

Series

NVIDIA L40S

Architecture

NVIDIA Ada Lovelace

GPU Memory

48 GB GDDR6 with ECC

Memory Bandwidth

864 GB/s

CUDA Cores

18,176

Tensor Cores

568 (4th generation)

RT Cores

142 (3rd generation)

FP32

91.6 TFLOPS

Tensor FP16

362.05 TFLOPS (733 sparse)

Tensor FP8

733 TFLOPS (1,466 sparse)

INT8 Tensor TOPS

733 (1,466 sparse)

Interconnect

PCIe Gen4 x16 — 64 GB/s bidirectional

Form Factor

4.4" H x 10.5" L, dual slot

Thermal

Passive

Max Power

350 W

Power Connector

16-pin (12V-2x6)

Display Outputs

4x DisplayPort 1.4a

vGPU Support

Yes

NVLink

Not supported

NEBS

Level 3 ready

NVIDIA L40 family comparison

ProductArchitectureMemoryMemory BandwidthMax PowerForm Factor
NVIDIA L40SAda Lovelace48 GB GDDR6 ECC864 GB/s350 WDual slot, passive
NVIDIA L40Ada Lovelace48 GB GDDR6 ECC864 GB/s300 WDual slot, passive
NVIDIA A100 80GB PCIeAmpere80 GB HBM2e1,935 GB/s300 WDual slot, passive

Use Case

The L40S sits between the A100 and the L40 in NVIDIA's line-up: it keeps the 48 GB GDDR6 frame buffer of the L40 but roughly doubles the Tensor Core throughput, which makes it the usual pick for LLM inference at moderate batch sizes, fine-tuning of 7B–13B models, and mixed graphics-plus-AI workloads in a 4U server. Because it is a passive, 350 W dual-slot card, it is deployed in NVIDIA-Certified Systems where chassis airflow is provided by the host.

Need NVIDIA L40S 48GB PCIe GPU Accelerator?

Enterprise hardware — configured, tested, deployed. Contact us for pricing and availability.

Request Quote