Specifications
MPN
900-2G133-0080-000
Series
NVIDIA L40S
Architecture
NVIDIA Ada Lovelace
GPU Memory
48 GB GDDR6 with ECC
Memory Bandwidth
864 GB/s
CUDA Cores
18,176
Tensor Cores
568 (4th generation)
RT Cores
142 (3rd generation)
FP32
91.6 TFLOPS
Tensor FP16
362.05 TFLOPS (733 sparse)
Tensor FP8
733 TFLOPS (1,466 sparse)
INT8 Tensor TOPS
733 (1,466 sparse)
Interconnect
PCIe Gen4 x16 — 64 GB/s bidirectional
Form Factor
4.4" H x 10.5" L, dual slot
Thermal
Passive
Max Power
350 W
Power Connector
16-pin (12V-2x6)
Display Outputs
4x DisplayPort 1.4a
vGPU Support
Yes
NVLink
Not supported
NEBS
Level 3 ready
NVIDIA L40 family comparison
| Product | Architecture | Memory | Memory Bandwidth | Max Power | Form Factor |
|---|---|---|---|---|---|
| NVIDIA L40S | Ada Lovelace | 48 GB GDDR6 ECC | 864 GB/s | 350 W | Dual slot, passive |
| NVIDIA L40 | Ada Lovelace | 48 GB GDDR6 ECC | 864 GB/s | 300 W | Dual slot, passive |
| NVIDIA A100 80GB PCIe | Ampere | 80 GB HBM2e | 1,935 GB/s | 300 W | Dual slot, passive |
Use Case
The L40S sits between the A100 and the L40 in NVIDIA's line-up: it keeps the 48 GB GDDR6 frame buffer of the L40 but roughly doubles the Tensor Core throughput, which makes it the usual pick for LLM inference at moderate batch sizes, fine-tuning of 7B–13B models, and mixed graphics-plus-AI workloads in a 4U server. Because it is a passive, 350 W dual-slot card, it is deployed in NVIDIA-Certified Systems where chassis airflow is provided by the host.
Need NVIDIA L40S 48GB PCIe GPU Accelerator?
Enterprise hardware — configured, tested, deployed. Contact us for pricing and availability.
Request Quote