NVIDIA
NVIDIA L40S 48GB AI accelerator card 900-2G133-0080-000
AI accelerator card with 48 GB GDDR6 ECC memory, 18,176 CUDA cores and Tensor cores delivering 733 teraFLOPS FP16 performance in a dual-slot PCIe 4.0 x16 form factor
- InterfacePCI Express 4.0 x16
- Chipset ManufacturerNVIDIA
- GPUL40S
- CUDA Cores18176
- Memory Size48GB
- Memory Clock864 GB/s
- 18176 CUDA cores accelerate parallel workloads
- 48GB GDDR6 memory with ECC supports large datasets
- 6 teraFLOPS FP32 performance drives compute tasks
- 1466 teraFLOPS FP8 Tensor throughput speeds AI inference
- PCIe 4.0 x16 interface provides high bandwidth connectivity
NVIDIA L40S 48GB AI accelerator
The NVIDIA L40S is a professional AI accelerator card built on the Ada architecture. It delivers 18176 CUDA cores and 48GB of GDDR6 memory with ECC for data integrity. The dual-slot design fits standard server chassis and targets machine-learning training, inference and high-performance visualisation workloads.
PCI Express 4.0 x16 interface limits
The card connects through a PCI Express 4.0 x16 slot providing up to 864 GB/s of memory bandwidth. Host systems must supply a compatible 4.0 x16 link to avoid throughput bottlenecks. The 350W maximum power draw requires adequate auxiliary connectors and chassis cooling capacity. Dual-slot width restricts adjacent slot availability in dense configurations.
Upgrade from lower-memory compute cards
This model replaces accelerators with less VRAM when 48GB is needed for larger models or batch sizes. The jump is worth it when workloads exceed the frame-buffer capacity of smaller cards and demand the 91.6 teraFLOPS FP32 or 1466 teraFLOPS FP8 throughput. Buyers who need more than 48GB or require multi-GPU scaling beyond a single dual-slot form factor should evaluate alternative platforms.
Highlights
- 18176 CUDA cores accelerate parallel workloads
- 48GB GDDR6 memory with ECC supports large datasets
- 6 teraFLOPS FP32 performance drives compute tasks
- 1466 teraFLOPS FP8 Tensor throughput speeds AI inference
- PCIe 4.0 x16 interface provides high bandwidth connectivity
Specifications
| Brand | NVIDIA |
|---|---|
| Model | L40S |
| Part Number | 900-2G133-0080-000 |
| Interface | PCI Express 4.0 x16 |
| Chipset Manufacturer | NVIDIA |
| GPU | L40S |
| CUDA Cores | 18176 |
| Memory Size | 48GB |
| Memory Clock | 864 GB/s |
| Memory Type | GDDR6 w/ECC |
| DirectX | DirectX 12.07 |
| DisplayPort | 4 |
| Dimensions (L x H) | Height: 4.4" Length:10.5" |
| Slot Width | Dual |
| Shipping weight | 1.00 kg |
| Package size | 279 × 102 × 51 mm |
Questions about this item
Should I pick the 48GB GDDR6 w/ECC version for large-model training?
The 48GB frame buffer with 864 GB/s bandwidth lets you hold larger batches and bigger layers on a single card, reducing the need for multi-GPU splitting.
What does the 91.6 teraFLOPS FP32 rating mean for real-time inference?
That peak throughput translates to high single-precision compute density for inference pipelines, letting you serve more requests per second within the 350W power envelope.
How does the PCI Express 4.0 x16 interface affect data throughput?
The x16 link provides up to 32 GB/s bidirectional bandwidth, which caps host-to-device transfer rates for dataset loading and model checkpointing.
New to this? Buying server hardware with crypto — what to check, what it costs to pay, and what is never asked for.
Also in this aisle
People compared these
Same shelf, same checkout — eight coins and a 30-minute rate lock.
Lenovo 4X60N04886 Nvidia Quadro P5000 16GB GDDR5X ai accelerator card
- GDDR5
NVIDIA H200 NVL AI Accelerator Card 141GB HBM3e PCIe 5.0 x16 900-21010-0040-000
- PCI Express 5.0 x16
- NVIDIA



