We take cryptocurrency — so the coins you hold can buy real hardware.Pay in crypto, spend it on real hardware. Free insured shipping worldwide over $399. Plain, unbranded boxes, sent from Germany.Free insured shipping over $399. Your rate is locked for 30 minutes once checkout opens.Rate locked 30 minutes. Every item sealed, serial-checked and covered by a 24-month warranty.Sealed, 24-month warranty.

Saved items Sign in

NVIDIA

NVIDIA L40S 48GB AI accelerator card 900-2G133-0080-000

SKU CH-9EV2-EGKCE MPN 900-2G133-0080-000

AI accelerator card with 48 GB GDDR6 ECC memory, 18,176 CUDA cores and Tensor cores delivering 733 teraFLOPS FP16 performance in a dual-slot PCIe 4.0 x16 form factor

  • InterfacePCI Express 4.0 x16
  • Chipset ManufacturerNVIDIA
  • GPUL40S
  • CUDA Cores18176
  • Memory Size48GB
  • Memory Clock864 GB/s
  • 18176 CUDA cores accelerate parallel workloads
  • 48GB GDDR6 memory with ECC supports large datasets
  • 6 teraFLOPS FP32 performance drives compute tasks
  • 1466 teraFLOPS FP8 Tensor throughput speeds AI inference
  • PCIe 4.0 x16 interface provides high bandwidth connectivity

NVIDIA L40S 48GB AI accelerator

The NVIDIA L40S is a professional AI accelerator card built on the Ada architecture. It delivers 18176 CUDA cores and 48GB of GDDR6 memory with ECC for data integrity. The dual-slot design fits standard server chassis and targets machine-learning training, inference and high-performance visualisation workloads.

PCI Express 4.0 x16 interface limits

The card connects through a PCI Express 4.0 x16 slot providing up to 864 GB/s of memory bandwidth. Host systems must supply a compatible 4.0 x16 link to avoid throughput bottlenecks. The 350W maximum power draw requires adequate auxiliary connectors and chassis cooling capacity. Dual-slot width restricts adjacent slot availability in dense configurations.

Upgrade from lower-memory compute cards

This model replaces accelerators with less VRAM when 48GB is needed for larger models or batch sizes. The jump is worth it when workloads exceed the frame-buffer capacity of smaller cards and demand the 91.6 teraFLOPS FP32 or 1466 teraFLOPS FP8 throughput. Buyers who need more than 48GB or require multi-GPU scaling beyond a single dual-slot form factor should evaluate alternative platforms.

Highlights

  • 18176 CUDA cores accelerate parallel workloads
  • 48GB GDDR6 memory with ECC supports large datasets
  • 6 teraFLOPS FP32 performance drives compute tasks
  • 1466 teraFLOPS FP8 Tensor throughput speeds AI inference
  • PCIe 4.0 x16 interface provides high bandwidth connectivity

Specifications

BrandNVIDIA
ModelL40S
Part Number900-2G133-0080-000
InterfacePCI Express 4.0 x16
Chipset ManufacturerNVIDIA
GPUL40S
CUDA Cores18176
Memory Size48GB
Memory Clock864 GB/s
Memory TypeGDDR6 w/ECC
DirectXDirectX 12.07
DisplayPort4
Dimensions (L x H)Height: 4.4"
Length:10.5"
Slot WidthDual
Shipping weight1.00 kg
Package size279 × 102 × 51 mm

Questions about this item

Should I pick the 48GB GDDR6 w/ECC version for large-model training?

The 48GB frame buffer with 864 GB/s bandwidth lets you hold larger batches and bigger layers on a single card, reducing the need for multi-GPU splitting.

What does the 91.6 teraFLOPS FP32 rating mean for real-time inference?

That peak throughput translates to high single-precision compute density for inference pipelines, letting you serve more requests per second within the 350W power envelope.

How does the PCI Express 4.0 x16 interface affect data throughput?

The x16 link provides up to 32 GB/s bidirectional bandwidth, which caps host-to-device transfer rates for dataset loading and model checkpointing.

New to this? Buying server hardware with crypto — what to check, what it costs to pay, and what is never asked for.

Also in this aisle

People compared these

Same shelf, same checkout — eight coins and a 30-minute rate lock.