We take cryptocurrency — so the coins you hold can buy real hardware.Pay in crypto, spend it on real hardware. Free insured shipping worldwide over $399. Plain, unbranded boxes, sent from Germany.Free insured shipping over $399. Your rate is locked for 30 minutes once checkout opens.Rate locked 30 minutes. Every item sealed, serial-checked and covered by a 24-month warranty.Sealed, 24-month warranty.

Saved items Sign in

PNY Technologies, Inc.

PNY Technologies, Inc. 699-2G133-0200-C00 NVIDIA A40 48GB GDDR6 ai accelerator card

SKU CH-NMG7-NX0MF MPN 699-2G133-0200-C00

AI accelerator with 48 GB GDDR6, 10,752 CUDA cores, 336 tensor cores and 696 GB/s bandwidth on PCIe 4.0 x16

  • 48 GB GDDR6 memory handles large AI models
  • 10,752 CUDA cores accelerate parallel workloads
  • 336 third-generation tensor cores speed mixed-precision training
  • 696 GB/s memory bandwidth feeds data fast
  • PCIe 4.0 x16 interface delivers high host throughput

48GB GDDR6 AI Accelerator for Data-Centres

The PNY NVIDIA A40 is a new 48GB GDDR6 accelerator card built on the Ampere architecture. It delivers 10,752 CUDA cores, 336 third-generation Tensor cores and 84 second-generation RT cores. A 384-bit memory interface provides 696GB/s bandwidth across 48GB of ECC-protected GDDR6 memory. The PCIe 4.0 x16 bus and NVLink at 112.5GB/s bidirectional speed enable multi-GPU scaling for large-model training and inference workloads.

Suited for High-Memory AI Workloads.

This card targets researchers and engineers training large language models, running high-resolution rendering or performing scientific simulation where 48GB of frame buffer and ECC are required. The passive heatsink demands server chassis with directed high-volume airflow; it will not operate reliably in workstations or servers that rely on active GPU fans. Buyers needing active cooling, a lower memory footprint or a single-slot form factor should evaluate alternative accelerator designs.

384-Bit GDDR6 Bus Determines Model Fit

The single specification that decides fit is the 384-bit GDDR6 memory interface delivering 696GB/s bandwidth. That width and bandwidth define the maximum model size and batch throughput the card can sustain without offloading to system memory. If the target workload fits within 48GB and can saturate the 384-bit bus, the A40 matches; if the dataset exceeds 48GB or requires higher memory bandwidth, a different accelerator class is needed.

Highlights

  • 48 GB GDDR6 memory handles large AI models
  • 10,752 CUDA cores accelerate parallel workloads
  • 336 third-generation tensor cores speed mixed-precision training
  • 696 GB/s memory bandwidth feeds data fast
  • PCIe 4.0 x16 interface delivers high host throughput

Specifications

No published specifications for this item yet.

Questions about this item

How do I install the card in a PCIe 4.0 x16 slot and what step is often missed?

Seat the card fully in a PCIe 4.0 x16 slot, then secure the bracket screw. The step often missed is verifying that the motherboard BIOS has PCIe 4.0 enabled for that slot so the x16 link trains at full speed.

What maintenance will the passive heatsink need over time?

The passive heatsink has no moving parts. Maintenance is limited to periodic dust removal from the fin stack using compressed air so airflow through the server chassis remains unrestricted.

At what point does the 696 GB/s memory bandwidth become the limiting factor?

When kernels saturate the 384-bit GDDR6 interface at 7251 MHz, the 696 GB/s ceiling is reached before the 10,752 CUDA cores or 336 third-gen Tensor cores hit their compute peaks.

New to this? Buying server hardware with crypto — what to check, what it costs to pay, and what is never asked for.

Also in this aisle

People compared these

Same shelf, same checkout — eight coins and a 30-minute rate lock.