AMD MI350X memory bandwidth and CDNA 4 architecture explained

The AMD MI350X features 288 GB of HBM3E memory and 8.0 TB/s bandwidth, allowing it to run 400B+ parameter models using only three GPUs. This CDNA 4 architecture offers significant performance advantages over NVIDIA alternatives for large-scale AI…

AMD MI350X memory bandwidth and CDNA 4 architecture explained

CDNA 4 memory and chiplet scale

AMD’s MI350X carries 288 GB of HBM3E memory for large-scale AI workloads. This capacity is 50% higher than the 192 GB on the MI300X. The memory bandwidth reaches 8.0 TB/s, which is a significant increase over the 5.3 TB/s on the MI300X. The CDNA 4 architecture uses TSMC’s N3P 3nm process node for its compute dies and utilizes eight Accelerator Compute Dies stacked atop four base dies to reach 185 billion transistors. Each XCD contains 32 CDNA 4 Compute Units, giving the full accelerator 256 CUs. This design handles 400B+ parameter models using only three MI350X GPUs at FP16, whereas the NVIDIA B200 requires five GPUs for the same model. You know that fitting large models into VRAM often matters more than raw TFLOPS. AMD uses Infinity Fabric for multi-GPU scaling, which provides 1,075 GB/s of bidirectional aggregate bandwidth per socket. The MI355X variant uses liquid cooling to reach a 1,400W TBP, which provides a 20% performance advantage over the air-cooled MI350X. The MI350X provides 256 MB of Infinity Cache on the base dies. This cache includes 128 channels, each with 2 MB of capacity. The MI350X provides 72.1 TFLOPS of FP64 throughput, which is over double the 34 TFLOPS from the H100. This 2.1x advantage makes the MI350X a strong candidate for scientific and HPC workloads.

Compute density and precision types

The MI350X delivers 4,614 TFLOPS of FP8 dense throughput. This figure is over twice the 1,979 TFLOPS provided by the NVIDIA H100. CDNA 4 includes native hardware support for MXFP4 and MXFP6 formats. These formats allow for aggressive quantization. The MI350P, a PCIe variant, carries 144 GB of HBM3E memory and has a 4.0 TB/s bandwidth. It has a 600W TBP, which is configurable down to 450W. The MI350P delivers 4,600 TFLOPS at MXFP4 precision. The MI350P contains 128 compute units, 512 matrix cores, and 8,192 stream processors running at 2.2 GHz. This PCIe card delivers 39% better FP8 performance and 43% better FP16 performance than the NVIDIA H200 NVL. The CDNA 4 architecture adds a new vector ALU to the Compute Unit. This ALU supports 2-bit operations and can accumulate BF16 results into FP32. The Local Data Share capacity is 160 KB per Compute Unit, which is an increase from the 64 KB in CDNA 3. This expanded capacity doubles the read bandwidth to 256 bytes per clock cycle.

Specification AMD MI350X NVIDIA B200 NVIDIA H100
Memory Capacity 288 GB 192 GB 80 GB
Memory Bandwidth 8.0 TB/s 8.0 TB/s 2.0 TB/s
FP8 Dense Throughput 4,614 TFLOPS – 1,513 TFLOPS
FP4 Dense Throughput 9,227.5 TFLOPS – Not supported

The MI350X provides 2,309.6 TFLOPS at FP16 precision. The H100 provides 989.5 TFLOPS at the same precision. The MI350X produces 2.2x more FP32 throughput than the H100.

Software and deployment economics

AMD’s ROCm 7.x supports PyTorch, JAX, vLLM, and SGLang. The software stack includes HIP for porting CUDA code. ROCm still trails CUDA 13.x in ecosystem depth. The MI350X provides a 33% lower hardware cost per token for FP8 deployments of large models compared to the B200. A Llama 4 Maverick deployment at FP16 requires 800 GB of memory. This model needs 3 MI350X GPUs but needs 5 B200 GPUs. This memory advantage creates a 40% lower hardware cost per token for FP16 workloads. The MI350P is a standard full-height, full-length, dual-slot PCIe card. It fits into off-the-shelf air-cooled servers. Companies like Dell, HPE, and Supermicro provide qualified servers for the MI350P. DigitalOcean lists MI350X at $6.16 per GPU-hour. Spheron lists the B200 at $4.42 per GPU-hour. Will the ROCm ecosystem catch up to CUDA’s maturity for specialized kernels? The MI350X is the better choice for large-parameter inference workloads where memory capacity dictates deployment scale.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories