Meta’s Llama 4 models require massive memory capacity to manage their expanded context windows and parameter counts. Llama 4 Maverick, an expert-based model with 400 billion total parameters, needs significantly more VRAM than previous generations. To run Llama 4 Maverick at FP16 precision with a 4K token context, a system requires 8.4 TB of VRAM. For a 128K token context, the requirement climbs to 1.45 TB of VRAM. I see a clear divergence in total cost of ownership between AMD’s MI350X and Intel’s Gaudi 3 when scaling these specific workloads. AMD’s MI350X provides 288GB of HBM3e memory per GPU, while Intel’s Gaudi 3 provides 128GB of HBM2e memory.
Scaling the Llama 4 Herd
The Llama 4 model suite consists of three primary variants designed for different scales of intelligence. Llama 4 Scout has 17 billion active parameters and 109 billion total parameters. It supports a 10 million token context window. Llama 4 Maverick has 17 billion active parameters and 400 billion total parameters. It uses a mixture-of-experts architecture with 128 routed experts and one shared expert. Llama 4 Behemoth is the largest model, with 288 billion active parameters and 2 trillion total parameters.
Training Behemoth requires specialized infrastructure. For 4K tokens at FP8 weights and FP16 intermediates, Behemoth needs 63 TB of VRAM. This workload requires 32,000 or more GPUs and advanced memory-sharding techniques. Llama 4 Maverick is designed for data-center setups where inference occurs on multi-GPU clusters or H100 DGX hosts. You should look at the memory bandwidth if you plan to run Maverick at scale.
AMD MI350X Architecture and Compute
The AMD Instinct MI350X uses the CDNA 4 architecture built on a TSMC 3nm process. It contains 185 billion transistors. The GPU package includes eight Accelerator Compute Chiplets (XCDs). Each XCD has four shader engines, and every shader engine has eight active CDNA 4 compute units. This configuration results in 32 CUs per XCD and 256 CUs for the full accelerator. The I/O Die layer has two tiles instead of the four used in previous generations. This change doubles the Infinity Fabric bus width and reduces power consumption.
The MI350X delivers 2.3 PFLOPS of FP16 performance. It delivers 9.2 PFLOPS at MXFP4 and MXFP6. It also provides 4.6 PFLOPS at OCP-FP8. The device carries 288GB of HBM3e memory per GPU. The aggregate bandwidth for the HBM3e memory reaches 8 TB/s. The MI350X has a TBP of 1,000W. In an eight-GPU server like the Supermicro H14, the system provides 2.3 TB of total GPU memory. This single node can handle large models with over 500 billion parameters.
Intel Gaudi 3 Economics
Intel Gaudi 3 targets budget-sensitive enterprises. The market price for a single Gaudi 3 unit is $16,500. This is significantly lower than the $37,000 market price for the AMD MI350X. Gaudi 3 provides 1,835 TFLOPS of BF16/FP8 compute power. It has 128GB of HBM2e memory. The memory bandwidth for Gaudi 3 is 3.7 TB/s. It includes 64 Tensor Processing Cores and 8 Matrix Multiplication Engines.
The Gaudi 3 architecture uses a PCIe 5.0 x16 interface. It also uses Ethernet for chip-to-chip connectivity via RoCE. This allows Gaudi 3 to work with existing Ethernet switches. For the Chinese market, Intel produces a special edition of Gaudi 3. This version has a 150 TFLOPS limit in FP16 to comply with US export regulations. The special edition also reduces the number of cores and operating frequencies.
VRAM Demands for Maverick
Llama 4 Maverick presents a massive memory challenge for standard clusters. Because Maverick uses 400 billion total parameters, the weight storage alone is enormous. At FP16 precision, the 4K token requirement is 8.4 TB of VRAM. For a 128K context, the requirement is 1.45 TB of VRAM.
The AMD MI350X carries 288GB of memory per GPU. An eight-GPU node with MI350X provides 2,304 GB of VRAM. This allows a single AMD node to handle the 1.45 TB requirement for Maverick at 128K context. Intel Gaudi 3 provides 128GB per GPU. An eight-GPU node with Gaudi 3 provides 1,024 GB of VRAM. An Intel cluster would need two nodes to meet the 1.45 TB VRAM requirement for Maverick at 128K context.
| Feature | AMD MI350X | Intel Gaudi 3 |
|---|---|---|
| Memory Capacity | 288 GB HBM3e | 128 GB HBM2e |
| Memory Bandwidth | 8 TB/s | 3.7 TB/s |
| FP16 Performance | 2.3 PFLOPS | 1.835 PFLOPS |
| FP8 Performance | 4.6 PFLOPS | 1.835 PFLOPS |
| TDP | 1,000W | 900W |
| Market Price | $37,000 | $16,500 |
Memory Bandwidth and the Maverick Bottleneck
Memory bandwidth dictates how fast a model can process tokens during inference. The MI350X provides 8 TB/s of bandwidth per GPU. The Gaudi 3 provides 3.7 TB/s of bandwidth per GPU. This difference is critical for Llama 4 Maverick. The high parameter count means the system must move large amounts of data from memory to the compute units for every token.
AMD’s MI350X uses eight HBM3e memory stacks. Each stack is a 36GB device. The MI350X provides 1.3x higher memory bandwidth per watt compared to the MI300 series. Intel’s Gaudi 3 relies on HBM2e. While Gaudi 3 is a strong contender for cost-sensitive inference, it lacks the raw bandwidth of the MI350X. This bandwidth gap will become more apparent as context lengths increase.
Will Intel’s Ethernet-based approach eventually match AMD’s Infinity Fabric in large-scale clusters?
Connectivity and Software Ecosystems
AMD uses the Infinity Fabric for its interconnect. This fabric provides 896 GB/s of bandwidth. AMD’s ROCm 7.2 software stack includes support for PyTorch, JAX, TensorFlow, and vLLM. AMD and Meta announced a multi-year, multi-generation 6-gigawatt GPU deployment agreement in February 2026. This agreement builds on Meta’s existing production deployments of MI300 and MI350 series hardware. The vLLM project added a dedicated AMD ROCm CI pipeline in late 2025.
Intel Gaudi 3 uses Ethernet for connectivity. The Gaudi ecosystem grows through the Habana SynapseAI software. Intel supports PyTorch and other mainstream frameworks. Gaudi 3 can run on IBM Cloud and Intel Tiber Developer Cloud. Users can manage Gaudi clusters using existing Ethernet networking knowledge. However, AMD’s ROCm is catching up to the industry dominance of CUDA.
Thermal Management and Power Delivery
The power requirements for these chips affect the total cost of ownership in a data center. The AMD MI350X has a TDP of 1,000W. There is a liquid-cooled variant called the MI355X. The MI355X has a TBP of 1,400W. It provides a 20% performance advantage over the air-cooled MI350X.
Intel Gaudi 3 has a TDP of 900W. Lower power versions of these cards consume 600W. This lower power draw can reduce cooling costs in air-cooled data centers. For enterprises using existing air-cooled servers, the AMD MI350P is a relevant option. The MI350P is a dual-slot PCIe card with a 600W TBP. It is designed to fit into existing infrastructure without requiring liquid cooling.
The MI350P carries 144GB of HBM3e memory. It delivers 1,150 TFLOPS at FP16 precision. The MI350P is 43% faster in FP16 performance than the Nvidia H200 NVL. The price for the MI350P lands in the $30,000 to $40,000 range.
TCO Analysis and Final Verdict
The decision between AMD and Intel depends on the specific Llama 4 workload. If the priority is maximizing VRAM for Maverick’s context windows, AMD is the winner. The MI350X’s 288GB of HBM3e allows for much higher density. An AMD node can host the 1.45 TB VRAM requirement for Maverick at 128K context on a single server. An Intel cluster would require two nodes to provide the same capacity.
Intel Gaudi 3 is a strong choice for budget-sensitive training and inference. The $16,500 price point is much lower than the $37,000 price for the AMD MI350X. Gaudi 3 is ideal for enterprises that want to reduce costs for Llama 4 Scout. However, the lower memory bandwidth and capacity of Gaudi 3 make it harder to scale for the largest Llama 4 models.
I recommend the AMD MI350X for Meta’s Llama 4 clusters. The 288GB of HBM3e memory and 8 TB/s bandwidth provide the necessary headroom for Maverick’s massive parameter count. The higher upfront cost per GPU is offset by the ability to run larger models on fewer nodes. AMD’s growing software support and Meta’s massive multi-year commitment to the hardware make it the most reliable choice for large-scale Llama 4 deployments.




