The capital requirements for AI reasoning infrastructure create a pricing squeeze for Microsoft and Meta. These companies must invest hundreds of billions to secure the compute capacity necessary for next-generation models. The industry must absorb these costs to maintain lead in AI reasoning. NVIDIA reports that its data center revenue reached $89 billion in the second quarter of fiscal 2026. This figure includes the ramp of Blackwell Ultra infrastructure. The company also reported a net income of $26.4 billion for that same period. This growth stems from the shift toward accelerated computing.
The Blackwell Ultra B300 technical specifications
The Blackwell Ultra B300 architecture targets AI reasoning through higher memory and throughput. The B300 has 288GB of HBM3e memory, which is 50% more than the 192GB in the standard B200. It provides 8 TB/s of memory bandwidth. The B300 has a 1,400W TDP, which is higher than the 1,000W TDP of the B200. This increased power draw necessitates liquid cooling solutions. The B300 also provides 1.5x more AI performance than the GB200 NVL72. This hardware is designed to handle large Mixture-of-Experts models with fewer GPUs. The additional memory bandwidth helps with decode tasks for large language models.
| Item | Price |
|---|---|
| B200 Cloud Rate | $5.99 – $16.11 per GPU-hour |
| B300 Cloud Rate | $7.10 – $17.80 per GPU-hour |
| DGX B200 (8 GPUs) | $280,000 – $320,000 (OEM) |
| DGX B200 (8 GPUs) | $515,000 (NVIDIA Official) |
| DGX B300 (8 GPUs) | $300,000 – $350,000 |
| DGX Station (1 GPU) | $80,000 – $125,000 |
| GB200 NVL72 (72 GPUs) | $2,000,000 – $3,000,000 |
| GB300 NVL72 (72 GPUs) | $3,000,000 – $4,000,000 |
Hardware costs for DGX and NVL72 systems
Hardware costs for DGX and NVL72 systems vary based on configuration and vendor. An 8-GPU DGX B300 system costs between $300,000 and $350,000. This works out to $37,500 to $43,750 per GPU. A single B300 GPU costs about $53,000. The B200 system costs $280,000 to $320,000 when purchased through an OEM. NVIDIA lists its official DGX B200 appliance at $515,000. This makes the per-GPU price $64,375 for the manufacturer-direct version. A GB300 NVL72 rack containing 72 GPUs costs between $3 million and $4 million. A desktop DGX Station with one GPU costs between $80,000 and $125,000. The price for an eight-way server node built on B200 or B300 sits in the $400,000 to $500,000 range.
Cloud rental and on-demand market rates
Cloud providers charge premiums for Blackwell access. B300 rates range from $7.10 to $17.80 per GPU-hour across five named on-demand providers. On-demand access to a B300 can cost as much as $18.00 per hour for managed DGX B300 stacks. A 48-month reservation for B300 brings the rate down to $3.13 per hour. B200 cloud rates range from $5.99 to $16.11 per GPU-hour across 11 named providers. Lambda Labs offers B200 at $5.29 per hour. The B300 price is higher than the B200 because of the extra memory. These rates reflect the difficulty of sourcing high-end SXM parts.
Microsoft Azure deployment of GB300 clusters
Microsoft Azure deploys the first at-scale production cluster using the GB300 NVL72. This cluster contains more than 4,600 units. Each rack in the system includes 72 Blackwell Ultra GPUs and 36 Grace CPUs. The system provides 130TB of per-second NVLink bandwidth. Each rack has 37TB of fast memory. The system also provides up to 1,440 petaflops of FP4 Tensor Core performance. While Microsoft Azure scales its deployment of the GB300 NVL72 to hundreds of thousands of Blackwell Ultra GPUs, the massive power requirements of these liquid-cooled racks force the company to invest heavily in new power distribution models and specialized cooling infrastructure. Microsoft plans to spend $120 billion on AI infrastructure in fiscal 2026. Azure uses ND GB300 v6 VMs to support reasoning models and agentic AI.
Meta infrastructure spending and MTIA roadmap
Meta plans to spend $145 billion on AI infrastructure this year. This spending supports a four-generation roadmap for its MTIA accelerators. Meta aims to deploy 7 gigawatts of computing infrastructure this year. The company plans to double this capacity by 2027. Meta uses Blackwell clusters to train Llama-4 and Llama-5 models. Meta also has a 6 gigawatt commitment to AMD MI450 GPUs starting in late 2026. The MTIA 300 is already deployed in production for internal recommendation models. The MTIA 450 will reach mass deployment in early 2027. Meta is working with Broadcom to develop these chips on a 2-nanometer process.
Supply chain constraints in packaging and memory
TSMC CoWoS packaging and SK Hynix HBM production limit the availability of NVIDIA chips. Packaging capacity is fully allocated through mid-2027. HBM3e production is difficult because higher die stacks require tighter tolerances. Each additional die in an HBM stack increases the loss rate by 5% to 8%. You already know that HBM capacity dictates the pace of AI scaling. These constraints mean that supply remains a bottleneck through at least the end of 2026. Samsung and Micron are adding capacity to help ease the shortage. However, these additions will not meaningfully ease the shortage before late 2026. The shortage of packaging slots and memory stacks creates a ceiling for how many GPUs can exist.
Competitive pressure from Google and Cerebras
Google and Cerebras provide alternatives to NVIDIA hardware. Google’s TPU 8t uses 9,600 chips in a superpod with 2 petabytes of HBM. The TPU 8i has 288GB of HBM per chip. The TPU 8i delivers 10.1 FP4 petaflops per chip. Cerebras CS-4 uses three WSE-3 Turbo chips per rack. The system has 750 petaFLOPs of compute. OpenAI uses Cerebras hardware to serve GPT-5.6 Sol. OpenAI’s agreement with Cerebras is valued at more than $20 billion. Amazon has also committed to spending $25 billion on Anthropic. This investment includes 5 gigawatts of Trainium capacity. Will custom silicon from Meta and Google eventually erode NVIDIA’s data center dominance?




