xAI expanded the Memphis Colossus site to a 2 gigawatt capacity with 555,000 GPUs. This $18 billion investment includes NVIDIA GB200 and GB300 models at an average cost of $32,400 per GPU. The company uses seven 35-megawatt gas turbines in Southaven to provide electricity. xAI also uses 168 Tesla Megapacks to stabilize voltage during GPU power spikes. The Memphis site bypasses traditional utility constraints through on-site generation using gas-fired power plants and Tesla Megapacks to manage the 2 gigawatt load required for 555,000 GPUs.
Colossus 1 houses 230,000 accelerators. This site includes 30,000 GB200 units. The facility provides 500 megawatts of power. Anthropic leases this first cluster for $5 billion to $6 billion annually. This lease revenue offsets a $6 billion annual net loss.
The Colossus 2 project began on March 7, 2025. The facility includes 110,000 GB300 processors. This cluster provides 220 megawatts of compute power. The company acquired a 1 million square foot warehouse in Memphis. This site is 4x the power of the next-largest dedicated AI training site. Colossus 1 had only 11% GPU utilization because mixed GPU types created architectural issues. The project has 460 megawatts of installed or under-construction capacity in Southaven.
Grok model development and roadmap
Grok 3 uses 100,000 H100 GPUs at the Memphis data center. The model costs $22 per month for Premium+ subscribers. xAI will launch Grok 4.6 with 1.5 trillion parameters in August 2026. Grok 4.7 will arrive three to four weeks later with 2.1 trillion parameters. Grok 5 will ingest SpaceX engineering data by the end of 2026. The training of Grok 3 relies on 100,000 GPUs working day and night.
Grok 4.6 aims to maintain the cost-per-token profile of Grok 4.5. Grok 4.5 input pricing was $2 per million tokens and output pricing was $6 per million tokens. Grok 4.6 will compete with Kimi K3, which has 2.8 trillion parameters and a 1 million token context window. xAI’s combined valuation with SpaceX is $1.25 trillion.
Hardware throughput comparisons
The NVIDIA B200 delivers 2.2 times the FP8 throughput of the H200. You should know that the B200 has 192 GB of VRAM. The B200 also has 1.8 TB/s of NVLink bandwidth. This is more than double the 0.9 TB/s of the H200. The B200 has a 1,000W TDP. The H200 has a 700W TDP.
| GPU Model | VRAM | NVLink Bandwidth | TDP |
|---|---|---|---|
| H200 | 141 GB | 0.9 TB/s | 700W |
| B200 | 192 GB | 1.8 TB/s | 1,000W |
The B200 delivers 9,000 FP8 TFLOPS with sparsity. This is 2.27 times the 3,958 TFLOPS of the H200. The B200 provides the highest efficiency for large-scale model training. The B200 handles 405B parameter models more efficiently at high concurrency. High concurrency workloads justify the higher cost per GPU. The B200 reduces GPU count by 40 to 50 percent for 405B model serving. This reduction reduces total cluster cost by 15 to 25 percent.
| Facility | GPUs | Power |
|---|---|---|
| Colossus 1 | 230,000 | 500 MW |
| Colossus 2 | 550,000 | 1 GW |
| Total | 555,000+ | 2 GW |
The B200 handles MoE architectures and multi-model serving from a single GPU. The H200 handles sub-70B models and workloads with less than 60 percent utilization. The B200 delivers 18,000 TFLOPS at FP4 precision. Will the orbital compute strategy via Starlink satellites ever replace these massive terrestrial clusters?




