xAI Colossus vs Microsoft Stargate: AI supercomputer efficiency

xAI’s Colossus cluster in Memphis delivers better training cost per watt than the Stargate project due to its 2 gigawatt capacity and high-density liquid cooling. The comparison examines hardware density, power sources, and the economic impact of on-premise…

xAI Colossus vs Microsoft Stargate: AI supercomputer efficiency

Massive Compute Concentrations

xAI’s Colossus complex in Memphis provides 2 gigawatts of capacity. This facility houses 555,000 NVIDIA GPUs. The total investment in these GPUs reached $18 billion. The Stargate Project, led by OpenAI, SoftBank, and Oracle, targets a $500 billion total investment. The Stargate Abilene site in Texas expects 1.2 gigawatts of power. This site expects to deploy 64,000 B200 GPUs by 2026. While xAI repurposes industrial warehouses in Memphis, the Stargate Project targets distributed infrastructure across several states. The 1.5GW data center campus in Neom’s Oxagon will be built by G42 and operated in partnership with several US companies. The first 300MW phase of that site will go live in 2028.

Colossus Hardware Density and Thermal Management

Colossus 2 uses GB200 and GB300 processors. The first batch of 550,000 GB200 and GB300 processors went live at Colossus 2 in July 2025, following a build process that compressed traditional development timelines into a few short months. The massive $18 billion investment in 555,000 NVIDIA GPUs allows xAI to train Grok models with massive scale, as the single-site deployment exceeds the capacity of every other AI training facility currently in operation throughout the world. The site generates significant heat from 2 gigawatts of compute. xAI uses liquid cooling for these chips. Specifically, 2 gigawatts of GPU compute generates approximately 1.8 gigawatts of heat requiring dissipation. The facility requires a cooling capacity of over 50,000 gallons per minute. To manage this, xAI utilizes Supermicro liquid-cooled racks.

Stargate’s Energy and Distributed Buildout

Stargate utilizes SB Energy to develop solar and energy storage infrastructure. The agreement includes a $1 billion equity investment to fund these renewable assets. In Milam County, Texas, OpenAI selected SB Energy to build a 1.2 gigawatt data center site. The project focuses on grid modernization and the mitigation of solar intermittency. Stargate also relies on natural gas turbines for the Abilene site. The facility in Abilene expects to deliver 1.2 gigawatts of power. This project involves eight buildings, and each building can house up to 50,000 GPUs. The Stargate project involves a partnership between SoftBank, OpenAI, Oracle, and MGX. SoftBank and OpenAI lead the project, with SoftBank holding financial responsibility and OpenAI holding operational responsibility. Masayoshi Son serves as chairman.

Comparison of Hardware and Power

Feature xAI Colossus (Memphis) Stargate (Abilene/Phase 1)
Total Power 2 GW 1.2 GW
GPU Count 555,000 64,000
Primary GPU GB200/GB300 B200/GB200/GB300
Power Source Gas turbines/Megapacks Natural gas/Solar/Wind
Cooling Liquid Dry

The Colossus cluster in Memphis reaches 2 gigawatts of total power. This scale exceeds the 1.2 gigawatt capacity planned for the Stargate Abilene site. xAI’s deployment of 555,000 GPUs includes a mix of GB200 and GB300 models. The Stargate project plans for 64,000 B200 GPUs by 2026 at the Abilene location.

Economic Efficiency of On-Premise Ownership

Training costs fluctuate based on whether a company rents cloud capacity or owns hardware. Owning B300 infrastructure saves $4,746,771.5 per server over a 5-year lifecycle compared to the hourly cloud rate for an AWS p6-b300.48xlarge. The total five-year cost for an on-premise Lenovo Config D system with 8x B300 GPUs reaches $1,505,678.5. This total includes a CapEx of $785,606.5 and an hourly OpEx of $16.44. The Lenovo Config B system with 8x H200 GPUs has a CapEx of $397,801.6. The OpEx for that system is $9.80 per hour, which includes $5.45 for maintenance, $2.27 for power and cooling, and $2.08 for colocation. Against Azure’s on-demand rate of $114.656 per hour, the system reaches a breakeven point in 3,793 hours, or about 5.2 months. Against a 1-year reserved rate of $73.39 per hour, the breakeven takes 6,250 hours, or roughly 8.5 months. Against a 3-year reserved rate of $50.33 per hour, the Lenovo system pays for itself in 9,800 hours, or 13.4 months. Against a 5-year reserved rate of $46.56 per hour, the breakeven occurs in 10,800 hours, or 14.8 months, after which the user saves money for 45 months. You know the inefficiency of legacy data centers, so look at the numbers for Colossus. For B200 systems, the Lenovo SR680a V3 reaches a breakeven point against Google Cloud in just 5.3 hours of daily utilization.

The AI Factory Model and Vertical Integration

xAI treats Colossus as a Gigafactory of Compute. This model focuses on industrializing the production of intelligence. The company repurposes existing industrial buildings, such as the 785,000 square foot Electrolux factory, to accelerate deployment. The buildout relies on direct-to-chip liquid cooling to manage the 700W burn of each of the 200,000 GPUs. This strategy mirrors the production models used in automotive manufacturing. The company builds its own power generation on-site rather than waiting for utility interconnection. The deployment of 168 Tesla Megapacks helps stabilize power spikes at the Memphis site, providing a buffer while the company constructs a permanent 1.2 gigawatt simple-cycle gas plant to replace temporary units. This model allows xAI to bypass the years-long wait times for traditional utility grid upgrades. The system runs on NVIDIA’s Spectrum-X Ethernet platform, which provides 400Gbps per GPU and 400Gbps per CPU server, alongside BlueField-3 SuperNICs for RDMA at scale.

Power Generation and Environmental Friction

The rapid deployment of Colossus relied on temporary gas turbines. These units caused significant environmental concerns in Memphis, specifically regarding air quality, NOx, and formaldehyde. Local residents and environmental groups criticize the lack of public notice and the potential health impacts of the gas turbines. The operation of gas turbines in Black neighborhoods triggered federal lawsuits and significant local opposition. In April 2026, the NAACP sued xAI in federal court, alleging the operation of the Southaven gas plant without required Clean Air Act permits. The suit also alleged a disproportionate impact on nearby Black communities. The U.S. Department of Justice moved to intervene in the case in June 2026. Will the environmental litigation in Mississippi halt the expansion of the Memphis clusters?

The Verdict on Training Cost per Watt

xAI’s Colossus delivers better training cost per watt than the Stargate Phase 1 plan. The concentration of 555,000 GPUs in a single 2-gigawatt site allows for extreme density. This density facilitates liquid cooling and direct-to-chip thermal management, which minimizes energy waste during massive training runs. The $18 billion investment in 555,000 NVIDIA GPUs allows xAI to train Grok models with massive scale, as the single-site deployment exceeds the capacity of every other AI training facility currently in operation throughout the world. The vertical integration of on-site gas turbines and Tesla Megapacks reduces the reliance on expensive, intermittent grid connections. Stargate’s distributed model requires more extensive energy coordination across different sites and renewable sources. The ability of xAI to control the physical substrate of intelligence – GPUs, electrons, and cooling water – through a single, high-density site provides a clear advantage in energy efficiency.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories