The Colossus infrastructure and the reality of Grok training

xAI is expanding its Colossus GPU cluster toward 200,000 units using H100 and H200 chips. While the hardware scales massively, Grok 4.5 faces competition from frontier models despite its 1.5 trillion parameter foundation and aggressive pricing.

The Colossus infrastructure and the reality of Grok training

xAI operates Colossus, a GPU training cluster in South Memphis, Tennessee, and Southaven, Mississippi. This facility uses 100,000 NVIDIA H100 GPUs connected through NVIDIA Spectrum-X Ethernet networking. The team built this high-performance system in 122 days. Each rack in the cluster contains eight 4U servers. Each of these servers contains eight NVIDIA H100 GPUs, which results in 64 GPUs per rack. These GPU servers, along with cooling distribution units and networking components, form mini-clusters of 512 GPUs. The configuration includes over 1,500 racks across nearly 200 arrays. This setup aims to train the Grok family of models, which need massive processing power for large-scale training.

The Hardware Scale of Colossus

The expansion of Colossus moves toward a 200,000 GPU cluster. This second phase adds 50,000 NVIDIA H100 GPUs and 50,000 next-generation H200 GPUs. These additions push the power consumption beyond the capacity of the 14 diesel generators installed earlier this year. The cluster architecture relies on the Spectrum-X platform, which uses Spectrum SN5600 Ethernet switches. These switches support 800Gbps port speeds and run on the Spectrum-4 ASIC. BlueField-3 DPUs offload network tasks from CPUs to prevent overloading the GPUs. Each H100 server includes a dedicated 400Gb NIC and an additional NIC. This setup achieves a combined bandwidth of 3.6 Tbps per server.

The cluster uses advanced liquid cooling to manage heat. This system improves heat management and energy efficiency compared to traditional air cooling. However, the reliance on evaporative cooling introduces concerns regarding local water availability. The reported 250-megawatt figure requires careful scrutiny because a simple calculation multiplying the 100,000 H100 GPUs by their 700-watt thermal design power produces a significantly lower total than the capacity mentioned by industry analysts. The original Colossus 1 data center consumes 150 megawatts of electricity at full capacity.

Power Constraints and the Methane Turbine Ruling

The Environmental Protection Agency ruled that the portable methane gas turbines used by xAI are not exempt from air quality requirements. xAI used up to 35 unpermitted turbines at the Memphis site. The company claimed these turbines were temporary, but the EPA policy requires air permits even for portable or temporary machines. This ruling follows a lawsuit from the NAACP and the Southern Environmental Law Center. Methane gas turbines emit nitrogen oxides, which link to asthma and respiratory diseases. The EPA estimates that enforcing these standards will reduce nitrogen oxide emissions by 296 tons annually by 2032.

The construction of Colossus 2 began in March 2025. This expansion uses dozens of gas turbines. xAI developed a gigawatt-scale energy hub in Mississippi to manage the pushback from Memphis. The state granted xAI temporary approval to run gas turbines without a permit for up to 12 months. xAI uses rental turbines from Solaris Energy Infrastructure. The total power needs for the expansion could reach 2 gigawatts. Will the power grid accommodate this gigawatt-scale expansion?

Grok Model Capabilities and Pricing

xAI has released several iterations of the Grok model, each with different context windows and parameter counts. The models vary in their intended use, such as reasoning, coding, or general intelligence.

Model Context Window Parameters Primary Focus
Grok 3 1,000,000 tokens Not disclosed Reasoning and Mathematics
Grok 4 256,000 tokens Not disclosed General Intelligence
Grok 4.1 Not disclosed Not disclosed Agentic Reasoning
Grok 4.3 Not disclosed Not disclosed Scale and Efficiency
Grok 4.5 500,000 tokens 1.5 Trillion Coding and Agents

Grok 4.5 is the first model built specifically for coding and agentic work. It uses a 1.5-trillion-parameter foundation. The pricing for Grok 4.5 remains low compared to competitors. The standard rate is $2 per million input tokens and $6 per million output tokens. Users can use cached input tokens for $0.50 per million, which is a 75% discount. Tool calls involve additional fees. Web Search, X Search, and Code Execution cost $5 per 1,000 calls. File attachments cost $10 per 1,000 calls. Collections search costs $2.50 per 1,000 invocations.

Training Performance and Reasoning Modes

Grok 3 utilizes a "Think" mode to spend seconds or minutes reasoning before it answers. At its highest test-time compute setting, the model reached 93.3% on the 2025 American Invitational Mathematics Examination. This model uses a self-consistency method with 64 attempts per problem. Grok 4.1 introduced a two-mode design with a reasoning variant called quasarflux and a non-reasoning variant called tensor.

Grok 4.5 provides a compelling option for high-volume, low-reasoning tasks, but it fails to compete with frontier models on complex agentic workflows. The model uses three reasoning-effort settings: low, medium, and high. High effort is the default setting. Dropping to medium or low effort reduces the number of reasoning tokens, which lowers the effective output cost. The model reaches a score of 54 on the Artificial Analysis Intelligence Index. This score follows a 16-point jump over Grok 4.3.

The Intelligence Gap in Benchmarks

Grok 4.3 trails OpenAI models, Anthropic models, Google models, Kimi, and MIMO on the Artificial Analysis composite intelligence ranking. This gap is not a rounding error. OpenAI models hold the top positions, followed by Anthropic and then Google models. Grok 4.5 also stays behind the frontier in certain areas. On the SWE-Bench Pro benchmark, Grok 4.5 scores 64.7%, while Claude Opus 4.6 scores 69.2%.

The cost advantage of Grok 4.5 comes from a different cost structure. xAI uses roughly 11% of its available compute for Grok models. This means xAI has a large amount of idle infrastructure. This excess compute allows for aggressive pricing to increase utilization. You should focus on tokens per second per dollar when evaluating these models. For tasks like classification or structured extraction, a lower reasoning score does not prevent high accuracy.

The Fallacy of GPU Utilization Metrics

Monitoring tools often provide misleading data regarding how well a cluster operates. nvidia-smi reports the percentage of time during the last sample period that one or more kernels executed on the GPU. This metric stays binary: a kernel is running or it is not. This metric says nothing about how efficiently the kernel uses the GPU’s compute units. A kernel that launches, performs a small memory copy, and waits for data reports 99% utilization while the shader processors sit idle for most of that execution. This makes the metric useless for modern deep learning workloads.

The metrics that matter include SM active cycles and memory bandwidth utilization. Memory-bandwidth-bound workloads, such as LLM inference, typically achieve 40% to 70% of peak memory bandwidth. If a user sees a value below 30%, a configuration problem exists. Small batch sizes cause the GPU to spend time waiting for memory rather than computing. Increasing the batch size keeps the compute units busy during memory loads.

GPU Model Architecture FP16 Performance Memory Bandwidth
H100 Hopper 1,671 TFLOPS 80 GB 3.9 TB/s
A100 Ampere 312 TFLOPS 80 GB 2,039 GB/s
L40S Ada Lovelace 731 TFLOPS 48 GB 864 GB/s

Scaling Economics and the Megawatt Metric

OpenAI CFO Sarah Friar suggested a correlation between compute power and revenue growth. She noted that OpenAI’s compute grew 9.5 times from 2023 to 2025. This suggests a relationship where $1GW of compute generates $10B in annual recurring revenue. This Megawatt-to-Revenue metric treats AI as an industrial commodity. However, this ignores the difference between training compute and inference compute. Training is the upfront capital expense for foundation models. Inference is the operational expense for serving users.

xAI’s Colossus 2 aims for a gigawatt capacity. The company can use this capacity for training Grok or for external contracts. Reflection, an open-source startup, signed an agreement to use Colossus 2 hardware. The arrangement includes Nvidia GB300 processors. Anthropic and Google also expect to use Musk-controlled computing. The ability to monetize idle capacity through third-party deals helps offset the massive cost of the hardware.

Execution Risk and the Infrastructure Bottleneck

The primary tension in the AI industry exists between the scale promises of executives and the physical work of deployment. Nvidia can ship processors and networking products. xAI must turn those components into a reliable computing service. A delay in a transformer, switchgear package, or cooling installation prevents servers from running.

The demand for processors like the H100 and B200 remains high. This demand puts pressure on the supply chain for HBM and CoWoS packaging. TSMC doubles its CoWoS capacity in 2026, but the capacity remains tight. Data center construction involves utility coordination and environmental permits. These processes move on a different schedule than chip deliveries. The success of Colossus 2 depends on the ability to manage power, cooling, and networking at an unprecedented scale.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories