Microsoft’s deployment of the Maia 200 AI accelerator improves the economics of AI token generation in Azure. This second-generation chip provides 30% more performance per dollar than the Maia 100. The Maia 200 uses TSMC 3nm technology and contains 140 billion transistors. This release follows the earlier deployment of the Maia 100, which arrived in 2023. The Maia 200 targets inference tasks, specifically helping customers who require high-volume AI token generation. The Maia 200 provides a higher performance-per-dollar ratio than previous hardware, even though development delays and staffing issues slowed its arrival.
Technical specifications and architecture of the Maia 200
The Maia 200 architecture organizes into tiles, which act as the smallest autonomous units. Each tile contains a Tile Tensor Unit (TTU) for matrix multiply and convolution operations. This TTU supports mixed precision like FP8 activations with FP4 weights. Each tile also houses a Tile Vector Processor (TVP), which functions as a programmable SIMD engine. The TVP supports FP8 and also BF16, FP16, and FP32 formats. To prevent compute stalls, a Tile DMA moves data into and out of the Tile SRAM (TSRAM). A Tile Control Processor (TCP) orchestrates the issuance of work to the TTU and DMA. The Maia 200 memory subsystem targets the bottleneck of data movement by integrating a specialized DMA engine, on-die SRAM, and a specialized NoC fabric to maintain high bandwidth for massive models.
At the cluster level, multiple tiles share a Cluster SRAM (CSRAM) and use a Cluster DMA to stage traffic between the CSRAM and co-packaged HBM. This hierarchical model uses redundancy schemes for tiles and SRAM to improve yield. The network design uses an Ethernet-based approach with a custom Microsoft AI Transport Layer (ATL) protocol. This protocol supports packet spraying and congestion-resistant flow control. The chip also uses an intra-node topology with Fully Connected Quads (FCQ) between accelerator packages. This provides direct, non-switched links within a tray of four accelerators. Each accelerator exposes 2.8 TB/s of bidirectional, dedicated scaleup bandwidth.
| Metric | Azure Maia 100 | Azure Maia 200 |
|---|---|---|
| Process Technology | 5nm | 3nm |
| Transistor Count | 105 Billion | 140 Billion |
| FP4 Compute | N/A | >10 PFLOPS |
| FP8 Compute | N/A | >5 PFLOPS |
| HBM Memory | N/A | 216 GB HBM3e |
| HBM Bandwidth | N/A | 7 TB/s |
| On-die SRAM | N/A | 272 MB |
| TDP | 500W (in operation) | 750W |
Production delays and design challenges
Microsoft delayed mass production of the Braga chip, which becomes the Maia 200, until 2026. The company originally hoped to deploy the Braga chip into data centers in 2025. Design changes requested by OpenAI caused instability in simulations and pushed the project back several months. Staffing constraints and high turnover also hindered progress, as one-fifth of some chip design teams left the company. The company also scrapped plans to design a training-focused chip in early 2024.
The Maia 100 was designed for image processing rather than generative AI. Because of this, the Maia 100 does not power any of the company’s AI services. Instead, Microsoft used the Maia 100 internally for staff training purposes. This limitation led to concerns about whether the Braga chip can compete with Nvidia’s Blackwell offering. Will the Maia 300 design eventually overcome the competition with Nvidia’s Blackwell?
Liquid cooling and custom infrastructure
Microsoft uses a liquid cooling solution to manage heat in the Maia 200. This system includes a "sidekick" design that sits next to the server rack. Cold liquid flows from the sidekick to cold plates that attach to the surface of the Maia 100 and Maia 200 chips. Each plate contains channels through which liquid circulates to absorb and transport heat. This heat moves back to the sidekick, which removes the heat from the liquid and returns it to the rack.
The implementation of this cooling method required custom hardware. No existing racks held the unique requirements of the Maia 100 server boards. Consequently, Microsoft built wider racks from scratch to provide space for power and networking cables. The design of these racks and sidekicks allows Microsoft to control the entire infrastructure stack. You already know the basics of the AI hardware market, so focus on how Microsoft’s silicon strategy impacts Azure costs.
The AI at Work roadmap transition
The AI at Work roadmap replaced the Microsoft 365 Roadmap on August 25, 2026. Starting in September 2026, Dynamics 365, Power Platform, and Dataverse content also move to this single destination. This change replaces the old twice-yearly release wave model with continuous publishing. Microsoft publishes new capabilities as soon as the company commits to them. Users can filter roadmap views by product or environment.
The platform provides several tools for organizational planning. Users can export filtered views to a CSV file. The system also supports RSS subscriptions for updates. Organizations can use the Release Communications MCP Server to pull roadmap data into AI tools and workflows. This server pulls information directly into custom tools and workflows. This transition changes how Microsoft communicates innovation without changing how the company builds or deploys products.
Deployment and OpenAI collaboration
Microsoft deployed the Maia 200 in the US Central datacenter region near Des Moines, Iowa. Future deployments include the US West 3 datacenter region near Phoenix, Arizona. The Maia 200 serves multiple models, including the latest GPT-5.2 models from OpenAI. This deployment helps improve the economics of AI token generation for Microsoft’s customers.
OpenAI has collaborated with Microsoft since 2020 to co-design an Azure-hosted AI supercomputer. OpenAI provides feedback on the design of the Maia chip to help refine future hardware. This synergy powers multibillion-dollar workloads on Azure. The Maia 200 also helps the Microsoft Superintelligence team with synthetic data generation and reinforcement learning. This process helps improve next-generation in-house models.
The Azure silicon ecosystem and Cobalt 100
Microsoft aims to optimize every layer of the infrastructure stack to maximize performance and diversify its supply chain. The Azure Cobalt 100 CPU uses Arm architecture to optimize performance per watt in data centers. This 128-core chip follows an Arm Neoverse CSS design. The Cobalt 100 is optimized for cloud-native offerings and powers new virtual machines for customers.
The company maintains a heterogeneous deployment strategy. This strategy uses both in-house silicon and hardware from partners like Nvidia, AMD, Intel, and Qualcomm. Microsoft adds AMD MI300X accelerated VMs to Azure to provide more choice. The company also provides customers with the latest Nvidia H200 Tensor Core GPUs to support larger model inferencing. Microsoft tests GPT-4 on AMD chips to ensure performance across different hardware.
Competition in the AI accelerator market
The Maia 200 provides higher performance in several key metrics compared to other hyperscaler chips. It delivers over 10 petaFLOPS in 4-bit precision, which is three times the performance of the third-generation Amazon Trainium3. The Maia 200 also provides higher FP8 performance than the seventh-generation Google TPU. The 216GB of HBM3e memory on the Maia 200 provides 7 TB/s of bandwidth, which exceeds the 4.9 TB/s bandwidth of the Amazon Trainium3.
| Metric | Azure Maia 200 | AWS Trainium3 | Nvidia Blackwell B300 Ultra |
|---|---|---|---|
| Process Technology | 3nm | 3nm | 4nm |
| FP4 PetaFLOPS | 10.14 | 2.517 | 15 |
| FP8 PetaFLOPS | 5.072 | 2.517 | 2.5 |
| HBM Memory Size | 216 GB | 144 GB | 288 GB |
| HBM Bandwidth | 7 TB/s | 4.9 TB/s | 8 TB/s |
| TDP | 750 W | Unknown | 1400 W |
| Bi-directional Bandwidth | 2.8 TB/s | 2.56 TB/s | 1.8 TB/s |
Microsoft faces significant competition from Nvidia, which holds more than 70 percent of the AI chip market. The Blackwell B300 Ultra provides 15 petaFLOPS of FP4 compute, which is higher than the 10.14 petaFLOPS of the Maia 200. However, the Maia 200 operates at 750W, which is nearly half of the 1400W TDP of the Blackwell B300 Ultra. This efficiency helps Microsoft address environmental concerns and reduce the cost of running AI at scale.




