Massive compute scaling via Condor Galaxy
Cerebras and G42 deployed Condor Galaxy 3 to provide 8 exaflops of AI compute through 64 Cerebras CS-3 systems. Each CS-3 system uses the Wafer Scale Engine 3 chip, which packs 4 trillion transistors and 900,000 AI-optimized cores. This hardware supports the training of neural network models up to 24 trillion parameters. The Condor Galaxy network targets a total capacity of 36 exaflops. I find the scaling of these systems far simpler than traditional GPU clusters. While many companies build massive GPU clusters that cost billions, distributing a single model over thousands of tiny GPUs often takes months of work from dozens of experts. The WSE-3 chip achieves 125 petaflops of peak AI performance and 21 petabytes of memory bandwidth. This performance follows a progression of wafer-scale technology. The first WSE product in 2019 used 14nm technology and 400,000 cores to reach 9 petabytes per second of bandwidth. The second version in 2021 moved to 7nm and 850,000 cores to reach 20 petabytes per second of bandwidth. The Condor Galaxy 1, using the older WSE-2, provides 4 exaflops of compute and 54 million cores, fed by 72,704 AMD EPYC processor cores.
National lab budget realities
The Department of Energy fiscal year 2027 budget request creates a difficult environment for national lab research because it proposes a $1.26 billion decrease for the Office of Science compared to the $8.40 billion approved for the previous year. The proposal asks for $7.14 billion for the Office of Science, a 15 percent decrease from the $8.40 billion approved for fiscal year 2026. This $1.26 billion cut impacts the Genesis Mission. The Advanced Scientific Computing Research program would receive $1.10 billion, a $20 million reduction from the previous year. The administration also proposes a new Office of Artificial Intelligence and Quantum with $1.2 billion to coordinate AI and quantum research. Does this new office provide enough support to offset the massive cuts to basic science research? I suspect the refocus on infrastructure development may not address the immediate needs of researchers facing these shrinking budgets. The Energy Sciences Coalition urges Congress to reject these cuts and instead appropriate $9.5 billion to maintain U.S. leadership. Within the Office of Science, the Biological and Environmental Research program faces a 54 percent cut, while Basic Energy Sciences faces a 20 percent reduction. ARPA-E funding drops 43 percent to $200 million from the previous $350 million. These reductions hit the very programs that drive the innovation required for national security and energy dominance.
Economics of the WSE-3 architecture
The WSE-3 architecture reduces the costs that typically plague AI scaling. The CS-3 delivers 32 percent lower cost per token and one third lower energy use than the Nvidia DGX B200 according to data from SemiAnalysis. The WSE-3 also delivers 21 times faster inference on models such as Llama 3 70B than the Blackwell architecture. I observe that the ability to perform deeper chain-of-thought iterations in the same response time gives Cerebras an edge in reasoning-intensive workloads. This capacity helps solve problems in healthcare, energy, and climate action. The CS-3 can train a one-trillion parameter model as easily as a one-billion parameter model trained on GPUs. While Nvidia’s B200 provides 20 PetaFLOPS of FP4 AI performance and 8 TB/s of memory bandwidth, the Cerebras system avoids the interchip communication bottlenecks found in multichip architectures. The WSE-3 uses 5nm process technology to achieve its density. The H100, a leading GPU, uses 80 billion transistors and 80 GB of HBM3 memory with 3 TB/s of bandwidth. I would skip the massive GPU clusters if you could achieve 2,700 tokens/s with the Cerebras system instead of 900 tokens/s with Blackwell. The Condor Galaxy network has already trained models like Jais-30B and Med42, which surpassed MedPaLM in accuracy.




