Cerebras Systems maintains a valuation of roughly $50.5 billion as of September 2026. This figure follows a massive climb from its $4 billion valuation in November 2021. The company reported $510 million in revenue for 2025, which grew 76% from the $290.3 million it earned in 2024. This represents a massive leap from the $78.7 million it recorded in 2023 and the $25 million it earned in 2022. Management expects to recognize 15% of its $24.6 billion revenue backlog during 2026 and 2027. The company faces intense customer concentration, as two clients, G42 and MBZUAI, accounted for 86% of all sales in 2025. G42 contributed 24% of those sales, while MBZUAI provided 62%. The company also reported a net income of $237.8 million in 2025, but this total includes a $363.3 million one-time gain from extinguishing a forward contract liability. If you remove that accounting event, the company lost $75.7 million on a non-GAAP basis. The operating loss also widened from $101.4 million in 2024 to $145.9 million in 2025. In 2024, the company saw $640.3 million in customer prepayments from G42. Investors also eye the massive OpenAI agreement, which involves providing 750 megawatts of computing power through 2028 in a deal valued at more than $10 billion. OpenAI also lent Cerebras $1 billion in exchange for 33 million warrants worth roughly $10 billion. At the top end of its IPO range, Cerebras implied a $34.4 billion market cap. The company raised $5.55 billion during its debut.
SambaNova Systems holds an $11 billion valuation after completing a $1 billion Series F round in July 2026. This capital infusion aims to scale business operations and secure supply chain capacity for high demand. The SN40L Reconfigurable Dataflow Unit (RDU) is a compact solution for data centers compared to larger architectures. SambaNova achieves 10X better performance per area than Cerebras on the Llama 3.1 70B model. Its 70B configuration uses 16 chips to deliver 457 output tokens per second per user. This performance exceeds Cerebras, which delivers 445 tokens per second for the same model. On the largest Llama 3.1 model, the 405B version, SambaNova holds a world record with 129 output tokens per second per user. On the smallest 8B model, Cerebras delivers 1837 tokens per second, which beats the 1042 tokens per second produced by SambaNova. SambaNova also provides up to 40X better performance per area than Groq on the Llama 3.1 70B model. JPMorgan Chase already selected the SN40L and SN50 systems for its on-premises AI inference workloads. The SN40L rack includes a standard 19-inch form factor with air cooling and delivers full utilization of its 10.2 PFLOPs of performance. This configuration reduces machine footprint by up to 19X for Composition of Experts deployments compared to a DGX H100. You should look closely at how these companies scale. The SN40L allows for faster model switching, providing speedups between 15X and 31X.
| SN40L Component | Specification |
|---|---|
| Compute Performance | 10.2 bf16 PFLOPS |
| On-chip SRAM | 520 MiB |
| On-package HBM | 64 GiB |
| Off-package DDR DRAM | Up to 1.5 TiB |
| Die Technology | TSMC 5nm |
The industry focuses heavily on the transition from model training to the inference stage. Cerebras uses wafer-scale engines that contain 4 trillion transistors and 900,000 cores. This single piece of silicon provides 44 gigabytes of memory to run models directly. Cerebras claims its CS-3 system works up to 21 times faster than Nvidia’s B200 for certain inference jobs. SambaNova uses a different approach with its RDU architecture. It combines 520 MiB of on-chip SRAM, 64 GiB of HBM, and up to 1.5 TiB of DDR DRAM. This three-tier memory system helps overcome the memory wall that limits many other accelerators. The SN40L provides 10.2 bf16 PFLOPS of peak compute performance. For Composition of Experts inference deployments, the SN40L Node reduces machine footprint by up to 19X. It also achieves an overall speedup of 3.7X over a DGX H100 and 6.6X over a DGX A100. The SN40L utilizes 1040 distributed Pattern Compute Units (PCUs) and 1040 distributed Pattern Memory Units (PMUs). These components provide hundreds of TBps of on-chip memory bandwidth and high bank-level parallelism. The RDU maps each model to optimize performance using nested combinations of data, tensor, and pipeline parallelisms within each chip and across chips. While Cerebras builds massive single chips, SambaNova uses 16 chips interconnected with a peer-to-peer network. Will the industry ultimately favor the massive scale of a single wafer or the flexibility of multi-chip systems?




