Nvidia’s acquisition and the Groq 3 LPU
Nvidia signed a $20 billion licensing agreement with Groq on December 24, 2025. This deal moved founder Jonathan Ross, President Sunny Madra, and 90% of the engineering team to Nvidia. Groq now seeks $300 million to $500 million in funding at a $6 billion valuation. This follows a $6.9 billion valuation in September 2025. Groq intends to use these funds to satisfy a $1.5 billion contract with Saudi Arabia, which should generate $500 million in revenue this year. The Groq 3 LPU, part of the Nvidia Vera Rubin platform, ships in the third quarter of 2026 at a target price of $45 per million tokens. Each LP30 chip contains 512MB of on-chip SRAM, which yields 150 TB/s of memory bandwidth per chip. The Groq 3 LPX rack houses 256 LPUs for 128GB of total SRAM and 40 PB/s of aggregate bandwidth. I find the concentration of Groq talent within Nvidia creates a significant barrier for independent LPU developers. The Groq 3 LPU focuses on the decode phase. It delivers 1,500 tokens per second throughput to support autonomous AI agents that require much faster communication than the 100 tokens per second humans require for interaction. Nvidia’s Groq 3 LPX, which focuses on the decode phase, delivers 35 times higher inference throughput per megawatt for trillion-parameter models than a standalone Blackwell NVL72 system can actually provide. The Groq 3 LPU is manufactured by Samsung on a 4nm process and sits on Nvidia’s MGX infrastructure.
| Component | Specification |
|---|---|
| Total LPUs (LP30 chips) | 256 |
| SRAM per LPU | 512 MB |
| Total SRAM | 128 GB |
| SRAM Bandwidth per LPU | 150 TB/s |
| Aggregate Rack Bandwidth | 40 PB/s |
| Fabric Links per LPU | 96 x 112 Gbps C2C |
| Compute per LPU | 1.23 FP8 PFLOPS |
| Manufacturing Process | Samsung 4nm |
| Cooling | Full Liquid Cooling |
| Platform | Nvidia MGX |
| Target Availability | Q3 2026 |
| Target Price/Performance | ~$45 per million tokens |
Llama 4 Maverick deployment costs
Meta released Llama 4 Maverick with 17 billion active parameters and 128 routed experts. The model has 400 billion total parameters. It requires 206GB of VRAM for deployment. You can run Llama 4 Scout, a 17 billion active parameter model, on a single NVIDIA H100 80GB GPU with INT4 quantization. Llama 4 Maverick requires two to four H100 GPUs or a full H100 DGX host. DeepInfra provides Llama 4 Maverick at $0.15 per million input tokens and $0.60 per million output tokens. AWS Bedrock charges $0.50 per million input tokens for the same model. This price difference shows that third-party providers undercut hyperscalers. I would skip the high costs of AWS Bedrock if you use DeepInfra instead. You already know the cost of training, so focus on inference. Llama 4 Maverick deployments on a single NVIDIA H100 DGX host cost between $8 and $16 per hour in cloud rental. Llama 4 Maverick provides better results than Gemini 2.0 Flash across several benchmarks. Llama 4 Scout has a 10M context window. The Llama Community License requires a separate commercial license if your monthly active users exceed 700 million. You cannot use Llama outputs to train non-Llama models. Llama 4 Behemoth, which is still training, has 288 billion active parameters and 16 experts. It outperforms GPT-4.5 and Claude Sonnet 3.7 on STEM benchmarks like MATH-500. Llama 4 Maverick provides high performance in image and text understanding. Llama 4 Maverick uses alternating dense and mixture-of-experts (MoE) layers for inference efficiency. This makes it more compute efficient for training and inference. Llama 4 Maverick achieves comparable results to the new DeepSeek v3 on reasoning and coding. Llama 4 Scout is more powerful than all previous generation Llama models.
Decoding the inference market reality
The Groq 3 LPU uses a deterministic VLIW architecture. This design eliminates branch prediction and out-of-order execution to favor sequential token generation. Groq’s architecture avoids the cache misses seen in traditional designs through high bandwidth on-chip SRAMs. The LPU reaches 394 to 1,000 tokens per second depending on the model size. A trillion-parameter model requires roughly 70 racks of SRAM to store the weights. This does not include the memory needed for KV cache growth. A TB-class model requires enormous physical space and increases the total cost of ownership. Does the speed advantage of the LPU outweigh the massive hardware footprint required for large-scale deployments? I see that the market rewards speed, yet the memory density of LPUs remains a problem for massive models. Groq’s architecture is a tiled architecture where interleaved vertical columns of SRAM and compute operate as a unified, synchronous grid. The on-chip SRAM has 230MB capacity. For Llama 3.1 8B, input tokens cost $0.05 per million and output tokens cost $0.08 per million. For Llama 3.3 70B, input tokens cost $0.59 per million and output tokens cost $0.79 per million.




