Groq scaling through Nvidia while Llama 4 shifts deployment economics

Groq targets a $6 billion valuation following a $20 billion Nvidia licensing deal, while Meta’s Llama 4 Maverick offers competitive inference costs through providers like DeepInfra at $0.60 per million output tokens.

Groq scaling through Nvidia while Llama 4 shifts deployment economics

Nvidia’s acquisition and the Groq 3 LPU

Nvidia signed a $20 billion licensing agreement with Groq on December 24, 2025. This deal moved founder Jonathan Ross, President Sunny Madra, and 90% of the engineering team to Nvidia. Groq now seeks $300 million to $500 million in funding at a $6 billion valuation. This follows a $6.9 billion valuation in September 2025. Groq intends to use these funds to satisfy a $1.5 billion contract with Saudi Arabia, which should generate $500 million in revenue this year. The Groq 3 LPU, part of the Nvidia Vera Rubin platform, ships in the third quarter of 2026 at a target price of $45 per million tokens. Each LP30 chip contains 512MB of on-chip SRAM, which yields 150 TB/s of memory bandwidth per chip. The Groq 3 LPX rack houses 256 LPUs for 128GB of total SRAM and 40 PB/s of aggregate bandwidth. I find the concentration of Groq talent within Nvidia creates a significant barrier for independent LPU developers. The Groq 3 LPU focuses on the decode phase. It delivers 1,500 tokens per second throughput to support autonomous AI agents that require much faster communication than the 100 tokens per second humans require for interaction. Nvidia’s Groq 3 LPX, which focuses on the decode phase, delivers 35 times higher inference throughput per megawatt for trillion-parameter models than a standalone Blackwell NVL72 system can actually provide. The Groq 3 LPU is manufactured by Samsung on a 4nm process and sits on Nvidia’s MGX infrastructure.

Component Specification
Total LPUs (LP30 chips) 256
SRAM per LPU 512 MB
Total SRAM 128 GB
SRAM Bandwidth per LPU 150 TB/s
Aggregate Rack Bandwidth 40 PB/s
Fabric Links per LPU 96 x 112 Gbps C2C
Compute per LPU 1.23 FP8 PFLOPS
Manufacturing Process Samsung 4nm
Cooling Full Liquid Cooling
Platform Nvidia MGX
Target Availability Q3 2026
Target Price/Performance ~$45 per million tokens

Llama 4 Maverick deployment costs

Meta released Llama 4 Maverick with 17 billion active parameters and 128 routed experts. The model has 400 billion total parameters. It requires 206GB of VRAM for deployment. You can run Llama 4 Scout, a 17 billion active parameter model, on a single NVIDIA H100 80GB GPU with INT4 quantization. Llama 4 Maverick requires two to four H100 GPUs or a full H100 DGX host. DeepInfra provides Llama 4 Maverick at $0.15 per million input tokens and $0.60 per million output tokens. AWS Bedrock charges $0.50 per million input tokens for the same model. This price difference shows that third-party providers undercut hyperscalers. I would skip the high costs of AWS Bedrock if you use DeepInfra instead. You already know the cost of training, so focus on inference. Llama 4 Maverick deployments on a single NVIDIA H100 DGX host cost between $8 and $16 per hour in cloud rental. Llama 4 Maverick provides better results than Gemini 2.0 Flash across several benchmarks. Llama 4 Scout has a 10M context window. The Llama Community License requires a separate commercial license if your monthly active users exceed 700 million. You cannot use Llama outputs to train non-Llama models. Llama 4 Behemoth, which is still training, has 288 billion active parameters and 16 experts. It outperforms GPT-4.5 and Claude Sonnet 3.7 on STEM benchmarks like MATH-500. Llama 4 Maverick provides high performance in image and text understanding. Llama 4 Maverick uses alternating dense and mixture-of-experts (MoE) layers for inference efficiency. This makes it more compute efficient for training and inference. Llama 4 Maverick achieves comparable results to the new DeepSeek v3 on reasoning and coding. Llama 4 Scout is more powerful than all previous generation Llama models.

Decoding the inference market reality

The Groq 3 LPU uses a deterministic VLIW architecture. This design eliminates branch prediction and out-of-order execution to favor sequential token generation. Groq’s architecture avoids the cache misses seen in traditional designs through high bandwidth on-chip SRAMs. The LPU reaches 394 to 1,000 tokens per second depending on the model size. A trillion-parameter model requires roughly 70 racks of SRAM to store the weights. This does not include the memory needed for KV cache growth. A TB-class model requires enormous physical space and increases the total cost of ownership. Does the speed advantage of the LPU outweigh the massive hardware footprint required for large-scale deployments? I see that the market rewards speed, yet the memory density of LPUs remains a problem for massive models. Groq’s architecture is a tiled architecture where interleaved vertical columns of SRAM and compute operate as a unified, synchronous grid. The on-chip SRAM has 230MB capacity. For Llama 3.1 8B, input tokens cost $0.05 per million and output tokens cost $0.08 per million. For Llama 3.3 70B, input tokens cost $0.59 per million and output tokens cost $0.79 per million.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories