Groq LPU inference pricing edge over NVIDIA H100s

Groq provides a cost advantage for Llama 3.3 70B workloads compared to NVIDIA H100s when using highly optimized hardware. While Cerebras offers higher raw throughput, Groq remains the cheaper option per million tokens for specific model deployments.

Groq LPU inference pricing edge over NVIDIA H100s

Groq beats NVIDIA on cost per million tokens for Llama 3.3 70B when users run highly optimized, high-utilization workloads. Cerebras delivers higher raw throughput by 3x to 8x than Groq, but Groq is the cheaper option per million tokens across tracked models. This comparison evaluates the LPU and WSE architectures against NVIDIA H100 GPUs.

LPU and WSE architectures

Groq LPU carries 230MB of on-chip SRAM and draws 375W per chip. It runs natively at 8-bit precision, while 16-bit inference runs significantly slower. A single LPU cannot hold the weights for a 70B model, so Groq assembles hundreds of chips into an interconnected fabric. Cerebras uses the WSE-3, which holds 44GB of on-die SRAM and 900,000 AI cores on a single silicon wafer. The WSE-3 connects cores via an on-chip fabric running at 21 PB/s and supports 16-bit precision natively. NVIDIA H100 variants use 80GB of HBM3 memory.

Groq spreads workloads across many small, low-power chips wired together. Cerebras puts the entire workload on one enormous chip with no chip-to-chip hop at all. Cerebras claims roughly 3x higher compute-per-watt than an 8-GPU DGX system running the same class of workload. This advantage matters for users paying the power bill on a dedicated deployment rather than a per-token API call.

The decision between paying per token on Groq or managing a dedicated cluster of NVIDIA H100 GPUs depends on whether your workload prioritizes lowest possible latency or lowest possible cost per token. Groq’s LPU eliminates branch prediction, caches, and out-of-order execution. These features are necessary for GPUs in general-purpose workloads but add latency for the sequential token generation LLMs need.

Throughput and latency benchmarks

Cerebras is faster on every published benchmark. On Llama 3.3 70B, Cerebras puts Groq at 403 to 750 tokens per second against Cerebras’s 2,100 to 2,500 tokens per second. This represents a 3x to 6x spread. Independent benchmarking firm Artificial Analysis measured Groq lower at 293.6 tokens per second on the same model. This measurement widens the gap to 7x or 8x against Cerebras. On GPT-OSS 120B, Cerebras claims roughly 3,000 tokens per second, while Artificial Analysis measured a sustained 1,700 tokens per second. Groq lands at roughly 478 to 493 tokens per second on that same model.

Groq’s throughput stands up against most of the market. Artificial Analysis ranks Groq as the fastest of 14 benchmarked providers serving Llama 3.3 70B, placing it just ahead of SambaNova at 288.6 tokens per second. Cerebras CEO Andrew Feldman stated at the RAISE Summit 2026 that Cerebras is the fastest inference in the industry. His claim is measured against GPUs, not Groq.

You know that latency matters for real-time agents, but you must also weigh the cost of the hardware running those agents. Groq’s SRAM-resident weights eliminate the largest source of latency in transformer inference. The deterministic interconnect minimizes the variance.

Model availability and constraints

Neither Groq nor Cerebras is a general-purpose inference layer. Both platforms maintain narrow, curated catalogs. Groq’s non-enterprise catalog includes:

Model Input ($ / 1M) Output ($ / 1M)
Llama 3.1 8B Instant $0.05 $0.08
GPT-OSS 20B $0.075 $0.30
GPT-OSS 120B $0.15 $0.60
Llama 3.3 70B Versatile $0.59 $0.79
Qwen 3.6 27B $0.60 $3.00

Cerebras’s Inference API supports Llama 3.1 8B, Llama 3.1 70B, Llama 3.3 70B, Qwen 3-32B, Qwen 3-235B Instruct, Qwen 3-235B Thinking, GPT-OSS-120B, and ZAI GLM-4.7. Both platforms exclude DeepSeek, Gemma, Mistral, and all multimodal models. Neither platform offers self-serve fine-tuning. Groq’s LoRA fine-tune support is enterprise-only and requires a request process.

Model Input ($ / 1M) Output ($ / 1M)
Whisper Large v3 Turbo $0.04 (per hour) N/A
Canopy Labs Orpheus English $22.00 (per 1M chars) $22.00 (per 1M chars)

Groq’s speech models use different billing units. Whisper Large v3 Turbo costs $0.04 per hour of audio transcribed. Canopy Labs Orpheus English costs $22.00 per million characters.

Groq API pricing structures

Groq pricing follows a pay-as-you-go, tokens-as-a-service model. Input token prices for Llama 3.1 8B range from $0.05 per million tokens. Output tokens for the same model cost $0.08 per million. Users get a free tier with no credit card requirement. This tier limits users to 30 requests per minute and 6,000 tokens per minute.

The Developer tier allows users to add a credit card to unlock 10x the free tier rate limits and a 25% discount on all token costs. Groq also provides a Batch API and prompt caching. Both mechanisms cut rates by 50%. The Batch API processes asynchronous jobs within 24 hours to seven days. Prompt caching applies a 50% discount to repeated input tokens.

Service Tier Rate Limit Cost
Free 30 RPM / 6,000 TPM $0.00
Developer 300 RPM / 60,000 TPM 25% discount
Enterprise Custom Custom

The Batch API provides 50% off every rate. For Llama 3.3 70B, rates drop from $0.59/$0.79 to $0.30/$0.40 per million tokens. Prompt caching for GPT-OSS 120B drops the input rate from $0.15 to $0.075 per million tokens.

NVIDIA H100 rental and purchase costs

Renting NVIDIA H100 GPUs in 2026 costs between $2.00 and $10.00 per GPU-hour. GMI Cloud offers H100 PCIe at $2.00 per hour and H100 SXM at $2.40 per hour. Jarvislabs lists H100 SXM at $2.69 per hour. Hyperscalers like AWS and Google Cloud charge between $4.00 and $8.00 per GPU-hour. These providers often add fees for storage and data egress.

Buying hardware requires high upfront capital. A single NVIDIA H100 SXM costs between $35,000 and $40,000. An 8-GPU server with H100 SXM costs between $280,000 and $320,000 for the GPUs.

H100 Configuration Rental Price (per hour) Purchase Price (per GPU)
H100 PCIe (GMI) $2.00 $25,000 – $30,000
H100 SXM (Jarvislabs) $2.69 $35,000 – $40,000
H100 SXM (Lambda) $3.99 $35,000 – $40,000
8-GPU SXM Cluster $21.52 $280,000 – $320,000

Buying becomes cost-effective only if users run GPUs at consistently high utilization for multiple years. Renting avoids the large capital costs and procurement lead times of 6 to 12 months. For a 1,000 GPU-hour monthly workload, renting via GMI Cloud costs $2,000 per month. An on-premises purchase for that same workload requires hundreds of thousands of dollars in hardware and infrastructure.

The cost crossover math

The crossover point between Groq and NVIDIA exists where highly optimized GPU utilization meets Groq’s base pricing. Groq Llama 3.3 70B pricing hits $0.59 for input and $0.79 for output per million tokens. A typical 80/20 input-heavy split results in a blended rate of $0.63 to $0.67 per million tokens. A self-hosted deployment on rented H100s using vLLM can reach $1.24 per million tokens at 40% utilization. If a team optimizes a 2-GPU configuration to 85% utilization, the cost drops to roughly $0.62 per million tokens.

The deciding variable is utilization rather than volume. A rented H100 costs the same hourly rate whether it serves one request or 100 concurrent requests. High utilization amortizes the cost across more tokens.

Deployment Type Utilization Cost per 1M Tokens (Llama 3.3 70B)
Groq API (Blended) N/A $0.63 – $0.67
H100 (vLLM) 40% $1.24
H100 (Optimized) 85% $0.62

Groq’s baseline price matches the best-case number an optimized ML infra team reaches after significant work. The cost advantage of Groq disappears if a team can maintain near-maximum GPU utilization.

The NVIDIA and Groq relationship

NVIDIA signed a $20 billion non-exclusive licensing agreement for Groq’s inference technology on December 24, 2025. This deal moved Groq’s founder Jonathan Ross and 90% of the engineering team to NVIDIA. GroqCloud continues to operate independently under CEO Simon Edwards.

The NVIDIA deal changed how the industry views inference. NVIDIA produced the Groq 3 LPU at GTC 2026 as part of the Vera Rubin platform. This LPU technology serves as an inference co-processor. NVIDIA provides the Groq 3 LPX platform, a server rack powered by 128 individual Groq 3 LPUs. When used with the Vera Rubin NVL72 rack, customers see 35x higher throughput per megawatt of power.

Will NVIDIA’s massive infrastructure scale enough to stop the rise of these specialized inference insurgents? Groq remains a private company. NVIDIA’s deal was a licensing agreement, not a merger. Groq designers still focus on the LPU architecture to eliminate the cost of moving weights from external memory to compute.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories