The o3 pricing reduction and the OpenAI model hierarchy

OpenAI has implemented an 80% price reduction for the o3 model, dropping input costs from $10 to $2 per million tokens. This update introduces significant shifts in the model hierarchy, affecting how developers route workloads between o3, o3-pro,…

The o3 pricing reduction and the OpenAI model hierarchy

I see the 80% reduction in o3 pricing as a major change in the recent OpenAI updates. The company dropped the input price from $10 per million tokens to $2 and the output price from $40 to $8. This change targets developers who found the previous $10/$40 rate a barrier to production. I note that o3-pro remains at $20 per million input tokens and $80 per million output tokens. The massive gap between the $2.00 input cost of the o3 tier and the $20.00 input cost of the o3-pro tier forces developers to make much more difficult decisions about which model to route their specific workloads to. One user reported a billing error on June 10, 2025, when they ran a batch using o3-2025-04-16. Even though the price reduction was announced, the user still paid the old rates of $5 per million for input and $20 per million for output, which totaled $150.00 for the batch.

o3-mini provides a specialized alternative for technical domains with an input price of $1.10 per million tokens and an output price of $4.40 per million tokens. This model replaces o1-mini and provides three reasoning effort options: low, medium, and high. I find the medium reasoning effort to be the default for ChatGPT users. It provides higher rate limits and lower latency than the previous model. ChatGPT Plus and Team users receive 150 messages per day with o3-mini, which is a tripling of the 50 messages per day allowed with o1-mini. The model is available to select developers in API usage tiers 3-5. It handles science, math, and coding tasks with improved STEM capabilities.

The pricing structure changes once a request crosses the 272K token boundary. For the GPT-5.6 family, input prices double and output prices rise by 50% once the request exceeds this limit. I note that the GPT-5.6 Luna model has a context window of 1.05 million tokens. Using the Luna model at $0.20 per million input tokens remains much cheaper than using the Sol model at $5.00 per million input tokens. You probably noticed the sudden shift in pricing after the July updates. The jump in cost is steep when moving from short-context to long-context workloads. The GPT-5.4 Nano and Mini models do not have a long-context tier because they top out at standard context lengths.

OpenAI provides the GPT-5.6 family in three tiers: Sol, Terra, and Luna. Sol costs $5.00 for input and $30.00 for output per million tokens. Terra costs $2.00 for input and $12.00 for output. Luna costs $0.20 for input and $1.20 for output. The flagship GPT-6 Astra costs $10.00 for input and $50.00 for output per million tokens. Astra has different prices depending on the processing mode. Batch and Flex processing halve the standard rates to $5 and $25, while Fast mode doubles the rates to $20 and $100. Will the recent price cuts on o3 eventually force a permanent reduction across the entire GPT-5.6 family?

Reasoning models like o3 and o3-pro use internal "thinking" steps that increase costs. These internal reasoning tokens bill at the output rate. I observe that this causes effective costs to run 3x to 10x higher than the base rate. For a support assistant on GPT-5.6 Terra using a 1,500-token prompt where 1,200 tokens are cached, the math is specific. The uncached 300 tokens cost $0.0006 at a $2.00 rate. The cached 1,200 tokens cost $0.00024 at a $0.20 rate. The 400-token response costs $0.0048 at a $12.00 rate. The total cost per call is $0.0057. Non-reasoning models such as the GPT-5.4 series bill at a standard rate for input and output.

DeepSeek V4 Pro provides a price of $0.42 per million input tokens and $0.84 per million output tokens. This is significantly lower than OpenAI’s GPT-4o at $2.50 per million input tokens and $10.00 per million output tokens. Claude 3.5 Sonnet costs $3.00 for input and $15.00 for output per million tokens. DeepSeek V4 has a 1 million token context window and a 384,000 token maximum output. It also has a 98% cache hit discount. Gemini 2.5 Pro provides a 2 million token context window at $1.25 for input and $10.00 for output. The Gemini 2.5 Pro maximum output is 64,000 tokens.

Prompt caching reduces input costs by 90%. For GPT-5.4, cached input costs $0.075 compared to the $0.75 standard rate. The Batch API provides a 50% discount for asynchronous workloads. I find the 10,000 daily conversation scenario for a startup chatbot useful for math. With GPT-5.4 Mini at $0.75 for input and $4.50 for output, 120 million input tokens and 120 million output tokens cost $900 and $5,400 respectively. This brings the monthly cost to $6,300. For the GPT-4.1 family, processing 10,000 customer support tickets with 500 input and 200 output tokens costs $16 with GPT-4.1, $3.20 with GPT-4.1 mini, and $0.80 with GPT-4.1 nano.

Fine-tuning training costs depend on the model. GPT-4.1 trains at $3.00 per million tokens. Fine-tuned inference for GPT-4.1 costs $3.00 for input and $12.00 for output. GPT-4.1 Mini trains at $0.80 per million tokens. Inference for GPT-4.1 Mini is $0.80 for input and $3.20 for output. The inference premium persists for the life of the deployment.

Model Category Model Name Input Price (USD) Output Price (USD) Context Window
Reasoning o3 $2.00 $8.00 200,000
Reasoning o3-pro $20.00 $80.00 200,000
Reasoning o3-mini $1.10 $4.40 200,000
Flagship GPT-6 Astra $10.00 $50.00 1,050,000
Flagship GPT-5.6 Sol $5.00 $30.00 1,050,000
Flagship GPT-5.6 Terra $2.00 $12.00 1,050,000
Flagship GPT-5.6 Luna $0.20 $1.20 1,050,000
Budget GPT-5.4 Nano $0.20 $0.20 200,000
Budget GPT-5.4 Mini $0.75 $4.50 200,000
Budget GPT-4.1 mini $0.40 $1.60 1,000,000
Budget GPT-4.1 $2.00 $8.00 1,000,000
airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories