Skip to main content

OpenAI API pricing: the cheapest model is 50 times less than the flagship

GPT-6 Astra is $10 per million input tokens. GPT-5.6 Luna is $0.20. Same API, same call, a 50x gap. What each model costs and how caching and batching cut the bill further.

By Tashawar AwaisResearcher and editor
Last updated 4 min readPricing verified

Verdict

Pricing guide

GPT-6 Astra is $10 per million input tokens and $50 output. GPT-5.6 Luna is $0.20 and $1.20, a 50-fold difference on input. Cached input bills at a tenth, and the Batch API halves everything.

The whole rate card#

ModelInputOutput
GPT-6 Astra$10.00$50.00
GPT-5.6 Sol$5.00$30.00
GPT-5.6 Terra$2.00$12.00
GPT-5.6 Luna$0.20$1.20

Per million tokens. Cached input bills at roughly a tenth of the standard rate. The Batch API halves everything.

The number that should change how you build#

50x

Input price, Astra against Luna

$10.00 per million against $0.20. Same API, same call, one line of difference in your code.

This is the single most useful fact about OpenAI pricing and it is almost never stated plainly.

The gap between the most and least expensive model is fiftyfold. Not fifty percent. Fifty times.

That means the model string in your code is, for most applications, a bigger cost lever than anything else you could optimise. And there is no warning attached to it: the API will happily run classification work that Luna handles perfectly on Astra at fifty times the price.

What each model is actually for#

ModelUse it for
Astra, $10/$50Hard reasoning, long documents, work where a wrong answer is expensive
Sol, $5/$30General purpose, the sensible default for most production work
Terra, $2/$12Well-defined tasks with clear instructions
Luna, $0.20/$1.20Classification, extraction, routing, tagging, summarising short text

Most business applications are Luna and Terra work wearing Astra's price tag. Extracting a date from an email is not a reasoning problem.

The July cuts, and who got them#

On 30 July 2026 OpenAI cut two models and left one alone:

ModelCut
Luna80 percent
Terra20 percent
SolNothing

Luna's cut is the reason the fifty-times gap exists at all. Before it, the spread was much narrower.

This matters beyond one price. The cheap end is getting cheaper fast while the flagship end holds. If you built something a year ago on a mid-tier model and have not revisited it, the arithmetic that justified that choice has changed underneath you.

Two levers that cut the bill without changing models#

Caching. Repeated context bills at roughly a tenth of standard input. If every call sends the same system prompt, the same instructions, or the same reference document, you are paying full price for identical tokens over and over. This is the largest easy saving available.

Batching. The Batch API halves every rate if you can accept a delay. Anything scheduled, overnight, or not blocking a user should be on it.

Where the rate doubles#

Past the standard context window, rates roughly double.

That is worth knowing before you build something that leans on a million-token context. The big window exists, but it is not priced flat across its whole range, and a workload that routinely runs long is not paying the headline rate you budgeted with.

API or subscription?#

IfBuy
A person is typing the questionsA subscription. Plus at $20 covers far more than $20 of tokens
An application is making the callsThe API
A team of people typingBusiness, $20 a seat
BothBoth. They are different products

The confusion is common and expensive in one direction: people build an internal tool on the API to save on subscriptions, then discover that five staff generating text all day costs far more in tokens than five $20 seats.

What a small business should actually do#

  1. Default to Sol, not Astra. Move up only when you can point at a task it fails
  2. Try Luna on your highest-volume task. If quality holds, that one change is most of your bill
  3. Cache anything repeated. A tenth of the price for the same tokens
  4. Batch anything that can wait. Half price for patience

The full picture on the flagship, including the doubling against Sol, is in GPT-6 Astra pricing.

The short version

What works

  • The spread between models is enormous, so most workloads can be moved to a cheaper one without anyone noticing
  • Cached input at a tenth of standard rates is the single largest saving available and takes little work
  • The Batch API halves every rate if you can wait for results, which suits overnight and scheduled work

What does not

  • Pricing is per token, so the bill scales with usage in a way a subscription does not, and can surprise you
  • Rates roughly double once a prompt exceeds the standard context window
  • Model names give no indication of price, so picking the newest one is an easy and expensive default

Frequently asked questions

How much does the OpenAI API cost?
It depends entirely on the model. GPT-6 Astra is $10 per million input tokens and $50 output. GPT-5.6 Sol is $5 and $30, Terra is $2 and $12, and Luna is $0.20 and $1.20. Cached input bills at roughly a tenth of the standard rate and the Batch API halves everything.
What is the cheapest OpenAI model?
GPT-5.6 Luna at $0.20 per million input tokens and $1.20 output, after an 80 percent price cut on 30 July 2026. That is fifty times cheaper than GPT-6 Astra on input and roughly forty times cheaper on output.
Is the API cheaper than a ChatGPT subscription?
For a person, almost never. ChatGPT Plus at $20 a month covers far more usage than $20 of API tokens would at flagship rates. The API is for applications, not for people typing questions. If a human is doing the asking, buy the subscription.
How do I reduce an OpenAI API bill?
Three things, in order of effect. Move work to a cheaper model and measure whether anyone notices. Cache repeated context, which bills at about a tenth. Use the Batch API for anything that does not need an answer immediately, which halves the rate.
Portrait of Tashawar Awais

Written by

Tashawar Awais

Researcher and editor

Tashawar handles verification and editing. Every figure in a review is checked a second time before it goes out, and anything that cannot be traced to a vendor page or a documented source is either qualified or cut. Where pricing is genuinely unclear, as it is with Canva team plans or Close CRM tiers, the article says so and tells the reader to confirm directly instead of quoting a number with false confidence.

  • Second check on every published figure
  • Removes claims the sources do not support
  • Flags pricing that changes without notice