OpenAI API Pricing in 2026 — What You Actually Pay Per Token

OpenAI API Pricing in 2026 — What You Actually Pay Per Token — illustration

GPT-5.6 Sol costs $4.00 per million input tokens and $20.00 per million output right now, down from $5.00 and $30.00. That’s a promotional price, cut on 21 August 2026, which OpenAI says is “available at least through November 21, 2026.” Terra is $2.00 in and $12.00 out, Luna $0.20 and $1.20. All of those are standard-tier prices for prompts up to 272K input tokens. Go past that and input doubles while output rises 1.5x, which most pricing write-ups never mention.

There’s also a newer line on top. GPT-6 Sol, released 22 September, is $2.00 in and $10.00 out in the same short-context band: half what GPT-5.6 Sol costs even at its discounted price.

Last verified: 26 September 2026. This pass caught the 21 August Sol cut, which the previous version of this page missed and kept quoting the old $5.00/$30.00 for five weeks, and added the three GPT-6 models.

I re-check this page against OpenAI’s own pricing page, model pages and changelog, not against other people’s summaries. I aim for monthly. This time the gap was more than six weeks and a price moved inside it, which is exactly why the date above is at the top and not in a footer.

The GPT-5.6 pricing ladder

Everything below is the Standard service tier, in dollars per 1M tokens, from developers.openai.com/api/docs/pricing.

One thing you need to know before reading any GPT-5.6 price anywhere: OpenAI’s table is split into two context bands. The header tooltips define them precisely — short context is ≤272K input tokens, long context is >272K. Almost every summary you’ll find quotes only the short-context half. Here are both.

Short context (≤272K input tokens):

ModelInputCached inputOutput
gpt-5.6-sol (promotional)$4.00$0.40$20.00
gpt-5.6-terra$2.00$0.20$12.00
gpt-5.6-luna$0.20$0.02$1.20

Long context (>272K input tokens):

ModelInputCached inputOutput
gpt-5.6-sol (promotional)$8.00$0.80$30.00
gpt-5.6-terra$4.00$0.40$18.00
gpt-5.6-luna$0.40$0.04$1.80

The jump is not a rounding detail: input doubles and output goes up 1.5x the moment a request crosses 272K input tokens. If your workload is long documents — codebases, contracts, transcript piles — the long-context column is your real price, and budgeting off the short-context one understates your bill by around 2x on input.

GPT-5.6 Sol, Terra, and Luna API pricing compared, with Sol at its promotional price Log-scale bar chart comparing OpenAI's GPT-5.6 Standard-tier, short-context (up to 272K input tokens) API pricing in US dollars per 1 million tokens, verified 26 September 2026. Sol is shown at its promotional price, which OpenAI says is available at least through November 21, 2026. Long-context prices (over 272K input tokens) are higher and shown in the article tables. Input price per 1M tokens, short context: Sol $4.00, Terra $2.00, Luna $0.20. Output price per 1M tokens, short context: Sol $20.00, Terra $12.00, Luna $1.20. Source: developers.openai.com/api/docs/pricing. GPT-5.6 API Pricing Standard tier · short context (≤272K input) · $ per 1M tokens Log scale — each gridline is 10× the last Input Output $0.1 $1 $10 $100 Sol* — Input $4.00 Sol* — Output $20.00 Terra — Input $2.00 Terra — Output $12.00 Luna — Input $0.20 Luna — Output $1.20 Nearly 17× cheaper: Luna vs. Sol output $1.20 vs. $20.00 per 1M output tokens *Sol promotional price, through at least 21 Nov 2026 Verified 26 Sep 2026 · developers.openai.com/api/docs/pricing
GPT-5.6 Standard-tier, short-context (≤272K input tokens) API pricing for Sol, Terra and Luna, input and output cost in dollars per 1 million tokens, on a log scale. Sol is shown at its promotional price, which OpenAI says is available at least through November 21, 2026. Luna's output price is nearly 17x cheaper than Sol's ($1.20 vs $20.00 per 1M tokens). Verified 26 September 2026 against developers.openai.com/api/docs/pricing.

There’s a fourth column on OpenAI’s table that most write-ups drop, and that the two tables above leave out: cache writes, which cost more than a plain input token. In short context, Sol writes are $5.00 at the promotional price, Terra $2.50, Luna $0.25. That’s 1.25x the standard input price, so priming a cache costs you a 25% premium on the first pass in exchange for 90% off every read after it. Sol’s long-context write is $10.00.

All three prices have moved since late July. The July 30, 2026 changelog entry cut two of them, and not by the same amount: “Starting July 30, GPT-5.6 Luna costs 80% less, while GPT-5.6 Terra costs 20% less.” Luna got the headline. If you sized a budget off a Terra quote from July, your estimate is 25% too high. If you sized one off Luna, your estimate is five times too high, which is the nicer direction to be wrong in.

All three share the same envelope: 1.05M token context window, 128K max output, knowledge cutoff February 16, 2026. So any of them fits your documents, but fitting and costing the same are different things. Sol’s model page spells the band rule out: “Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request.” That applies to all three equally, so the choice between models is about cost and quality, while the choice about how much context you send is a separate pricing decision of its own.

Sol’s price is a promotion with an end date and no stated sequel

Sol moved on 21 August 2026. OpenAI’s community announcement said it was “dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months,” and that the cut would “also apply to Fast mode, long-context requests, and Batch and Flex processing.” Short-context input went from $5.00 to $4.00, cached input from $0.50 to $0.40, and output from $30.00 to $20.00. So the output cut is a third, not the fifth the headline suggests.

Both the pricing page and Sol’s own model page carry the same sentence: “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” Neither of those pages, nor the community announcement, says what Sol costs after that. It could go back to $5.00/$30.00, land somewhere new, or stay put. I’m not guessing. If you’re committing a budget past late November on GPT-5.6 Sol, model it at the old price and treat anything lower as a bonus.

GPT-6 Astra, Sol and Luna

OpenAI released GPT-6 Astra on 3 September 2026, then GPT-6 Sol and GPT-6 Luna on 22 September. The catalog taglines tell you the pecking order: Astra is “Our most capable model, built for the hardest end-to-end work,” Sol is “Built to power complex coding and agentic workflows,” and Luna is “Our most efficient model for focused, high-volume tasks.”

Same two-band structure as GPT-5.6, same 272K line. GPT-6 Sol’s model page puts it this way: “Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.” Standard tier, dollars per 1M tokens:

Short context (≤272K input tokens):

ModelInputCached inputCache writeOutput
gpt-6-astra$10.00$1.00$12.50$50.00
gpt-6-sol$2.00$0.20$2.50$10.00
gpt-6-luna$0.10$0.01$0.125$0.50

Long context (>272K input tokens):

ModelInputCached inputCache writeOutput
gpt-6-astra$20.00$2.00$25.00$75.00
gpt-6-sol$4.00$0.40$5.00$15.00
gpt-6-luna$0.20$0.02$0.25$0.75

Fast mode is exactly 2x Standard for all three, in both bands. One catch on the most expensive combination: OpenAI’s Fast mode guide says “Fast mode for GPT-6 Astra does not include a latency SLA.” In short context that’s Astra $20.00 in / $100.00 out, Sol $4.00/$20.00, Luna $0.20/$1.00. In long context, Astra $40.00/$150.00, Sol $8.00/$30.00, Luna $0.40/$1.50.

Batch and Flex are currently identical for the three GPT-6 models, both at half of Standard. Short context: Astra $5.00/$25.00, Sol $1.00/$5.00, Luna $0.05/$0.25. Long context: Astra $10.00/$37.50, Sol $2.00/$7.50, Luna $0.10/$0.375.

All three have a 1.05M context window and 128K max output. Knowledge cutoffs differ: Astra April 30, 2026, Sol April 20, 2026, Luna May 18, 2026.

The number that jumped out at me: GPT-6 Sol at $2.00/$10.00 is exactly half of GPT-5.6 Sol’s discounted $4.00/$20.00, and it has a later knowledge cutoff. I haven’t run the two against each other on my own work, so I can’t tell you they’re equal on quality. On price alone, I wouldn’t start anything new on GPT-5.6 Sol. Astra is the other story. At $50.00 output in short context and $150.00 on Fast long-context, it’s the priciest of the six flagship models by a distance, and you want a reason before you point a loop at it.

GPT-5.6 is not going anywhere yet

This one nearly caught me. The models overview page at developers.openai.com/api/docs/models now shows only the three GPT-6 models under “Flagship models.” Land there first and you’d assume GPT-5.6 is gone. It isn’t. The full catalog at developers.openai.com/api/docs/models/all lists all six under that same heading, none labelled deprecated or legacy, and each GPT-5.6 model still has its own live, priced page.

The deprecations page agrees. It lists no GPT-5.6 model as the thing being retired, and it still points retiring models at gpt-5.6-sol, terra and luna as replacements. I checked every row on 26 September and none of them names a GPT-6 model as a replacement yet.

Legacy models are still billed at their own rates

This matters more than it sounds. When a vendor retires a model slug, one of two things happens: the request errors, or it silently redirects and bills at the successor’s rate. I checked every older row on OpenAI’s table against its GPT-5.6 replacement and none of them have been quietly repriced, even the ones already carrying a shutdown date. You pay what the row says.

ModelInputCached inputOutput
gpt-5$1.25$0.125$10.00
gpt-5-mini$0.25$0.025$2.00
gpt-5-nano$0.05$0.005$0.40
gpt-5-pro$15.00—$120.00
gpt-4.1$2.00$0.50$8.00
gpt-4.1-mini$0.40$0.10$1.60
gpt-4.1-nano$0.10$0.025$0.40
gpt-4o$2.50$1.25$10.00
gpt-4o-mini$0.15$0.075$0.60
gpt-3.5-turbo$0.50—$1.50
o1$15.00$7.50$60.00
o1-pro$150.00—$600.00
o3$2.00$0.50$8.00
o3-pro$20.00—$80.00
o3-mini$1.10$0.55$4.40
o4-mini$1.10$0.275$4.40

o1-pro at $150 input and $600 output per million is still sitting there, 37.5 times GPT-5.6 Sol’s current input price for a model that shuts down in October. Nobody should be sending it traffic.

Two other things fall out of this table. Caching discounts vary a lot across the legacy rows: gpt-5, gpt-5-mini and gpt-5-nano keep the full 90%, gpt-4.1 (all three sizes), o3 and o4-mini get 75%, and gpt-4o, gpt-4o-mini, o1 and o3-mini only 50%. If you’re still on one of the 50% rows, caching is worth much less than the GPT-5.6 numbers suggest. And on OpenAI’s own table (the column isn’t reproduced here), the legacy rows show a dash under cache writes, so the write premium is a GPT-5.6-era addition rather than something you were already paying.

Fast mode is Priority Processing with a new name

Do not treat this as a new feature to evaluate. OpenAI’s own guide says: “Priority processing was renamed Fast mode on July 30, 2026.” Same tier, new label, available now. The old service_tier: "priority" value still works and is an alias for service_tier: "fast", so nothing in your code breaks.

What did change is speed. OpenAI says it “increased the speed at which Fast mode operates for gpt-5.6-sol to make it up to 2.5× faster than Standard processing,” with “more consistent latency while keeping pay-as-you-go flexibility.”

The price is the interesting part. For the entire GPT-5.6 family, Fast mode is exactly 2x Standard, to the cent, and that holds in both context bands and through Sol’s promotional cut. (The GPT-6 three follow the same 2x rule; their Fast prices are in the section above.) Short-context (≤272K input) Fast prices:

ModelFast inputFast cachedFast output
gpt-5.6-sol (promotional)$8.00$0.80$40.00
gpt-5.6-terra$4.00$0.40$24.00
gpt-5.6-luna$0.40$0.04$2.40

Above 272K input tokens the same 2x applies to the long-context rates: Sol runs $16.00 in / $60.00 out at the promotional price, Terra $8.00/$36.00, Luna $0.80/$3.60.

Older models get a smaller markup. gpt-4o goes from $2.50 to $4.25 input, which is 1.7x. o3 goes $2.00 to $3.50, or 1.75x. gpt-4o-mini goes $0.15 to $0.25, about 1.67x. If your latency-sensitive path is still on gpt-4o, Fast mode is proportionally cheaper there than it would be on the new family — the sort of asymmetry that quietly changes the shape of a migration plan.

Fast mode isn’t available on everything. When I went through the fast tier in August it listed gpt-5.6-sol, terra and luna, then gpt-5.5, gpt-5.4, gpt-5.4-mini, gpt-5.2, gpt-5.1, gpt-5, gpt-5-mini, gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-2024-05-13, gpt-4o-mini, o3 and o4-mini, and no other slugs had populated prices. The three GPT-6 models have Fast prices now too.

Batch and Flex: the same 50%, two different trades

Batch is 50% off with a 24-hour completion window, plus what OpenAI calls “a separate pool of significantly higher rate limits.” For GPT-5.6 Sol at the promotional price that’s $2.00 input and $10.00 output in short context (≤272K input tokens), and $4.00 in / $15.00 out in long context. Terra lands at $1.00/$0.10/$1.25/$6.00 (input, cached, cache write, output) and Luna at $0.10/$0.01/$0.125/$0.60, both short context. Exact halves, no rounding surprises.

Flex matches Batch for GPT-5.6 Sol in both bands: $2.00 in / $10.00 out in short context, and $4.00 in / $0.40 cached / $5.00 cache write / $15.00 out in long context. For the GPT-6 models Batch and Flex are identical in both bands.

The two tiers are not mirror images across the whole table, though: Batch prices noticeably more model rows than Flex, so plenty of slugs have a Batch price and no Flex option at all. (The two easy to miss on the Flex list are gpt-5.5-pro and gpt-5.4-pro, both at $15.00 in / $90.00 out — half their standard $30/$180.) And Flex carries a caveat Batch doesn’t: OpenAI’s guide says plainly that “Flex processing is in beta with limited model availability.” What you’re buying is different too. Batch trades turnaround for money; Flex trades per-request latency for money and warns about “occasional resource unavailability.” So Flex is for synchronous work where you can tolerate slow and can retry, and Batch is for work you can hand over and collect tomorrow. Same bill, very different failure mode.

Prompt caching

OpenAI’s caching guide says “Prompt caching is enabled by default for supported OpenAI models”, and “The minimum cacheable prompt length is 1,024 tokens for GPT-5.6 and later.” On the GPT-5.6 family that’s a 90% discount on the cached portion, which is the single biggest lever on this whole page if you run long system prompts.

A correction on this section: the August version warned that GPT-5.6 stopped falling back to shorter cached prefixes, quoting OpenAI’s caching guide. OpenAI has since rewritten that guide and the quoted wording is gone. It now describes walking the cache lookup boundaries “from longest prefix to shortest, looking for an available matching prefix already cached”, so I’ve dropped the warning. Watch your cached-token counts after any model switch anyway; it’s the cheapest check you can run.

One more line item that’s easy to miss: regional processing endpoints carry a 10% uplift for models released on or after March 5, 2026 that are eligible for data residency. Every price above assumes you’re not using them.

Usage tiers

Access is gated by cumulative spend, and promotion is automatic — the docs say “as your spend on our API goes up, we automatically graduate you to the next usage tier.”

TierQualificationMonthly usage limit
FreeAllowed geography$100 / month
Tier 1$5 paid$100 / month
Tier 2$50 paid$500 / month
Tier 3$100 paid$1,000 / month
Tier 4$250 paid$5,000 / month
Tier 5$1,000 paid$200,000 / month

Spend is the only variable on that page. There’s no waiting period listed, no account-age criterion, nothing about payment history.

Two things I went looking for and could not find. First, per-model RPM and TPM numbers: I checked the rate-limits guide, the models page and the pricing page on 10 August, then again on 26 September along with the full models catalog, and none of them carry a public per-tier RPM/TPM table. The rate-limits guide points you at the models page for “a high-level summary,” but the models page’s public HTML has no such figures, and the authoritative per-org numbers live behind login at platform.openai.com/settings/organization/limits, which returns 403 without an account. Second, a new-account free credit: I checked the rate-limits guide, the quickstart and the pricing page, and none of them state a dollar grant for new signups. The widely repeated “$5 free credit” isn’t on any of those pages, so I’m not printing it as a number. The only “Free” is that top table row, and it’s a ceiling, not a gift.

Dated shutdowns you need in your calendar

These are absolute dates from OpenAI’s deprecations page. I’m writing them out in full because a page that says “next month” is worthless six weeks later.

DateWhat goes awayReplacement
10 August 2026 (passed)gpt-5.2-chat-latest, gpt-5.3-chat-latestgpt-5.6-sol
26 August 2026 (passed)Assistants APIResponses API + Conversations API
24 September 2026 (passed)Videos API, sora-2, sora-2-pro and all their snapshotsnone listed
28 September 2026gpt-3.5-turbo-instruct, gpt-3.5-turbo-1106, babbage-002, davinci-002gpt-5.6-terra
1 October 2026gpt-5.4-cybergpt-5.6-cyber
23 October 2026gpt-3.5-turbo (base and 0125), gpt-4 and gpt-4-turbo families, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1, o1-pro, o3-mini, o4-mini, ft-o4-mini-2025-04-16GPT-5.6 family (ft-o4-mini → gpt-5.6-terra); gpt-image-1 → gpt-image-2
30 November 2026v1/prompts reusable prompt objects, the Evals dashboard and API (read-only from 31 October), Agent BuilderPromptfoo (Evals); Agents SDK or ChatGPT Workspace Agents (Agent Builder)
1 December 2026gpt-image-1.5, gpt-image-1-mini, chatgpt-image-latestgpt-image-2
11 December 2026gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07, gpt-5-pro-2025-10-06, o3-2025-04-16, o3-pro-2025-06-10Sol / Terra / Luna
20 January 2027gpt-realtime, gpt-audio, gpt-4o-audio, gpt-4o-realtime and the mini variantsgpt-realtime-2.1 (gpt-realtime-mini and gpt-4o-mini-realtime → gpt-realtime-2.1-mini), gpt-audio-1.5 (including gpt-audio-mini and gpt-4o-mini-audio); gpt-4o-mini-transcribe → its 2025-12-15 snapshot
26 February 2027whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarizegpt-live-transcribe or gpt-transcribe

Two of those rows are new since the last version of this page. The gpt-5.4-cyber removal was announced on 11 September with under three weeks’ notice. The transcription one was announced on 26 August and is a separate entry from the January snapshot line above it. If you’re moving off gpt-4o-mini-transcribe for January anyway, I’d plan one migration straight to gpt-transcribe or gpt-live-transcribe rather than two.

Mind the gpt-3.5 split: the -instruct and -1106 variants go on 28 September, a month before the rest of the family. If you’re on either — both are still individually priced today — the October date everyone quotes is a month late for you.

The two chat-latest snapshots went dark on 10 August 2026, the same day this article was drafted. Notice also that the December wave maps gpt-5-pro and o3-pro onto gpt-5.6-sol with reasoning.mode: pro rather than onto a separate pro model.

Fine-tuning is winding down too. The pricing page says the platform “is no longer accessible to new users,” and the deprecations page puts a hard date on the rest: active existing customers can no longer create new fine-tuning jobs after 6 January 2027. Fine-tuned models stay available for inference until their base models are deprecated, and for fine-tuned o4-mini that is soon: ft-o4-mini-2025-04-16 shuts down on 23 October 2026, with gpt-5.6-terra as the listed replacement.

Everything else on the bill

Embeddings are cheap enough to ignore in most budgets: text-embedding-3-small is $0.02 per 1M, 3-large is $0.13, and ada-002 is $0.10 — which means the old ada model now costs five times the current small one. Moderation via omni-moderation-latest is free, listed literally as “Free” in the input column.

Image generation bills tokens in two streams — and both older image models have a shutdown date. gpt-image-1.5 charges $8.00 input, $2.00 cached and $32.00 output per 1M image tokens, plus $5.00/$1.25/$10.00 on the text side. gpt-image-1-mini is $2.50/$0.25/$8.00 for image tokens and $2.00/$0.20 for text input and cached input. Both are removed from the API on 1 December 2026 along with chatgpt-image-latest, replaced by gpt-image-2 — which is already priced on the page at $8.00/$2.00/$30.00 per 1M image tokens ($5.00/$1.25 text). If you’re building on image generation in the autumn, build on gpt-image-2 from the start.

Last housekeeping note, and it trips up scripts more than people: platform.openai.com/docs/* now 301-redirects to developers.openai.com/api/docs/*. Old bookmarks resolve fine, so nothing looks broken, but if you’ve got a scraper or an agent pinned to the old host it’s following a redirect it may not be logging. Same class of quiet migration as the retired Grok slugs that keep billing without erroring.

If you’re comparing this against what other vendors charge, the free AI API rate limits page covers the zero-dollar end of the market, and whether the Claude API is free does the same for Anthropic.

Frequently Asked Questions

How much does the OpenAI API cost per million tokens in 2026?

For prompts up to 272K input tokens (short context), GPT-5.6 Sol is currently $4.00 per 1M input and $20.00 per 1M output on promotional pricing OpenAI says is available at least through November 21, 2026; Terra is $2.00/$12.00 and Luna $0.20/$1.20. Above 272K input tokens: Sol $8.00/$30.00, Terra $4.00/$18.00, Luna $0.40/$1.80.

How much does GPT-6 cost on the OpenAI API?

In short context (up to 272K input tokens), GPT-6 Astra is $10.00 input and $50.00 output per 1M tokens, GPT-6 Sol is $2.00/$10.00 and GPT-6 Luna is $0.10/$0.50. Above 272K input tokens they're $20.00/$75.00, $4.00/$15.00 and $0.20/$0.75.

Is GPT-5.6 deprecated now that GPT-6 is out?

No. GPT-5.6 Sol, Terra and Luna are still listed under Flagship models on OpenAI's full model catalog with no deprecated or legacy label, and the deprecations page doesn't list any of them.

How much extra does OpenAI Fast mode cost?

Exactly double the standard price for the GPT-5.6 and GPT-6 families, in both context bands. GPT-5.6 Sol goes to $8.00 input and $40.00 output per 1M in short context and $16.00/$60.00 above 272K input tokens; older models like gpt-4o get a smaller markup, $2.50 to $4.25 on input.

When do GPT-4 and GPT-3.5 shut down on the OpenAI API?

Most of the gpt-3.5-turbo, gpt-4 and gpt-4-turbo families shut down on October 23, 2026, along with o1, o1-pro, o3-mini and o4-mini — but gpt-3.5-turbo-instruct and gpt-3.5-turbo-1106 go earlier, on September 28, 2026. The original GPT-5 snapshots follow on December 11, 2026.

Written by Hirak Banerjee

Indie dev and maker. I build AI-powered apps and write about the tools I actually use. Follow on X · GitHub

Get told when these numbers change

Every figure here is re-checked against the vendor's own docs. Leave your email and we'll tell you when one moves — only when something actually changes.

Join builders who ship faster. No spam.

Comments