GPT-6 272K long-context pricing

GPT-6 Astra, Sol, and Luna share a 272K long-context pricing cliff. When prompt tokens cross 272K, OpenAI reprices the entire request at a higher band, not only the overflow. That is different from Grok 4.7 (200K cliff) and from Claude Opus 5.5 / Fable 5.1 (flat list rates across their published windows). Rates in TokenCalculator checked 23 Sep 2026.

Check your prompt size first

Paste the real system plus tools plus RAG blob into the cost calculator with a GPT-6 model selected. If you are near 272K, redesign the prefix before you treat the 1.05M window as a flat price.

What the 272K cliff means

For GPT-6 family models in this catalog, standard rates apply when prompt tokens stay at or under 272K. Above that threshold, input, cached input, and output switch to the long-context band for the whole request. Design agents so fat prefixes stay under the line unless you intentionally buy the higher band.

GPT-6 long-context rate cards

ModelStandard in / outLong-context in / outLong-context cached
GPT-6 Luna$0.100/1M / $0.500/1M$0.200/1M / $0.750/1M$0.020/1M
GPT-6 Sol$2.00/1M / $10.00/1M$4.00/1M / $15.00/1M$0.400/1M
GPT-6 Astra$10.00/1M / $50.00/1M$20.00/1M / $75.00/1M$2.00/1M

Pattern on all three: input and cached input roughly double above 272K. Output rises to about 1.5x of the standard band. Luna stays cheapest in absolute dollars, but the cliff still doubles its input line.

272K vs 200K vs flat Claude windows

Provider familyCliffWhat happens
OpenAI GPT-6 (Astra / Sol / Luna)272K prompt tokensWhole request moves to long-context band
xAI Grok 4.7200K prompt tokensWhole request moves to $4 / $1 / $12 band
Claude Opus 5.5 / Fable 5.1None on published cardStandard per-token rates across the 1M window

Grok hits its cliff earlier (200K). Between 200K and 272K, Sol often wins the rate card versus Grok even when Grok wins below 200K. Opus 5.5 lists $4.00/1M / $20.00/1M flat across 1M. Fable 5.1 lists $10.00/1M / $50.00/1M flat. Flat does not mean free: you still pay for every token.

Worked cliff examples (Sol)

  • Just under the line: 270K fresh input + 2K output about $0.54 + $0.02 = $0.56 on standard Sol rates.
  • Just over the line: 275K fresh input + 2K output uses long-context Sol ($4 / $15). About $1.10 + $0.03 = $1.13 for the whole request.
  • Crossing by a few thousand tokens can nearly double the bill. Trim RAG or tools before you shrug past 272K.

How to stay under 272K

  • Summarize or retrieve fewer chunks instead of stuffing full documents
  • Keep tool schemas and system prompts lean
  • Split multi-repo contexts across turns when quality allows
  • Prefer Batch offline jobs only after you confirm the band you are in
  • Price Exact OpenAI counts in TokenCalculator before production cutover

Fast mode and Batch still stack

Fast mode is about 2x the band you are in. Batch is about 50% of that band when latency can wait. A Fast request already over 272K pays the long-context band first, then the Fast multiplier. Confirm service tiers on OpenAI pricing before you ship.

Common mistakes

  • Treating the 1.05M window as one flat price
  • Comparing Sol under 272K to Grok over 200K without noting both cliffs
  • Assuming Claude has the same cliff because it also offers ~1M context
  • Ignoring that Luna and Astra share the cliff shape even at different stickers

Frequently asked questions

What happens when a GPT-6 prompt exceeds 272K input tokens?

OpenAI reprices the full request at the long-context band for input, cached input, and output. Overflow alone is not billed at a special rate.

Does GPT-6 Luna also have the 272K cliff?

Yes. Luna, Sol, and Astra share the cliff shape. Luna stays cheaper in absolute dollars but still doubles input above 272K.

How is this different from Grok 4.7?

Grok moves to its long-context band at 200K prompt tokens. GPT-6 waits until 272K. See Grok 4.7 pricing and Sol vs Grok for side-by-side math.

Does Claude Opus 5.5 charge extra for long context?

Anthropic documents standard per-token pricing across the 1M window for Opus 5.5 in this catalog. There is no OpenAI-style 272K whole-request cliff on that card.

Is the cliff on prompt tokens or total tokens?

Catalog notes frame it as prompt (input) tokens crossing 272K. Confirm the exact meter language on OpenAI pricing for your account.

Can Batch avoid the cliff?

Batch discounts the band you are in. It does not remove the cliff. An oversized prompt still enters the long-context band, then Batch applies.

Are these rates live?

No. Curated list prices last verified 23 Sep 2026. Confirm on OpenAI and sibling provider pages.

What should I open first?

Paste your largest production prefix into the cost calculator on GPT-6 Sol. If you are within ~20K of 272K, trim before you scale traffic.


Next steps


Related Sep 2026 pricing guides

Frontier rate cards, cost-per-task math, and long-context cliffs. Cross-link these when you compare Sol, Luna, Opus 5.5, Grok 4.7, or Astra.