Home / Pricing / Long context

Cheapest long context models

Models with at least 128,000 tokens of context, ranked by input rate. A note appears when the catalog publishes a higher long context tier. Example request stays 1000 input and 500 output, so it does not cross those tiers.

132 models · rates checked . Price rank only, not a quality score.

Ranked by published rates

#ModelProviderContextInput / 1MOutput / 1MExample10k requests
1Gemini 1.5 Flash-8BGoogle1M$0.037/1M$0.150/1M$0.000112$1.12Cost
2Command R7BCohere128K$0.037/1M$0.150/1M$0.000112$1.12Cost
3Llama 3.2 1B (Groq)Groq128K$0.040/1M$0.040/1M$0.00006$0.6Cost
4Llama 3.1 8B Instant (Groq)Groq128K$0.050/1M$0.080/1M$0.00009$0.9Cost
5GPT-5 NanoOpenAI128K$0.050/1M$0.400/1M$0.00025$2.5Cost
6Llama 3.2 3B (Groq)Groq128K$0.060/1M$0.060/1M$0.00009$0.9Cost
7Llama 3.2 3B (Together)Together AI128K$0.060/1M$0.060/1M$0.00009$0.9Cost
8Gemini 1.5 FlashGoogle1M$0.075/1M$0.300/1M$0.000225$2.25Cost
9Gemini 2.0 Flash-LiteGoogle1M$0.075/1M$0.300/1M$0.000225$2.25Cost
10Llama 3.2 3B (Fireworks)Fireworks128K$0.100/1M$0.100/1M$0.00015$1.5Cost
11GPT-4.1 NanoOpenAI1M$0.100/1M$0.400/1M$0.0003$3.00Cost
12Gemini 2.0 FlashGoogle1M$0.100/1M$0.400/1M$0.0003$3.00Cost
13Gemini 2.5 Flash-LiteGoogle1M$0.100/1M$0.400/1M$0.0003$3.00Cost
14GPT-6 Luna (higher rate after 272K)OpenAI1.1M$0.100/1M$0.500/1M$0.00035$3.5Cost
15Llama 4 Scout (Groq)Groq128K$0.110/1M$0.340/1M$0.00028$2.8Cost
16DeepSeek CoderDeepSeek128K$0.140/1M$0.280/1M$0.00028$2.8Cost
17Ministral 8BMistral128K$0.150/1M$0.150/1M$0.000225$2.25Cost
18Mistral NemoMistral128K$0.150/1M$0.150/1M$0.000225$2.25Cost
19Pixtral 12BMistral128K$0.150/1M$0.150/1M$0.000225$2.25Cost
20GPT-4o miniOpenAI128K$0.150/1M$0.600/1M$0.00045$4.5Cost
21Mistral SmallMistral128K$0.150/1M$0.600/1M$0.00045$4.5Cost
22Command RCohere128K$0.150/1M$0.600/1M$0.00045$4.5Cost
23Command R (08-2024)Cohere128K$0.150/1M$0.600/1M$0.00045$4.5Cost
24Llama 3.2 11B Vision (Groq)Groq128K$0.180/1M$0.180/1M$0.00027$2.7Cost
25Llama 3.1 8B (Together)Together AI128K$0.180/1M$0.180/1M$0.00027$2.7Cost
26Llama 3.1 8B (Fireworks)Fireworks128K$0.200/1M$0.200/1M$0.0003$3.00Cost
27Gemma 3 27BGoogle128K$0.200/1M$0.400/1M$0.0004$4.00Cost
28Grok 2 MinixAI131K$0.200/1M$0.500/1M$0.00045$4.5Cost
29GPT-5.6 Luna (higher rate after 272K)OpenAI1.1M$0.200/1M$1.20/1M$0.0008$8.00Cost
30Llama 4 Maverick (Fireworks)Fireworks1M$0.220/1M$0.880/1M$0.00066$6.6Cost
31Claude 3 HaikuAnthropic200K$0.250/1M$1.25/1M$0.000875$8.75Cost
32Gemini 3.1 Flash-LiteGoogle1M$0.250/1M$1.50/1M$0.001$10.00Cost
33GPT-5 MiniOpenAI128K$0.250/1M$2.00/1M$0.00125$12.5Cost
34Llama 4 Maverick (Together)Together AI1M$0.270/1M$0.850/1M$0.000695$6.95Cost
35DeepSeek V3DeepSeek128K$0.270/1M$1.10/1M$0.00082$8.2Cost
36DeepSeek ChatDeepSeek128K$0.280/1M$0.420/1M$0.00049$4.9Cost
37QwQ 32B (Groq)Groq128K$0.290/1M$0.390/1M$0.000485$4.85Cost
38Qwen3 32B (Groq)Groq131K$0.290/1M$0.590/1M$0.000585$5.85Cost
39Grok 3 MinixAI131K$0.300/1M$0.500/1M$0.00055$5.5Cost
40DeepSeek R1 Distill Qwen 32BDeepSeek128K$0.300/1M$0.600/1M$0.0006$6.00Cost
41CodestralMistral256K$0.300/1M$0.900/1M$0.00075$7.5Cost
42DeepSeek V4 FlashDeepSeek1M$0.300/1M$1.20/1M$0.0009$9.00Cost
43Gemini 2.5 FlashGoogle1M$0.300/1M$2.50/1M$0.00155$15.5Cost
44GPT-4.1 MiniOpenAI1M$0.400/1M$1.60/1M$0.0012$12.00Cost
45Mistral Large 3Mistral128K$0.500/1M$1.50/1M$0.00125$12.5Cost
46Gemini 3 FlashGoogle1M$0.500/1M$3.00/1M$0.002$20.00Cost
47DeepSeek R1DeepSeek128K$0.550/1M$2.19/1M$0.001645$16.45Cost
48DeepSeek ReasonerDeepSeek128K$0.550/1M$2.19/1M$0.001645$16.45Cost
49Llama 3.3 70B (Groq)Groq128K$0.590/1M$0.790/1M$0.000985$9.85Cost
50DeepSeek R1 Distill Llama 70B (Groq)Groq128K$0.750/1M$0.990/1M$0.001245$12.45Cost
51Gemini 3.8 FlashGoogle1.0M$0.750/1M$3.75/1M$0.002625$26.25Cost
52Claude 3.5 HaikuAnthropic200K$0.800/1M$4.00/1M$0.0028$28.00Cost
53Llama 3.1 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00132$13.2Cost
54Llama 3.3 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00132$13.2Cost
55DeepSeek R1 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.00135$13.5Cost
56DeepSeek V3 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.00135$13.5Cost
57Llama 3.1 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.00135$13.5Cost
58Llama 3.3 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.00135$13.5Cost
59Command NightlyCohere128K$1.00/1M$2.00/1M$0.002$20.00Cost
60Claude Haiku 4.5Anthropic200K$1.00/1M$5.00/1M$0.0035$35.00Cost
61o1-miniOpenAI128K$1.10/1M$4.40/1M$0.0033$33.00Cost
62o3-miniOpenAI200K$1.10/1M$4.40/1M$0.0033$33.00Cost
63o4-miniOpenAI200K$1.10/1M$4.40/1M$0.0033$33.00Cost
64DeepSeek V3 (Together)Together AI128K$1.25/1M$1.25/1M$0.001875$18.75Cost
65Grok 3xAI131K$1.25/1M$2.50/1M$0.0025$25.00Cost
66Grok 4xAI256K$1.25/1M$2.50/1M$0.0025$25.00Cost
67Grok 4 FastxAI256K$1.25/1M$2.50/1M$0.0025$25.00Cost
68Gemini 1.5 ProGoogle2M$1.25/1M$5.00/1M$0.00375$37.5Cost
69Gemini 2.0 Pro ExpGoogle2M$1.25/1M$5.00/1M$0.00375$37.5Cost
70GPT-5OpenAI400K$1.25/1M$10.00/1M$0.00625$62.5Cost
71Gemini 2.5 Pro (higher rate after 200K)Google1M$1.25/1M$10.00/1M$0.00625$62.5Cost
72DeepSeek V4 ProDeepSeek1M$1.32/1M$3.96/1M$0.0033$33.00Cost
73Mistral Medium 3Mistral128K$1.50/1M$7.50/1M$0.00525$52.5Cost
74Grok 4.7 (higher rate after 200K)xAI500K$2.00/1M$6.00/1M$0.005$50.00Cost
75Grok 4.6 (higher rate after 200K)xAI500K$2.00/1M$6.00/1M$0.005$50.00Cost
76Mistral Large (2407)Mistral128K$2.00/1M$6.00/1M$0.005$50.00Cost
77GPT-4.1OpenAI1M$2.00/1M$8.00/1M$0.006$60.00Cost
78GPT-4.1 (2025-04-14)OpenAI1M$2.00/1M$8.00/1M$0.006$60.00Cost
79o3OpenAI200K$2.00/1M$8.00/1M$0.006$60.00Cost
80GPT-6 Sol (higher rate after 272K)OpenAI1.1M$2.00/1M$10.00/1M$0.007$70.00Cost
81Claude Sonnet 5Anthropic1M$2.00/1M$10.00/1M$0.007$70.00Cost
82Grok 2xAI131K$2.00/1M$10.00/1M$0.007$70.00Cost
83GPT-5.6 Terra (higher rate after 272K)OpenAI1.1M$2.00/1M$12.00/1M$0.008$80.00Cost
84Gemini 3.1 Pro (higher rate after 200K)Google1M$2.00/1M$12.00/1M$0.008$80.00Cost
85GPT-4oOpenAI128K$2.50/1M$10.00/1M$0.0075$75.00Cost
86GPT-4o (2024-08-06)OpenAI128K$2.50/1M$10.00/1M$0.0075$75.00Cost
87Command ACohere256K$2.50/1M$10.00/1M$0.0075$75.00Cost
88Command R+Cohere128K$2.50/1M$10.00/1M$0.0075$75.00Cost
89Command R+ (08-2024)Cohere128K$2.50/1M$10.00/1M$0.0075$75.00Cost
90DeepSeek R1 (Together)Together AI128K$3.00/1M$7.00/1M$0.0065$65.00Cost
91Claude 3.5 SonnetAnthropic200K$3.00/1M$15.00/1M$0.0105$105.00Cost
92Claude 3 SonnetAnthropic200K$3.00/1M$15.00/1M$0.0105$105.00Cost
93Claude Sonnet 4.6Anthropic1M$3.00/1M$15.00/1M$0.0105$105.00Cost
94GPT-5.6 Sol (higher rate after 272K)OpenAI1.1M$4.00/1M$20.00/1M$0.014$140.00Cost
95Claude Opus 5.5Anthropic1M$4.00/1M$20.00/1M$0.014$140.00Cost
96ChatGPT-4o LatestOpenAI128K$5.00/1M$15.00/1M$0.0125$125.00Cost
97Grok BetaxAI131K$5.00/1M$15.00/1M$0.0125$125.00Cost
98Claude Opus 5Anthropic1M$5.00/1M$25.00/1M$0.0175$175.00Cost
99Claude Opus 4.6Anthropic1M$5.00/1M$25.00/1M$0.0175$175.00Cost
100Claude Opus 4.8Anthropic1M$5.00/1M$25.00/1M$0.0175$175.00Cost
101Claude 2.1Anthropic200K$8.00/1M$24.00/1M$0.02$200.00Cost
102GPT-4 TurboOpenAI128K$10.00/1M$30.00/1M$0.025$250.00Cost
103GPT-6 Astra (higher rate after 272K)OpenAI1.1M$10.00/1M$50.00/1M$0.035$350.00Cost
104Claude Fable 5.1Anthropic1M$10.00/1M$50.00/1M$0.035$350.00Cost
105o1OpenAI200K$15.00/1M$60.00/1M$0.045$450.00Cost
106o1-previewOpenAI128K$15.00/1M$60.00/1M$0.045$450.00Cost
107Claude 3 OpusAnthropic200K$15.00/1M$75.00/1M$0.0525$525.00Cost

Related tools

Context window calculator · RAG cost · Cost calculator

Related guides

Context windows and overflow · Will my document fit? Context window calculator · Context budgeting for RAG · What happens when you exceed the context window · Why you must reserve output tokens

Sources and references

Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.

Pricing rank FAQ

What counts as long context here?
A published context window of 128,000 tokens or more.
Why mention a higher rate after a threshold?
Some models charge more once the prompt crosses a published token line. The example request is short on purpose. Use the calculator with your real document size to see the tier.
Is a larger window always cheaper?
No. A wide window can still have a high per token rate, and some models raise rates on long prompts.