Home / Pricing / Long replies

Cheapest models for long replies

Ranked by output rate per 1M tokens. The example uses 1000 input tokens and 2000 output tokens, which is closer to support chat and writing than a short completion.

132 models · rates checked . Price rank only, not a quality score.

Ranked by published rates

#ModelProviderContextInput / 1MOutput / 1MLong example10k requests
1Llama 3.2 1B (Groq)Groq128K$0.040/1M$0.040/1M$0.00012$1.2Cost
2Llama 3.2 3B (Groq)Groq128K$0.060/1M$0.060/1M$0.00018$1.8Cost
3Llama 3.2 3B (Together)Together AI128K$0.060/1M$0.060/1M$0.00018$1.8Cost
4Gemma 7B (Groq)Groq8K$0.070/1M$0.070/1M$0.00021$2.1Cost
5Llama 3.1 8B Instant (Groq)Groq128K$0.050/1M$0.080/1M$0.00021$2.1Cost
6Llama 3.2 3B (Fireworks)Fireworks128K$0.100/1M$0.100/1M$0.0003$3.00Cost
7Gemini 1.5 Flash-8BGoogle1M$0.037/1M$0.150/1M$0.000337$3.37Cost
8Command R7BCohere128K$0.037/1M$0.150/1M$0.000337$3.37Cost
9Ministral 8BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
10Mistral NemoMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
11Pixtral 12BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
12Llama 3.2 11B Vision (Groq)Groq128K$0.180/1M$0.180/1M$0.00054$5.4Cost
13Llama 3.1 8B (Together)Together AI128K$0.180/1M$0.180/1M$0.00054$5.4Cost
14Gemma 2 9B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
15Llama Guard 3 8B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
16Mistral 7B (Together)Together AI33K$0.200/1M$0.200/1M$0.0006$6.00Cost
17Llama 3.1 8B (Fireworks)Fireworks128K$0.200/1M$0.200/1M$0.0006$6.00Cost
18MythoMax L2 13B (Fireworks)Fireworks4K$0.200/1M$0.200/1M$0.0006$6.00Cost
19Mixtral 8x7B (Groq)Groq33K$0.240/1M$0.240/1M$0.00072$7.2Cost
20Mistral TinyMistral33K$0.250/1M$0.250/1M$0.00075$7.5Cost
21DeepSeek CoderDeepSeek128K$0.140/1M$0.280/1M$0.0007$7.00Cost
22DeepSeek VLDeepSeek4K$0.140/1M$0.280/1M$0.0007$7.00Cost
23Gemini 1.5 FlashGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
24Gemini 2.0 Flash-LiteGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
25Qwen2.5 7B (Together)Together AI33K$0.300/1M$0.300/1M$0.0009$9.00Cost
26Llama 4 Scout (Groq)Groq128K$0.110/1M$0.340/1M$0.00079$7.9Cost
27QwQ 32B (Groq)Groq128K$0.290/1M$0.390/1M$0.00107$10.7Cost
28GPT-5 NanoOpenAI128K$0.050/1M$0.400/1M$0.00085$8.5Cost
29GPT-4.1 NanoOpenAI1M$0.100/1M$0.400/1M$0.0009$9.00Cost
30Gemini 2.0 FlashGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
31Gemini 2.5 Flash-LiteGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
32Gemma 3 27BGoogle128K$0.200/1M$0.400/1M$0.001$10.00Cost
33DeepSeek ChatDeepSeek128K$0.280/1M$0.420/1M$0.00112$11.2Cost
34GPT-6 LunaOpenAI1.1M$0.100/1M$0.500/1M$0.0011$11.00Cost
35Grok 2 MinixAI131K$0.200/1M$0.500/1M$0.0012$12.00Cost
36Grok 3 MinixAI131K$0.300/1M$0.500/1M$0.0013$13.00Cost
37Qwen3 32B (Groq)Groq131K$0.290/1M$0.590/1M$0.00147$14.7Cost
38GPT-4o miniOpenAI128K$0.150/1M$0.600/1M$0.00135$13.5Cost
39Mistral SmallMistral128K$0.150/1M$0.600/1M$0.00135$13.5Cost
40Command RCohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
41Command R (08-2024)Cohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
42Mistral SabaMistral33K$0.200/1M$0.600/1M$0.0014$14.00Cost
43DeepSeek R1 Distill Qwen 32BDeepSeek128K$0.300/1M$0.600/1M$0.0015$15.00Cost
44Command LightCohere4K$0.300/1M$0.600/1M$0.0015$15.00Cost
45Open Mixtral 8x7BMistral33K$0.700/1M$0.700/1M$0.0021$21.00Cost
46Llama 3.3 70B (Groq)Groq128K$0.590/1M$0.790/1M$0.00217$21.7Cost
47Qwen2.5 Coder 32B (Together)Together AI33K$0.800/1M$0.800/1M$0.0024$24.00Cost
48Llama 4 Maverick (Together)Together AI1M$0.270/1M$0.850/1M$0.00197$19.7Cost
49Llama 4 Maverick (Fireworks)Fireworks1M$0.220/1M$0.880/1M$0.00198$19.8Cost
50Llama 3.1 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
51Llama 3.3 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
52CodestralMistral256K$0.300/1M$0.900/1M$0.0021$21.00Cost
53DeepSeek R1 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
54DeepSeek V3 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
55Llama 3.1 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
56Llama 3.3 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
57Mixtral 8x22B (Fireworks)Fireworks66K$0.900/1M$0.900/1M$0.0027$27.00Cost
58Qwen2.5 32B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
59Qwen2.5 72B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
60Llama 3.3 70B SpecDec (Groq)Groq8K$0.590/1M$0.990/1M$0.00257$25.7Cost
61DeepSeek R1 Distill Llama 70B (Groq)Groq128K$0.750/1M$0.990/1M$0.00273$27.3Cost
62DeepSeek V3DeepSeek128K$0.270/1M$1.10/1M$0.00247$24.7Cost
63GPT-5.6 LunaOpenAI1.1M$0.200/1M$1.20/1M$0.0026$26.00Cost
64DeepSeek V4 FlashDeepSeek1M$0.300/1M$1.20/1M$0.0027$27.00Cost
65Mixtral 8x22B (Together)Together AI66K$1.20/1M$1.20/1M$0.0036$36.00Cost
66Qwen2.5 72B (Together)Together AI33K$1.20/1M$1.20/1M$0.0036$36.00Cost
67Claude 3 HaikuAnthropic200K$0.250/1M$1.25/1M$0.00275$27.5Cost
68DeepSeek V3 (Together)Together AI128K$1.25/1M$1.25/1M$0.00375$37.5Cost
69Gemini 3.1 Flash-LiteGoogle1M$0.250/1M$1.50/1M$0.00325$32.5Cost
70GPT-3.5 TurboOpenAI16K$0.500/1M$1.50/1M$0.0035$35.00Cost
71Gemini 1.0 ProGoogle33K$0.500/1M$1.50/1M$0.0035$35.00Cost
72Mistral Large 3Mistral128K$0.500/1M$1.50/1M$0.0035$35.00Cost
73GPT-4.1 MiniOpenAI1M$0.400/1M$1.60/1M$0.0036$36.00Cost
74GPT-5 MiniOpenAI128K$0.250/1M$2.00/1M$0.00425$42.5Cost
75Command NightlyCohere128K$1.00/1M$2.00/1M$0.005$50.00Cost
76DeepSeek R1DeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
77DeepSeek ReasonerDeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
78Gemini 2.5 FlashGoogle1M$0.300/1M$2.50/1M$0.0053$53.00Cost
79Grok 3xAI131K$1.25/1M$2.50/1M$0.00625$62.5Cost
80Grok 4xAI256K$1.25/1M$2.50/1M$0.00625$62.5Cost
81Grok 4 FastxAI256K$1.25/1M$2.50/1M$0.00625$62.5Cost
82Gemini 3 FlashGoogle1M$0.500/1M$3.00/1M$0.0065$65.00Cost
83Gemini 3.8 FlashGoogle1.0M$0.750/1M$3.75/1M$0.00825$82.5Cost
84DeepSeek V4 ProDeepSeek1M$1.32/1M$3.96/1M$0.00924$92.4Cost
85Claude 3.5 HaikuAnthropic200K$0.800/1M$4.00/1M$0.0088$88.00Cost
86o1-miniOpenAI128K$1.10/1M$4.40/1M$0.0099$99.00Cost
87o3-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
88o4-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
89Claude Haiku 4.5Anthropic200K$1.00/1M$5.00/1M$0.011$110.00Cost
90Gemini 1.5 ProGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
91Gemini 2.0 Pro ExpGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
92Grok 4.7xAI500K$2.00/1M$6.00/1M$0.014$140.00Cost
93Grok 4.6xAI500K$2.00/1M$6.00/1M$0.014$140.00Cost
94Mistral Large (2407)Mistral128K$2.00/1M$6.00/1M$0.014$140.00Cost
95Open Mixtral 8x22BMistral66K$2.00/1M$6.00/1M$0.014$140.00Cost
96DeepSeek R1 (Together)Together AI128K$3.00/1M$7.00/1M$0.017$170.00Cost
97Mistral Medium 3Mistral128K$1.50/1M$7.50/1M$0.0165$165.00Cost
98GPT-4.1OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
99GPT-4.1 (2025-04-14)OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
100o3OpenAI200K$2.00/1M$8.00/1M$0.018$180.00Cost
101GPT-5OpenAI400K$1.25/1M$10.00/1M$0.02125$212.5Cost
102Gemini 2.5 ProGoogle1M$1.25/1M$10.00/1M$0.02125$212.5Cost
103GPT-6 SolOpenAI1.1M$2.00/1M$10.00/1M$0.022$220.00Cost
104Claude Sonnet 5Anthropic1M$2.00/1M$10.00/1M$0.022$220.00Cost
105Grok 2xAI131K$2.00/1M$10.00/1M$0.022$220.00Cost
106Grok 2 VisionxAI33K$2.00/1M$10.00/1M$0.022$220.00Cost
107GPT-4oOpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
108GPT-4o (2024-08-06)OpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
109Command ACohere256K$2.50/1M$10.00/1M$0.0225$225.00Cost
110Command R+Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
111Command R+ (08-2024)Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
112GPT-5.6 TerraOpenAI1.1M$2.00/1M$12.00/1M$0.026$260.00Cost
113Gemini 3.1 ProGoogle1M$2.00/1M$12.00/1M$0.026$260.00Cost
114Claude 3.5 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
115Claude 3 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
116Claude Sonnet 4.6Anthropic1M$3.00/1M$15.00/1M$0.033$330.00Cost
117ChatGPT-4o LatestOpenAI128K$5.00/1M$15.00/1M$0.035$350.00Cost
118Grok BetaxAI131K$5.00/1M$15.00/1M$0.035$350.00Cost
119Grok Vision BetaxAI8K$5.00/1M$15.00/1M$0.035$350.00Cost
120GPT-5.6 SolOpenAI1.1M$4.00/1M$20.00/1M$0.044$440.00Cost
121Claude Opus 5.5Anthropic1M$4.00/1M$20.00/1M$0.044$440.00Cost
122Claude 2.1Anthropic200K$8.00/1M$24.00/1M$0.056$560.00Cost
123Claude Opus 5Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
124Claude Opus 4.6Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
125Claude Opus 4.8Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
126GPT-4 TurboOpenAI128K$10.00/1M$30.00/1M$0.07$700.00Cost
127GPT-6 AstraOpenAI1.1M$10.00/1M$50.00/1M$0.11$1100.00Cost
128Claude Fable 5.1Anthropic1M$10.00/1M$50.00/1M$0.11$1100.00Cost
129o1OpenAI200K$15.00/1M$60.00/1M$0.135$1350.00Cost
130o1-previewOpenAI128K$15.00/1M$60.00/1M$0.135$1350.00Cost
131GPT-4OpenAI8K$30.00/1M$60.00/1M$0.15$1500.00Cost
132Claude 3 OpusAnthropic200K$15.00/1M$75.00/1M$0.165$1650.00Cost

Related tools

LLM cost calculator · Cheapest LLM API · Compare model prices · Cache savings · Batch pricing

Related guides

How LLM API pricing works · Prompt caching explained · Exact vs Approx

Sources and references

Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.

Pricing rank FAQ

Why rank by output?
Long answers are billed on output tokens, which are usually more expensive than input. A cheap input rate can still be a costly chat model.
What is the long example?
1000 input tokens and 2000 output tokens at standard rates, no cache and no batch.
Can I change the mix?
Yes. Open Cost on any row and set your own output size and monthly volume.