Cheapest LLM API in TokenCalculator means the lowest published input rate in this catalog, then a fixed example of 1000 input tokens and 500 output tokens. Rates checked 12 Aug 2026. This is a price rank, not a quality score.
Use the live rank
The table below is the same catalog as the tools app. Open any row in the calculator before you treat a list price as your bill.
What cheapest means here
Providers publish USD prices per 1 million tokens. Input and output are separate. A model can look cheap on input and expensive once replies get long. TokenCalculator therefore ranks two ways: input rate first, then an example request so ties are not random.
The example is intentionally short. It is a planning unit, not your production mix. Change output size and monthly volume in the cost calculator when you have a real prompt.
Snapshot from this catalog
Top rows by input rate as of 12 Aug 2026. Example cost uses 1000 input tokens and 500 output tokens, no cache and no batch.
| # | Model | Provider | Input / 1M | Example |
|---|---|---|---|---|
| 1 | Gemini 1.5 Flash-8B | $0.037/1M | $0.000112 | |
| 2 | Command R7B | Cohere | $0.037/1M | $0.000112 |
| 3 | Llama 3.2 1B (Groq) | Groq | $0.040/1M | $0.00006 |
| 4 | Llama 3.1 8B Instant (Groq) | Groq | $0.050/1M | $0.00009 |
| 5 | GPT-5 Nano | OpenAI | $0.050/1M | $0.00025 |
| 6 | Llama 3.2 3B (Groq) | Groq | $0.060/1M | $0.00009 |
| 7 | Llama 3.2 3B (Together) | Together AI | $0.060/1M | $0.00009 |
| 8 | Mistral Small | Mistral | $0.060/1M | $0.00015 |
Why a single input number lies
Two models with the same input rate can diverge once you include output length, cache hits, batch discounts, and tokenizer differences. Identical UTF-8 text does not produce identical token counts across families. Cost comparison has to hold the text constant and compare both counts and rates.
- Output tokens are usually several times more expensive than input.
- Host copies of the same open weights can undercut the first party list price.
- Long context tiers can raise the rate after a published token threshold.
- Promotional credits and private contracts are not in this catalog.
Worked example
Suppose a support bot sends 1000 input tokens and receives 500 output tokens, 10,000 times a month. The 10k requests column on the cheapest page is that math. If the same bot starts writing 2000 token answers, switch to the long replies rank. Output, not input, will decide the bill.
Host copies versus first party
Llama and DeepSeek rows appear on Groq, Together, Fireworks, and native APIs. The cheapest host table groups those copies. Speed and terms differ even when the weights look the same. Confirm the API model id you will actually call.
When cheap is the wrong pick
A nano or flash model can lose money if it retries, rambles, or fails a task that a mid tier model finishes in one pass. Price rank is the shortlist. The calculator plus a golden prompt is the decision.
Frequently asked questions
Is the cheapest model the best model?
No. This page ranks published token rates. Quality, latency, and tool use are out of scope here.
Are these live API prices?
No. Rates are curated from official provider pages and stamped with a checked date. Confirm critical budgets on the provider page.
Can I change the example mix?
Yes. Open Cost on any row and set your own output size, cache hit rate, batch flag, and monthly volume.
Next steps
- Cheapest frontier models: Flagship class only
- Cheapest long replies: Ranked by output rate
- OpenAI API pricing
- Claude API pricing
- Gemini API pricing
- GPT vs Claude vs Gemini cost