Side by side pricing and context for GPT-4o mini and Gemini 2.5 Flash. Rates are curated catalog values with last verified dates on each model page.
GPT-4o mini
Inexpensive OpenAI chat model for customer support and simple tools. Popular cheap baseline. Pair with Exact token counts when budgeting.
Input $0.150 / Output $0.600 per 1M
Context 128K · Exact (o200k_base)
Gemini 2.5 Flash
Low-latency Gemini for chat and high throughput. Often the cost/speed sweet spot on Google. Approx token counts in the browser.
Input $0.300 / Output $2.50 per 1M
Context 1M · Approx (Gemini family)
Quick take
- GPT-4o mini is cheaper on input in this catalog snapshot.
- GPT-4o mini is cheaper on output in this catalog snapshot.
- Both last verified on their model pages (2026-08-02 / 2026-08-02).
Rate table
| Metric | GPT-4o mini | Gemini 2.5 Flash |
|---|---|---|
| Input / 1M | $0.150 | $0.300 |
| Output / 1M | $0.600 | $2.50 |
| Cached input / 1M | $0.075 | $0.030 |
| Context | 128K | 1M |
| Tokenizer | Exact (o200k_base) | Approx (Gemini family) |