Side by side pricing and context for GPT-4.1 Mini and Gemini 2.5 Flash. Rates are curated catalog values with last verified dates on each model page.
GPT-4.1 Mini
Cheaper GPT-4.1 variant for the same long-context shape. Use for RAG and assistants where Mini quality is enough and you want Exact OpenAI token counts.
Input $0.400 / Output $1.60 per 1M
Context 1M · Exact (o200k_base)
Gemini 2.5 Flash
Low-latency Gemini for chat and high throughput. Often the cost/speed sweet spot on Google. Approx token counts in the browser.
Input $0.300 / Output $2.50 per 1M
Context 1M · Approx (Gemini family)
Quick take
- Gemini 2.5 Flash is cheaper on input in this catalog snapshot.
- GPT-4.1 Mini is cheaper on output in this catalog snapshot.
- Both last verified on their model pages (2026-08-02 / 2026-08-02).
Rate table
| Metric | GPT-4.1 Mini | Gemini 2.5 Flash |
|---|---|---|
| Input / 1M | $0.400 | $0.300 |
| Output / 1M | $1.60 | $2.50 |
| Cached input / 1M | $0.100 | $0.030 |
| Context | 1M | 1M |
| Tokenizer | Exact (o200k_base) | Approx (Gemini family) |