Llama 3.1 8B Instant (Groq) vs Llama 3.1 8B (Fireworks)

Side by side pricing and context for Llama 3.1 8B Instant (Groq) and Llama 3.1 8B (Fireworks). Rates are curated catalog values with last verified dates on each model page.

Llama 3.1 8B Instant (Groq)

Tiny Llama on Groq for the cheapest fast responses. Ideal for routing and simple replies where 8B is enough.

Input $0.050 / Output $0.080 per 1M

Context 128K · Approx

Llama 3.1 8B (Fireworks)

Small Llama on Fireworks for low-cost OpenAI-compatible inference. Good control for routing layers in multi-host stacks.

Input $0.200 / Output $0.200 per 1M

Context 128K · Approx


Quick take

  • Llama 3.1 8B Instant (Groq) is cheaper on input in this catalog snapshot.
  • Llama 3.1 8B Instant (Groq) is cheaper on output in this catalog snapshot.
  • Both last verified on their model pages (2026-08-02 / 2026-08-02).

Rate table

MetricLlama 3.1 8B Instant (Groq)Llama 3.1 8B (Fireworks)
Input / 1M$0.050$0.200
Output / 1M$0.080$0.200
Cached input / 1Mn/an/a
Context128K128K
TokenizerApproxApprox