Groq Pricing Calculator
Updated August 2026Estimate Groq LPU inference costs for gpt-oss-120b, gpt-oss-20b, Llama 3.3 70B, Llama 3.1 8B, and Qwen 3.6 27B.
Groq pricing: Groq pricing 2026: GPT OSS 120B $0.15/$0.60, GPT OSS 20B $0.075/$0.30, Whisper Turbo $0.04/hr. Llama 3.1 8B & 3.3 70B now Enterprise-only, 50% Batch + cached. Use the free calculator below to enter your usage and get an instant all-in monthly estimate — including overages — and see which plan is cheapest for your needs.
How Groq prices: Pure-usage per-token serverless + per-hour transcription + per-1M-character text-to-speech + per-use tools + sales-led enterprise.
Best fit at your usage: Llama 3.1 8B Instant
$0.27/mo— saves $1.08/mo vs your selection
Token Usage
Average number of tokens you send per API call
Average number of tokens the model generates per call
Request Volume
How many API calls do you make daily?
Compare Plans at Your Usage
Monthly Estimate
gpt-oss-120b
$1.35
per month
Cost Breakdown
Daily
$0.04
Annual
$16.20
Why Llama 3.1 8B Instant — $0.27/mo
- Llama 3.1 8B Instant at $0.27/mo.
- Input tokens is 56% of your bill — the lever that matters most.
Budget range
Plan for usage swings, not just today's estimate.
Conservative
$1.13
500 tokens/mo
Expected
$1.35
1K tokens/mo
Aggressive
$1.80
2K tokens/mo
Need a calculator like this on your pricing page?
Embed interactive pricing calculators on your website to help customers understand costs and boost conversions.
Get StartedAbout this Groq calculator
This calculator estimates your Groq cost from publicly available pricing. Actual costs may vary with your specific agreement, volume discounts, and usage patterns — always verify on the provider's official pricing page for the most current rates.
Groq pricing — frequently asked questions
How much does Groq cost per month? ▼
Groq has no monthly subscription fee — you pay only for the tokens you consume on serverless inference, plus per-hour audio transcription, per-use built-in tools, and per-session agentic infrastructure. A small chat application using GPT OSS 20B (the cheapest self-serve model as of 2026-08-26) at 30M input + 10M output tokens would cost roughly $5.25/month — Groq's per-token rates are among the lowest in the market for small-model inference. Note that Llama 3.1 8B Instant, which previously anchored this estimate at a lower rate, moved to Enterprise-only "Contact Sales" pricing on 2026-08-26 and is no longer self-serve.
What are Groq's per-token rates for popular models? ▼
As of 2026-08-11, Groq's dedicated pricing page (groq.com/pricing/) has been removed and redirects to the homepage — rates are now published only in the GroqCloud developer docs (console.groq.com/docs/models). As of 2026-08-26, Llama 3.1 8B Instant and Llama 3.3 70B Versatile — Groq's two Llama models — are marked "Enterprise" with "Contact Sales" in place of a published rate, and both are absent from the Free and Developer rate-limit tables. The self-serve catalog per 1M tokens (input/output) is now: GPT OSS 20B and Safety GPT OSS 20B (formerly "GPT OSS Safeguard 20B") $0.075/$0.30 at 1,000 T/SEC; GPT OSS 120B $0.15/$0.60 at 500 T/SEC; Qwen 3.6 27B $0.60/$3.00 at 500 T/SEC; Qwen 3.8-27B $0.80/$4.00 at 450 T/SEC (newly priced as of 2026-08-27); and two Preview moderation models, Llama Prompt Guard 2 22M and Prompt Guard 2 86M, at $0.03/$0.03 and $0.04/$0.04. Kimi K2 Instruct 0905 was quoted on the marketing page's prompt-caching table at $1.00 uncached / $0.50 cached input and $3.00 output; it has never appeared in the docs model catalog, and with the marketing page gone that rate can no longer be verified against any live Groq source — confirm availability directly in the console before relying on it. Check availability as well as rate before committing a workload: on 2026-07-21 Groq removed Llama 4 Scout (17Bx16E) and Qwen3 32B from both the rate card and the docs catalog without changing any retained price, and Groq's docs warn that models in the Preview tier "may be discontinued at short notice" — a tier which as of 2026-08-27 currently includes Safety GPT OSS 20B, Qwen 3.6 27B, Qwen 3.8-27B (newly priced 2026-08-27), both Orpheus text-to-speech voices, and the two Prompt Guard models.
Does Groq have a free tier? ▼
Yes — Groq's free tier offers rate-capped access to most models for experimentation, though as of 2026-08-26 that no longer includes Llama 3.1 8B Instant or Llama 3.3 70B Versatile, which moved to Enterprise-only "Contact Sales" pricing and dropped off the published Free Plan Limits table entirely. Throughput limits apply on the remaining models (low requests-per-minute and tokens-per-minute caps), with sustained production usage requiring a paid plan. Free tier credit-card-on-file is not required for evaluation.
How does Groq's LPU differ from GPU-based inference? ▼
Groq's LPU (Language Processing Unit) is bespoke silicon designed exclusively for inference. Unlike GPUs which use HBM (high-bandwidth memory) and branch prediction, the LPU uses on-die SRAM and deterministic execution. The trade-off is lower per-chip memory capacity but vastly higher single-stream throughput — enabling 800+ tokens/second on Llama 3 8B and 500+ TPS on 120B-class models.
Is this Groq pricing calculator free and accurate? ▼
Yes — it's completely free with no signup. It uses the latest publicly available Groq pricing (verified August 2026). Actual costs may vary with volume discounts or enterprise terms — always confirm on Groq's official pricing page.