Ask
Packaging

Groq moves Llama 3.1 8B and Llama 3.3 70B to Enterprise-only pricing

Groq pricing

Groq's two Llama models lost their public per-token price and now show Contact Sales in the docs, dropping off the Free and Developer rate-limit tables entirely.

Before

Llama 3.1 8B Instant self-serve at $0.05 input / $0.08 output per 1M tokens (560 T/SEC); Llama 3.3 70B Versatile self-serve at $0.59 input / $0.79 output per 1M tokens (280 T/SEC); both listed with public Free and Developer plan rate limits.

After

Llama 3.1 8B Instant and Llama 3.3 70B Versatile marked "Enterprise" with "Contact Sales" in place of both price and rate limits; both models absent from the Free Plan Limits and Developer Plan Limits tables. GPT OSS 20B ($0.075/$0.30) is now the cheapest self-serve model.

Groq’s GroqCloud developer docs (console.groq.com/docs/models) gated its two Llama models behind an Enterprise sales conversation on this capture, the most consequential change since the dedicated groq.com/pricing/ page was retired on 2026-08-11. Llama 3.1 8B Instant and Llama 3.3 70B Versatile — previously the platform’s headline low-cost, high-throughput models — now show “Contact Sales” in both the price and rate-limit columns, and a cross-check against the docs’ Rate Limits page confirmed neither model appears on the Free Plan Limits or Developer Plan Limits tables anymore. Every other production price (GPT OSS 120B, GPT OSS 20B, Whisper, Whisper Turbo) held unchanged. The same capture added two small Preview-tier moderation models (Llama Prompt Guard 2 22M and Prompt Guard 2 86M, both priced at $0.03–$0.04 per 1M tokens) and formalized “Groq Compound” and “Compound Mini” as named Production Systems with no published per-token price.

From Groq's pricing timeline
Llama 3.1 8B Instant and Llama 3.3 70B Versatile Moved to Enterprise-Only Pricing

The two Llama models on Groq's rate card — previously self-serve at $0.05/$0.08 and $0.59/$0.79 per 1M tokens — are now marked "Enterprise" in the GroqCloud docs model catalog, with "Contact Sales" printed in place of both the price and rate-limit columns. Confirmed by a second surface: both models are also absent from the Free Plan Limits and Developer Plan Limits tables on the docs' Rate Limits page, where they previously had public RPM/TPM caps. GPT OSS 20B ($0.075/$0.30) is now the cheapest self-serve model on the card. Every other production price (GPT OSS 120B, Whisper, Orpheus, Qwen 3.6 27B) held unchanged. The docs catalog also grew two new Preview moderation models (Llama Prompt Guard 2 22M at $0.03/$0.03, Prompt Guard 2 86M at $0.04/$0.04 per 1M tokens) and formalized "Groq Compound" / "Compound Mini" as named Production Systems with no published per-token price.

About Groq
groq.com ↗

Groq runs a pure-usage per-token serverless inference API on its proprietary LPU silicon. As of 2026-08-11 the dedicated marketing pricing page (groq.com/pricing/) has been removed and redirects to the homepage; rates survive only in the GroqCloud developer docs. As of 2026-08-26, Llama 3.1 8B Instant and Llama 3.3 70B Versatile — previously $0.05/$0.08 and $0.59/$0.79 per 1M tokens — moved to Enterprise-only "Contact Sales" pricing and dropped off both the Free and Developer rate-limit tables. The self-serve catalog is now GPT OSS 20B and Safety GPT OSS 20B (formerly "GPT OSS Safeguard 20B") at $0.075/$0.30 (1,000 T/SEC), GPT OSS 120B at $0.15/$0.60 (500 T/SEC), Qwen 3.6 27B at $0.60/$3.00, and two Preview moderation models — Llama Prompt Guard 2 22M and Prompt Guard 2 86M — at $0.03/$0.03 and $0.04/$0.04 per 1M tokens. As of 2026-08-27, Qwen 3.8-27B joined the Preview catalog with a published price of $0.80/$4.00 per 1M tokens (450 T/SEC) — the only change versus the prior day's capture.

Free tier
Yes
Commits
Available
Transparency
public

Groq pricing history

  1. Aug 2026
    Qwen 3.8-27B Gains a Published Preview-Tier Price
  2. Aug 2026
    Llama 3.1 8B Instant and Llama 3.3 70B Versatile Moved to Enterprise-Only Pricing
  3. Aug 2026
    Public Pricing Page Removed — Rates Survive Only in Developer Docs
  4. Jul 2026
    Catalog Contraction: Two Models and the Browser Automation Tool Withdrawn
  5. Jul 2026
    Text-to-Speech SKU + Browser Automation Tool + New Models
Full Groq timeline

More Groq activity

All pricing activity