AI Summary
About
Groq is a Mountain View-based AI inference company founded in April 2016 by Jonathan Ross — the original engineer behind Google’s Tensor Processing Unit (TPU). The product is GroqCloud, an inference API built on Groq’s proprietary LPU (Language Processing Unit) silicon. The LPU was designed exclusively for inference workloads: deterministic execution, on-die SRAM (no HBM), and architectural choices that prioritize single-stream throughput over GPU-style parallel batch efficiency. The result is published throughput of 800+ tokens/second on Llama 3 8B and 500+ TPS on 120B-class models, with per-token pricing that holds competitive with GPU-based serverless inference.
By 2026 Groq serves a mix of latency-sensitive production AI customers (voice assistants, real-time agentic systems, customer-facing chatbots) and high-throughput batch workloads. The company raised a $640M Series D in August 2024 led by BlackRock at $2.8B valuation, with Tiger Global, Samsung Catalyst, KDDI, and others participating; a 2025 round led by Saudi PIF brought valuation to $6.9B. The Saudi partnership was paired with a strategic commitment to build Groq’s largest data center in Saudi Arabia, signaling the company’s intent to compete on global inference capacity, not just per-token economics. Groq’s homepage banner (re-confirmed 2026-08-26) now reads “GROQ CLOSES $350 MILLION SERIES A, BUILDING THE WORLD’S LEADING AI INFERENCE CLOUD,” and its public site is repositioned around “Premier NeoCloud for fast inference” — messaging that emphasizes bulk inference capacity (with a new LPX platform pairing Groq’s LPU alongside NVIDIA’s next-generation GPUs) over the per-token developer pricing that previously anchored the homepage. (Note: this banner figure differs from the $650M figure this page previously reported for the same August 2026 raise — treat the banner text above as the currently observed source of record and confirm the round’s actual size and structure directly with Groq before citing it.)
As of 2026-08-11, groq.com/pricing/ redirects to the Groq homepage and no longer displays any pricing content — the site’s main navigation has dropped its “Pricing” link entirely (re-confirmed 2026-08-26). Most published per-token, per-hour, and per-character rates are unchanged and still available inside the GroqCloud developer docs (console.groq.com/docs/models), but as of 2026-08-26 two of Groq’s six production LLMs — Llama 3.1 8B Instant and Llama 3.3 70B Versatile — no longer carry a public per-token price at all: the docs mark both “Enterprise” and list “Contact Sales” in place of a rate, and both models have also dropped off the Free and Developer plan rate-limit tables entirely.
Groq competes with Fireworks AI, Together AI, Baseten, and Replicate for the managed-inference market, plus first-party providers (OpenAI, Anthropic) for general-purpose API customers. Its differentiation is bespoke LPU silicon (the only commercial inference platform built on inference-specific hardware), TPU-pioneer founder credibility (Jonathan Ross designed the original TPU), and the highest published throughput-per-token in the category — making Groq the canonical choice for latency-critical AI workloads where time-to-first-token matters more than raw cost.
Pricing summary : How Groq’s LPU + per-token + tools stack works
As of 2026-08-11, groq.com/pricing/ no longer exists as a pricing page — it 308-redirects to the Groq homepage, which carries no pricing content and no “Pricing” link in its nav (re-confirmed 2026-08-26). The rate card now lives only inside the GroqCloud developer docs model catalog (console.groq.com/docs/models), and as of 2026-08-26 that catalog itself changed: Llama 3.1 8B Instant and Llama 3.3 70B Versatile — which previously carried a public self-serve per-token rate (see Pricing evolution for the last-confirmed figures) — are now marked “Enterprise” with “Contact Sales” in place of both the price and the rate-limit columns, and both models have disappeared entirely from the Free Plan Limits and Developer Plan Limits tables on the docs’ Rate Limits page (console.groq.com/docs/rate-limits). The four LLMs still self-serve priced are GPT OSS 20B and GPT OSS 120B at $0.075/$0.30 and $0.15/$0.60 (1,000 and 500 T/SEC), plus three Preview-tier models — Safety GPT OSS 20B (renamed from “GPT OSS Safeguard 20B”; same model id and $0.075/$0.30 price), Qwen 3.6 27B at $0.60/$3.00, and — newly priced as of 2026-08-27 — Qwen 3.8-27B at $0.80/$4.00 (450 T/SEC). Two Preview-tier moderation models also joined the catalog on 2026-08-26: Llama Prompt Guard 2 22M at $0.03/$0.03 per 1M tokens and Prompt Guard 2 86M at $0.04/$0.04 per 1M tokens. Alongside the LLMs sit Whisper for audio transcription (per hour transcribed, 10-second minimum per request, unchanged) and Canopy Labs Orpheus text-to-speech (per 1M characters, unchanged). MiniMax M2.7 remains enterprise-only at Contact Sales, and the docs now also list “Groq Compound” and “Compound Mini” as first-class Production Systems (groq/compound, groq/compound-mini, ~450 T/SEC) with no published per-token price (”-” in the price column) — folding the previously ad-hoc built-in tools into named systems without disclosing a rate. Kimi K2 Instruct 0905’s rate remains unknown: it was quoted on the now-removed marketing page’s prompt-caching table (see Pricing evolution for the last-confirmed figures) and has never appeared in the docs model catalog.
Cached input tokens get a 50% discount on standard rates (e.g. GPT OSS 20B cached input at $0.037 vs $0.075 uncached, GPT OSS 120B at $0.075 vs $0.15) — the docs limit caching support to the three GPT-OSS models — and the Batch API offers a 50% discount for asynchronous workloads with a 24-hour to 7-day processing window. The two 50% discounts do not stack: the docs state batch requests already receive a 50% discount and “no additional discount is applied to cached tokens in batch requests”. Built-in agentic tools previously had transparent per-use pricing published as two families — Compound tools (basic search, advanced search, visit website, code execution) and GPT-OSS tools (browser search, visit website, Python code execution) — but as of 2026-08-11 the marketing pricing page that carried those rates is gone and the GroqCloud docs for each tool now defer back to that dead page instead of printing a number, so current tool pricing is unknown (see Pricing evolution for the last-confirmed 2026-07-21 figures). Free tier provides rate-capped access for evaluation on the still-self-serve models; Enterprise plans add custom SLAs, dedicated LPU capacity reservations, volume discount commits, VPC deployment, and — as of 2026-08-26 — the only path to Llama 3.1 8B Instant and Llama 3.3 70B Versatile.
This single-SKU plus per-tool structure — token / hour-audio / per-use-tool — was one of the most legible usage-based rate cards in AI inference middleware. Through 2026-07-21 the marketing pricing page explicitly framed this as “linear and predictable” pricing with no hidden costs; that copy, and the page carrying it, is gone as of 2026-08-11 — and as of 2026-08-26 the “predictable” self-serve catalog itself is two flagship models smaller.
What makes this different: Groq’s LPU silicon delivers throughput that GPU-based competitors cannot match at the same price point. Llama 3.1 8B runs at 560 T/SEC per the current GroqCloud docs (the retired marketing page had claimed 840 TPS as recently as 2026-07-21) — but as of 2026-08-26 that throughput figure is no longer paired with a public price: reaching Llama 3.1 8B or Llama 3.3 70B now requires an Enterprise sales conversation, so the self-serve latency advantage for multi-turn agentic workflows currently rests on the GPT OSS and Qwen models instead.
Pricing by product
Serverless per-token inference (with docs-published T/SEC)
The marketing pricing page these rows were originally read from (groq.com/pricing/) is gone as of 2026-08-11 — it now 308-redirects to the Groq homepage (re-confirmed 2026-08-26). Every price below is re-verified against the sole surviving public source, the GroqCloud docs model catalog. As of 2026-08-26, two of the six production models on this table lost their public price: Llama 3.1 8B Instant and Llama 3.3 70B Versatile are now marked “Enterprise” with “Contact Sales” printed in both the price and rate-limit columns, and the GroqCloud docs Rate Limits page confirms neither model appears on the Free Plan Limits or Developer Plan Limits tables anymore. Every other production price is unchanged since 2026-07-21.
| Model | Input ($/1M) | Output ($/1M) | Throughput (docs T/SEC) |
|---|---|---|---|
| GPT OSS 20B 128k | $0.075 | $0.30 | 1,000 |
| GPT OSS 120B 128k | $0.15 | $0.60 | 500 |
| Llama 3.3 70B Versatile 128k | Contact Sales (Enterprise) | Contact Sales (Enterprise) | 280 |
| Llama 3.1 8B Instant 128k | Contact Sales (Enterprise) | Contact Sales (Enterprise) | 560 |
Production Systems (new category as of 2026-08-26 — no published per-token price, shown as ”-” in the docs):
| System | Model ID | Throughput (docs T/SEC) | Price per 1M tokens |
|---|---|---|---|
| Groq Compound | groq/compound | 450 | Not published (”-“) |
| Compound Mini | groq/compound-mini | 450 | Not published (”-”) |
Preview models with a published price (subject to Groq’s short-notice discontinuation warning):
| Model | Input ($/1M) | Output ($/1M) | Throughput (docs T/SEC) |
|---|---|---|---|
| Safety GPT OSS 20B (was “GPT OSS Safeguard 20B”) | $0.075 | $0.30 | 1,000 |
| Qwen/Qwen3.6-27B | $0.60 | $3.00 | 500 |
| Qwen/Qwen3.8-27B | $0.80 | $4.00 | 450 |
| Llama Prompt Guard 2 22M | $0.03 | $0.03 | Not published |
| Prompt Guard 2 86M | $0.04 | $0.04 | Not published |
Llama Prompt Guard 2 22M and Prompt Guard 2 86M are new as of the 2026-08-26 capture — both are small guard/moderation models, not general chat models. As of 2026-08-27, qwen/qwen3.8-27b gained a published price — $0.80 input / $4.00 output per 1M tokens at 450 T/SEC, 131,042-token context window, 16,384 max completion tokens, 20 MB max file size — joining the Preview Models table on the GroqCloud docs model catalog. It had previously only appeared on the Rate Limits page (30 RPM / 8K TPM / 2M TPD on the Developer plan) with no price or context-window spec in the models catalog; this is now resolved. This was the only change between the 2026-08-26 and 2026-08-27 captures — every other price, model, and page held exactly.
Enterprise-only model (contact sales for pricing): MiniMax M2.7 (260 T/SEC, unchanged).
Kimi K2 Instruct 0905’s price no longer has a live public source. moonshotai/kimi-k2-instruct-0905 was printed on the now-removed marketing pricing page’s Prompt Caching table (last confirmed 2026-07-21 — see Pricing evolution for the specific figures); it has never appeared in the GroqCloud docs model catalog. With the marketing page gone as of 2026-08-11, that rate cannot currently be re-verified against any live Groq page — confirm current pricing and availability directly in the console before building on it.
No longer on the public rate card as of 2026-07-21 (unchanged as of 2026-08-26): Llama 4 Scout (17Bx16E) 128k, Qwen3 32B 131k, and the enterprise-only Qwen3-VL 32B — all three were listed the week before and remain absent from the docs model catalog. Groq’s docs warn that Preview models “may be discontinued at short notice”.
Text-to-speech (per 1M characters)
| Model | Rate per 1M characters | Characters/s |
|---|---|---|
| Canopy Labs Orpheus English | $22.00 | 100 |
| Canopy Labs Orpheus Arabic Saudi | $40.00 | 100 |
Audio transcription (per hour, 10-second minimum)
| Model | Rate per hour | Speed factor | Notes |
|---|---|---|---|
| Whisper V3 Large | $0.111 | 217x | Higher accuracy, slower |
| Whisper Large v3 Turbo | $0.04 | 228x | Faster, slightly lower accuracy |
Built-in tools (Compound)
Rates below are unknown as of 2026-08-11. The GroqCloud docs for these tools (Web Search, Code Execution) each explicitly say “Please see the Pricing page for more information” rather than printing a rate — and that Pricing page (groq.com/pricing/) now 308-redirects to the homepage with no pricing content. These tools’ rates have no live public source anywhere on groq.com or console.groq.com as of this capture; see Pricing evolution for the last-confirmed figures (dated 2026-07-21) and confirm current pricing directly with Groq before budgeting against it.
| Tool | Rate | Unit | API parameter |
|---|---|---|---|
| Basic Search | Unknown | per 1,000 requests | web_search |
| Advanced Search | Unknown | per 1,000 requests | web_search |
| Visit Website | Unknown | per 1,000 requests | visit_website |
| Code Execution | Unknown | per hour | code_interpreter |
Built-in tools (GPT-OSS)
| Tool | Rate | Unit | API parameter |
|---|---|---|---|
| Browser Search — Basic Search | Unknown | per 1,000 requests | browser_search - browser.search |
| Browser Search — Visit Website | Unknown | per 1,000 requests | browser_search - browser.open |
| Code Execution — Python | Unknown | per hour | code_interpreter - python |
Prompt caching (cached input discount)
| Model | Uncached input ($/1M) | Cached input ($/1M) | Output ($/1M) | Source |
|---|---|---|---|---|
| Kimi K2 Instruct 0905 | Unknown | Unknown | Unknown | No live source (see Pricing evolution for the last-confirmed figures) |
| GPT OSS 120B | $0.15 | $0.075 | $0.60 | GroqCloud docs — per-model detail page prints this cached rate directly |
| GPT OSS 20B | $0.075 | $0.037 | $0.30 | GroqCloud docs — per-model detail page prints this cached rate directly (rounded down from a naive 50%-off calculation, not equal to it) |
Prompt caching is not catalog-wide: the docs (last checked 2026-07-21) list caching support for exactly three models — GPT OSS 20B, GPT OSS 120B, and GPT OSS Safeguard 20B (the Models catalog now displays this third model as “Safety GPT OSS 20B” as of 2026-08-26, same model id openai/gpt-oss-safeguard-20b) — and the model catalog table itself confirms only the 50%-off mechanism, not a printed per-model cached dollar rate. Each model’s own GroqCloud docs detail page (e.g. console.groq.com/docs/model/openai/gpt-oss-20b) does print a cached-input price, and for GPT OSS 20B that printed figure is $0.037 — the exact 50%-off arithmetic on the $0.075 input rate would suggest a slightly higher figure, but $0.037 is the number Groq actually publishes, so cite that. Kimi K2’s caching rate sat on the now-removed marketing page’s caching table, never in the caching docs, and as of 2026-08-26 has no remaining live source.
Discount mechanics
| Mechanism | Discount | Applies to |
|---|---|---|
| Cached input tokens | 50% off input rate | Only the three GPT-OSS models (20B, 120B, Safeguard 20B) per the docs — no extra fee; discount applies only on a cache hit |
| Batch API | 50% off standard rate | Asynchronous workloads (24-hour to 7-day processing window); does not stack with the caching discount |
Sales motions across products: PLG / self-serve for Free and Developer tier (GPT OSS, Qwen, Whisper, Orpheus, Prompt Guard); sales-led for Enterprise annual commits, dedicated LPU capacity, VPC deployments, MiniMax M2.7, and — as of 2026-08-26 — Llama 3.1 8B Instant and Llama 3.3 70B Versatile.
Hidden costs : What Groq customers actually pay beyond the per-token rate
Archetype A: Voice assistant startup on Llama 3.1 8B Instant
A latency-critical voice assistant startup serving ~200K requests/day at average 1K input + 250 output tokens, using Whisper Turbo for upstream audio transcription:
| Line item | Monthly cost |
|---|---|
| Input tokens (6M/day × 30 = 180M, Llama 8B at $0.05/1M) | $9 |
| Output tokens (1.5M/day × 30 = 45M, Llama 8B at $0.08/1M) | $3.60 |
| Whisper Turbo transcription (avg 4 hours audio/day × 30 × $0.04) | $4.80 |
| Web search tool (occasional, 5K requests/mo × $0.005) | $25 |
| Estimated total | ~$42/month |
For latency-critical voice workloads, Groq’s per-token rates plus Whisper Turbo deliver one of the lowest published total costs in the category. The same workload on a comparable GPU-based platform would typically cost 2–4× more — and would not match the latency.
Archetype B: Mid-market team running mixed Llama 70B + GPT OSS 120B with agentic tools
A 50-person team building an internal AI assistant using Llama 3.3 70B Versatile for general queries and GPT OSS 120B for reasoning-heavy tasks, with frequent agentic tool usage:
| Line item | Monthly cost |
|---|---|
| Llama 3.3 70B (30M input + 10M output) | $17.70 + $7.90 = $25.60 |
| GPT OSS 120B (10M input + 5M output) | $1.50 + $3.00 = $4.50 |
| Cached input savings (40% cache hit on system prompts) | -$7 |
| Web search ($5/1K × 40K requests/mo) | $200 |
| Code execution (50 hours × $0.18) | $9 |
| Estimated total | ~$232/month |
Built-in tool usage dominates the bill at this scale — and the transparency of the tool rate card lets finance teams forecast tool spend accurately. Most platforms either bundle tools into opaque tier pricing or force customers onto third-party tool providers, making forecasting harder.
Want to estimate your own Groq bill? Use the Groq pricing calculator to model token spend by model, audio transcription hours, and built-in tool usage.
Pricing evolution : Groq’s pricing history from LPU prototype to commercial inference category
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2016 Q2 | 0 | 1 | Groq founded; LPU silicon design begins |
| 2023 Q1 | 0 | 1 | GroqCloud public launch with Llama 2 + Mistral |
| 2024 Q1 | 0 | 1 | Llama 3 day-one support; published TPS rates |
| 2024 Q3 | 0 | 0 | Series D ($640M) at $2.8B valuation |
| 2025 Q1 | 0 | 1 | Whisper Large v3 + Turbo audio SKUs launched |
| 2025 Q2 | 1 | 1 | Batch API + cached input 50% discounts launched |
| 2025 Q3 | 0 | 0 | Saudi PIF investment at $6.9B valuation |
| 2026 Q1 | 0 | 1 | Built-in tools per-use pricing published |
| 2026 Q3 | 0 | 6 | 07-14 Orpheus text-to-speech SKU (per 1M chars), Browser Automation tool ($0.08/hr), new models (GPT OSS Safeguard 20B, Qwen 3.6 27B, Kimi K2 + enterprise-only Minimax M2.5, Qwen3-VL 32B); 07-21 five SKUs withdrawn — Llama 4 Scout, Qwen3 32B, the Browser Automation tool, Qwen3-VL 32B, and Minimax M2.5 (replaced by M2.7) — with every retained price unchanged; 08-26 Llama 3.1 8B Instant and Llama 3.3 70B Versatile moved from published self-serve rates to Enterprise-only “Contact Sales”, two new Preview moderation models added (Llama Prompt Guard 2 22M, Prompt Guard 2 86M), and “Groq Compound” / “Compound Mini” formalized as unpriced Production Systems; 08-27 Qwen/Qwen3.8-27B gained a published Preview-tier price ($0.80/$4.00, 450 T/SEC) — the only change versus the prior day’s capture |
Tracked range: 2016 Q2–2026 Q3. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).
Notable changes
- 2023-03-22 — GroqCloud public launch with Llama 2 + Mistral; LPU silicon as the structural differentiator.
- 2024-02-19 — Llama 3 day-one support at 800+ TPS for the 8B variant; established Groq as the throughput leader.
- 2025-03-04 — Whisper Large v3 ($0.111/hr) and Whisper Large v3 Turbo ($0.04/hr) for audio transcription with 10-second minimum billing.
- 2025-05-13 — Batch API and cached input 50% discounts launched; brought Groq to parity with Fireworks and Anthropic on discount mechanics.
- 2026-01-22 — Built-in tools per-use pricing published: web search $5–$8/1K, website visits $1/1K, code execution $0.18/hour.
- 2026-07-14 — Text-to-speech launched as a new billing dimension: Canopy Labs Orpheus V1 English at $22.00/1M characters, Orpheus Arabic (Saudi) at $40.00/1M characters — pairing with Whisper ASR to close a full audio round-trip on the LPU. A Browser Automation built-in tool joined the per-use rate card at $0.08/hour, and the model catalog grew (GPT OSS Safeguard 20B $0.075/$0.30, Qwen 3.6 27B $0.60/$3.00, Kimi K2 $1.00/$3.00 with $0.50 cached, plus enterprise-only Minimax M2.5 and Qwen3-VL 32B). Core per-token rates for Llama, GPT OSS, and Qwen3 held stable — a packaging-and-catalog expansion, not a repricing.
- 2026-07-21 — Catalog contraction with no repricing. Llama 4 Scout (17Bx16E, $0.11/$0.34) and Qwen3 32B ($0.29/$0.59) left both the rate card and the docs model catalog, the Browser Automation tool ($0.08/hour) was pulled from Built-In Tools (Compound) seven days after launching, Qwen3-VL 32B was dropped from the enterprise-only list, and Minimax M2.5 was replaced by M2.7. Every retained price held exactly — the published catalog shrank from eight priced LLMs to six and from five Compound tools to four.
- 2026-08-26 — Llama 3.1 8B Instant and Llama 3.3 70B Versatile — previously self-serve at $0.05/$0.08 and $0.59/$0.79 per 1M tokens — moved to Enterprise-only “Contact Sales” pricing in the docs model catalog, and both dropped off the Free Plan Limits and Developer Plan Limits tables on the docs’ Rate Limits page. GPT OSS 20B ($0.075/$0.30) became the cheapest self-serve model. Two new Preview moderation models joined the catalog (Llama Prompt Guard 2 22M at $0.03/$0.03, Prompt Guard 2 86M at $0.04/$0.04 per 1M tokens), “GPT OSS Safeguard 20B” was redisplayed as “Safety GPT OSS 20B” (same model id, same price), and “Groq Compound” / “Compound Mini” were formalized as named Production Systems with no published per-token price.
- 2026-08-27 — Qwen/Qwen3.8-27B gained a published price in the Preview Models table: $0.80 input / $4.00 output per 1M tokens at 450 T/SEC (131,042-token context window, 16,384 max completion tokens, 20 MB max file size). The model had appeared on the Rate Limits docs page since at least 2026-08-26 with no price in the models catalog; this capture is the first to show a price. No other model, rate, or page changed versus the prior day’s capture.
The July 2026 catalog contraction in detail
Groq’s two July captures bracket a complete expand-and-retract cycle. On 2026-07-14 the rate card grew to eight priced LLMs, five Compound tools, and an entirely new text-to-speech line. By 2026-07-21 it was back to six LLMs and four Compound tools, and three of the withdrawn SKUs — Browser Automation, Minimax M2.5, and Qwen3-VL 32B — had been published for exactly one week.
The withdrawals were not arbitrary. Both delisted models, Llama 4 Scout (17Bx16E) and Qwen3 32B, sat in the docs’ Preview tier, which carries an explicit warning that preview models “are intended for evaluation purposes only and should not be used in production environments as they may be discontinued at short notice.” The pricing page draws no such line: Preview and Production models are printed in one undifferentiated rate card, at the same decimal precision, with the same throughput figures beside them. A buyer reading groq.com/pricing cannot tell which half of the catalog carries a discontinuation warning without cross-referencing the docs — and that is not a small half. As of 2026-07-21, GPT OSS Safeguard 20B, Qwen 3.6 27B, and both Orpheus text-to-speech voices are all Preview-tier, meaning a third of the priced LLM catalog and the entire TTS product line sit in the tier the July withdrawals came from.
For anyone who had already built on the withdrawn models, the cost impact arrives as substitution rather than as a price change. The nearest retained replacement for Llama 4 Scout ($0.11/$0.34) is GPT OSS 120B at $0.15/$0.60 — 36% more on input and 76% more on output. The only remaining Qwen is Qwen 3.6 27B at $0.60/$3.00, against Qwen3 32B’s $0.29/$0.59: 2.1× on input and 5.1× on output. Groq did not move a single published number, and an output-heavy Qwen workload’s bill can still roughly quintuple. That is the honest reading of “no price changes” inside a contracting catalog — and it is the reason the forecasting exposure here sits in availability rather than in rate.
What’s unique : Groq’s distinctive pricing mechanics
1. Throughput (TPS) published alongside per-token rates. Groq is one of the only inference platforms that publishes throughput (tokens per second) next to every per-token rate. For latency-critical workloads (voice agents, real-time interactive UIs), this transparency lets customers evaluate the latency-cost trade-off in a single view rather than running benchmark calls. Most competitors hide TPS behind documentation pages or marketing claims.
2. Bespoke LPU silicon as the structural cost moat. The LPU is the only commercial inference platform built on inference-specific silicon (rather than repurposed GPU). This delivers throughput that GPU-based competitors cannot match at comparable prices — currently best evidenced by GPT OSS 20B at $0.075/$0.30 running 1,000 T/SEC, a rate genuinely hard to replicate without similar silicon investment. The moat’s original headline proof point has weakened, though: Llama 3.1 8B Instant, long cited here at $0.05/$0.08 and 800+ TPS, moved to Enterprise-only “Contact Sales” pricing on 2026-08-26, and the GroqCloud docs now show it running 560 T/SEC — below the 840 TPS the retired marketing page had claimed as recently as 2026-07-21. The silicon advantage is still real; the number a self-serve buyer can actually check has changed.
3. Per-use agentic tool pricing in line-item form. Web search ($5–$8/1K), website visits ($1/1K), and code execution ($0.18/hour) were priced individually rather than bundled into opaque tiers through 2026-07-21, and for finance teams forecasting agent-loop costs a transparent tool rate card reduces budget uncertainty in a way that bundled-tools competitors cannot match. The line-item structure cuts both ways: because every tool is its own row, Groq can add or withdraw one without disturbing anything else — Browser Automation appeared at $0.08/hour on 2026-07-14 and was gone by 2026-07-21, never on the rate card long enough to become a forecastable line. As of 2026-08-11 the mechanism went further than a withdrawn line item: the marketing pricing page carrying every tool rate was itself removed, and the GroqCloud docs for each tool now just point back to that dead page instead of printing a number — so the “transparent tool rate card” this point describes currently has no live public source at all.
4. Full audio round-trip (Whisper ASR in, Orpheus TTS out) on one silicon platform. The 2.8× price spread ($0.111/hr vs $0.04/hr) between Whisper Large v3 and Whisper Large v3 Turbo lets customers pick accuracy-vs-speed explicitly on transcription. As of 2026-07-14, the Orpheus text-to-speech launch ($22.00/1M chars English, $40.00/1M chars Arabic) closes the loop — speech-to-text and text-to-speech now run on the same LPU, so a low-latency voice agent can do both legs without leaving Groq or bolting on a second vendor. Note that TTS introduces a fourth billing unit (per 1M characters) alongside per-token, per-hour, and per-request — each modality is metered in its own native unit rather than being force-fit into tokens. The caveat is durability rather than price: both Orpheus voices are listed under Preview in the docs — the same tier the models withdrawn on 2026-07-21 came from — so the return leg of that round-trip is the newest and least-committed part of the catalog.
5. Cached input + Batch API discount mechanics in standard rate card. Both discounts apply at 50% off base rates, bringing Groq to parity with Fireworks, OpenAI, and Anthropic. The discount mechanics are the same across all eligible models — making them legible to customers without requiring per-model discount lookup. This discount-as-default architecture is becoming canonical for token-priced inference.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| TPS published alongside per-token rates in the GroqCloud docs — best latency-cost transparency in category | As of 2026-08-11, that transparency lives only in developer docs — the public marketing pricing page (groq.com/pricing/) was removed and now redirects to the homepage, raising the discovery bar for non-technical buyers |
| Bespoke LPU silicon delivers 4–10× higher TPS than GPU competitors at similar rates (GPT OSS 20B: 1,000 T/SEC at $0.075/$0.30) | Two of the six production LLMs — Llama 3.1 8B Instant and Llama 3.3 70B Versatile, previously the cheapest and fastest self-serve options — moved to Enterprise-only “Contact Sales” pricing on 2026-08-26 |
| Docs model catalog labels a distinct “Preview Models” table, and its Preview tier keeps taking new priced entries (Qwen/Qwen3.8-27B at $0.80/$4.00, added 2026-08-27) rather than sitting static | Built-in agentic tool pricing (web search, visits, code exec) has had no live public source since 2026-08-11 — the docs for each tool defer back to the now-dead marketing page instead of printing a rate |
| Full audio round-trip on one platform — Whisper ASR in (two variants, a 2.8× accuracy-vs-speed spread), Orpheus TTS out, no second speech vendor | Audio is Whisper-only for ASR and Orpheus-only for TTS (two voices/languages), and both Orpheus SKUs sit in the Preview tier |
| TPU-pioneer founder credibility (Jonathan Ross) | Free-tier rate caps are tight for early-stage evaluation, and no fine-tuning SKU is published |
| Every retained per-token, per-hour, and per-character rate has held exactly since 2026-07-21 — no repricing, only availability changes | LPU silicon production scaling is a structural risk — supply must keep pace with demand |
Billing UX : Groq’s account controls and payment experience
- Self-serve signup — Sign up at
console.groq.comwith email; free tier access enabled immediately, no credit card required. - Per-token usage console — Real-time view of per-model token consumption (input, output, cached) and per-request latency metrics.
- TPS transparency in console — Per-request TPS shown alongside latency, letting developers verify throughput SLAs in real time.
- Spend Limits — A named console control that “blocks API access when you reach your monthly cap,” with recommended alerts at 50%, 75% and 90% of the limit and a documented 10–15 minute tracking delay.
- Projects — Per-project cost centers for tracking spending, usage patterns, and resource consumption at project granularity (showback / chargeback without a third-party FinOps tool).
- Model Permissions — Console-level control over which models an org or project may call, so preview-model spend can be fenced off from production keys.
- Service Tiers (Performance / Flex / Batch) — Documented processing tiers that trade latency for price; Batch Processing runs asynchronously at 50% lower cost with a 24-hour to 7-day window and no impact on standard rate limits.
- Rate Limits page — Per-model RPM/RPD/TPM/TPD limits published per plan, with a visible “Free Plan Limits” / “Developer Plan Limits” toggle so both tiers’ caps are legible before signup (e.g. GPT OSS 120B: 30 RPM / 8K TPM / 200K TPD on Free, 1K RPM / 500K TPM / 250K TPD on Developer). As of 2026-08-26 neither Llama 3.1 8B Instant nor Llama 3.3 70B Versatile appears on either table — consistent with both models moving to Enterprise/Contact-Sales-only on the Models page.
- Progressive invoicing thresholds — Self-serve accounts are charged automatically as cumulative usage crosses a documented ladder of spend thresholds, after which the account moves to monthly arrears (India uses its own threshold ladder).
- Payment methods — Credit cards (Visa, MasterCard, American Express, Discover), US bank accounts, and SEPA debit on self-serve; wire/invoice billing on Enterprise.
- Prometheus Metrics — Documented metrics endpoint for exporting usage and latency telemetry into a customer’s own observability stack.
- Cached input transparency — Prompt-caching rates are published per model and the page states there is “no extra fee for the caching feature itself. The discount only applies when a cache hit occurs.” The docs scope the feature to three models (GPT OSS 20B, 120B, and Safeguard 20B) and confirm the caching discount does not stack with the Batch API discount.
Strategic wins : Why Groq’s pricing decisions worked
1. Bespoke LPU silicon as the structural cost moat
By investing in inference-specific silicon rather than repurposing GPU, Groq created a structural cost advantage that GPU-based competitors cannot replicate without comparable silicon investment. GPT OSS 20B running 1,000 T/SEC at $0.075/$0.30 is the current canonical proof that inference-specific silicon outperforms training-derived GPU silicon for sustained inference workloads — a title that used to belong to Llama 3.1 8B Instant at 800+ TPS and $0.05/$0.08. That original proof point has moved: on 2026-08-26 Llama 3.1 8B Instant went Enterprise-only (“Contact Sales”), and the docs now show it running 560 T/SEC, well below the 840 TPS the retired marketing page had claimed through 2026-07-21. The silicon moat itself hasn’t weakened; the model Groq lets self-serve buyers point to as evidence of it has.
2. TPS published alongside per-token rates removed evaluation friction
By publishing throughput (TPS) next to every per-token rate, Groq made the latency-cost trade-off legible without benchmark calls. For latency-critical workloads (voice agents, interactive UIs), this transparency converts more self-serve customers and reduces sales-led overhead. Most competitors hide TPS behind benchmark blog posts or marketing claims.
3. TPU-pioneer founder credibility as the silicon-trust anchor
Jonathan Ross’s role as the original Google TPU engineer gives Groq unusual credibility on silicon architecture — making the LPU’s throughput claims believable in a way that pure-engineering teams cannot replicate. For enterprise procurement leaders evaluating bespoke-silicon inference, the founder credentials act as a trust multiplier that distinguishes Groq from generic inference middleware.
4. Per-use agentic tool pricing preempted bundled-tool competitor lock-in
By publishing per-use rates for web search, website visits, and code execution as line items on the rate card, Groq let customers forecast agent-loop costs without negotiating contracts or accepting opaque tool bundling. Through 2026-07-21 this transparent agentic pricing was a meaningful competitive advantage as agentic workflows scale and tool usage becomes the dominant cost line. The 2026-07-21 withdrawal of Browser Automation showed one flip side of the same design — a line item is as cheap to remove as it was to add — and customers could at least watch the $0.08/hour tool arrive and leave rather than absorbing a silent change inside a bundled tier price. As of 2026-08-11 the design showed a second flip side: the marketing page carrying every tool rate was removed entirely, and the docs for each tool now defer back to that dead page instead of printing a number. A line-item rate card only beats bundling while the rate card is published somewhere live — right now it isn’t, so this win is on hold pending re-verification with Groq directly.
Areas to improve : Gaps in Groq’s pricing approach
1. Cached input + Batch API discounts should stack
Fireworks lets cached input (50%) and Batch API (50%) discounts stack — making batched cached workloads land at 25% of standard. Groq’s discounts do not stack — customers must pick one. Adding stacking would close a meaningful competitive gap with Fireworks for RAG and agent-loop workloads at scale.
2. No published fine-tuning SKU
Customers wanting custom-trained models must use Fireworks, Together, Anyscale, or first-party model providers. Adding a fine-tuning SKU (even at limited model coverage initially) would let Groq capture the model-customization workload that currently goes elsewhere. The LPU’s deterministic execution may make fine-tuning architecturally awkward, but customers want the option.
3. Free tier rate caps may be too restrictive
Free tier rate limits are tight enough that meaningful evaluation requires a credit card and Developer-tier upgrade. Loosening free-tier caps (or offering a $5–$10 trial credit) would let self-serve developers reach proof-of-value without paid commitment and likely accelerate self-serve revenue.
4. Audio transcription only via Whisper
Whisper is the dominant open-source ASR model, but customers wanting Deepgram-comparable accuracy or specialized vertical models (medical, legal) must use other providers. Adding a second ASR model option (or a transcription quality tier above Whisper) would broaden the audio TAM and reduce vendor sprawl for audio-heavy workloads.
5. Groq needs a public rate card somewhere outside the developer docs
Through 2026-07-21 the complaint here was that groq.com/pricing printed Production and Preview models in one undifferentiated table despite the docs’ own warning that Preview models “may be discontinued at short notice” — on 2026-07-21 that gap turned concrete when Llama 4 Scout and Qwen3 32B, both Preview, disappeared from the rate card with the nearest retained Qwen substitute costing 5.1× more on output. As of 2026-08-11 that specific complaint is largely moot: the marketing pricing page is gone, and the sole surviving rate card — the GroqCloud docs model catalog — already segments Production Systems from a distinctly labeled “Preview Models” table, which continues taking new entries (Qwen/Qwen3.8-27B priced 2026-08-27) as well as losing them. The gap that replaced it is worse for evaluation-stage buyers: there is no longer any public, marketing-facing page that states Groq’s rates at all, and as of 2026-08-26 two of six production LLMs (Llama 3.1 8B Instant, Llama 3.3 70B Versatile) also require an Enterprise sales conversation to get a price. A buyer researching Groq without console access today cannot find a published rate anywhere on groq.com — reinstating even a static, docs-linking pricing page would let prospects price migration and discovery risk before they talk to sales.
Monetization stack & signals : how Groq builds & buys its revenue engine
Buys 0 Builds 3
Groq builds the revenue engine behind its own price: the token meter, progressive/region-aware usage billing with org-wide spend caps, and project-level cost allocation are all first-party in the GroqCloud console — no metering or FinOps vendor named. The only unconfirmed piece is the PSP behind self-serve card/ACH/SEPA checkout.
-
“Groq meters consumption by tokens across models; customers "monitor your usage and charges in near real-time directly within your Groq Cloud dashboard" (spend tracking refreshes "every 10-15 minutes" per the spend-limits docs) — a first-party meter, no third-party metering vendor named.”
-
“Bespoke billing logic: an invoice "is automatically triggered and payment is deducted when your cumulative usage reaches specific thresholds: $1, $10, $100, $500, and $1,000," after which the account moves to monthly arrears; India uses "$1, $10, and then $100 recurring," with a $0.50 minimum — region-specific rules a generic billing SKU does not ship.”
-
“Projects give per-project cost centers: "track spending, usage patterns, and resource consumption at the project level. This granular visibility enables accurate cost allocation, budget management, and ROI analysis" — first-party showback/chargeback, not a bought FinOps tool.”
-
“Self-serve checkout accepts "credit cards (Visa, MasterCard, American Express, Discover), United States bank accounts, and SEPA debit accounts" — card + ACH + SEPA rails plus India recurring thresholds imply a payment orchestrator, but no PSP (Stripe etc.) is named anywhere in the docs.”
Signals reviewed · derived from product docs
Key takeaways
-
Bespoke silicon as the structural moat — but the proof point has moved. Groq’s LPU is the only commercial inference platform built on inference-specific silicon, and the throughput advantage is genuinely hard to replicate without comparable silicon investment — a durable competitive moat that pure-software optimization cannot match. The model buyers used to cite as evidence, Llama 3.1 8B Instant at 800+ TPS, moved to Enterprise-only “Contact Sales” pricing on 2026-08-26 and now shows 560 T/SEC in the docs (below the 840 TPS the retired marketing page had claimed through 2026-07-21); the current self-serve proof point is GPT OSS 20B at 1,000 T/SEC and $0.075/$0.30.
-
Throughput transparency converts latency-critical buyers. Publishing TPS alongside per-token rates makes the latency-cost trade-off legible for voice-agent, interactive-UI, and real-time agentic workloads where time-to-first-token determines user experience. Inference platforms targeting latency-sensitive workloads should publish TPS on the rate card.
-
Per-use agentic tool pricing is the next transparency frontier. As agent loops scale, web search and code execution costs become the dominant bill line. Platforms that publish per-tool rates in line-item form win finance-team trust over competitors that bundle tools into opaque tiers.
-
A published price is a promise about the rate, not about availability. Two-variant Whisper transcription (Large v3 vs Turbo, a 2.8× spread wide enough to matter and narrow enough to be a real choice) plus the July 2026 Orpheus launch give a voice agent a full round-trip on one platform, each modality metered in its own native unit rather than a token proxy. But on 2026-07-21 Groq withdrew two priced models and a priced tool without moving a single rate, and the nearest substitute for the delisted Qwen costs 5.1× more on output. The pattern deepened rather than reversed: on 2026-08-11 the marketing page publishing every rate was itself removed, and on 2026-08-26 Groq gated its two highest-profile self-serve models — Llama 3.1 8B Instant and Llama 3.3 70B Versatile — behind an Enterprise sales conversation, again without moving a single retained price. Three consecutive changes across five weeks, none of them a repricing, all of them a narrowing of who can see or buy the catalog — so a rate card that is precise about price can still be silent about how long that price, or the page publishing it, will exist. Pricing teams shipping fast-moving catalogs should publish a durability tier next to the number, not only in the docs, and should treat “keeping a public pricing page live at all” as its own commitment.
-
TPU-pioneer founder credibility is the most durable silicon-platform trust anchor. Jonathan Ross’s role as the original Google TPU engineer gives Groq unusual credibility on architectural claims — distinguishing it from generic inference middleware and converting silicon-skeptical enterprise procurement leaders.
UBP implications
-
Throughput (TPS) as a published metric on token-priced rate cards lets usage-based platforms compete on latency-cost rather than just absolute per-token price. Inference platforms that hide TPS behind documentation lose latency-sensitive buyers who cannot self-qualify — and as of 2026-08-11, Groq itself now only publishes that pairing inside developer docs rather than on a marketing-facing page, meaning less-technical procurement stakeholders lose the one-view comparison entirely unless someone hands them a docs URL.
-
Per-use line-item tool pricing is becoming the canonical structure for agentic UBP, and its defining property is reversibility — a property that turns out to apply to the whole packaging layer, not just individual SKUs. Bundled-tool pricing creates forecasting uncertainty that finance teams reject, so Groq meters four distinct billing units — tokens, hours, requests, characters — each in its native unit rather than normalizing everything to tokens, which keeps every SKU independently forecastable. Groq’s last five weeks show what that modularity costs at increasing scope: the same structure that let it add a per-1M-character text-to-speech unit and a per-hour browser tool on 2026-07-14 let it withdraw that tool and two models on 2026-07-21 without touching any other line; on 2026-08-11 the modularity extended to the publishing surface itself, when the marketing page listing every tool rate disappeared rather than a single tool row; and on 2026-08-26 it extended to whole SKUs changing sales motion — two of six production LLMs were pulled behind an Enterprise “Contact Sales” gate, shifting them from self-serve UBP into sales-led custom pricing without moving a rate. As usage-based platforms mature toward enterprise revenue mix, the durability of public self-serve pricing — not just its initial existence — is becoming a metric buyers should track.
-
Bespoke silicon as a UBP moat is the structural cost advantage that pure-software platforms cannot replicate. Groq’s LPU is the canonical example — but the precedent suggests that other inference-specific silicon investments (Cerebras, Tenstorrent, Etched) will reshape competitive dynamics over the coming decade.
Sources
- Groq pricing page (accessed 2026-07-21)
- Groq docs — models catalog (accessed 2026-07-21)
- Groq docs — prompt caching (accessed 2026-07-21)
- Groq docs — batch processing (accessed 2026-07-21)
- Groq docs — rate limits (accessed 2026-07-21)
- Groq docs — web search tool (accessed 2026-07-21)
- Groq docs — code execution tool (accessed 2026-07-21)
- Groq docs — per-model pricing (GPT OSS 20B) (accessed 2026-08-28)
- Groq docs — per-model pricing (GPT OSS 120B) (accessed 2026-08-28)
- Groq docs — speech-to-text pricing (accessed 2026-08-28)
- Groq blog — Llama 3 launch (accessed 2026-05-29)
- Groq blog — Series D announcement (accessed 2026-05-29)
- Groq Saudi Arabia partnership (accessed 2026-05-29)
- Related infra blueprint — Cerebras
- Related infra blueprint — Fireworks AI
- Blueprint corpus index
Bottom line
Groq priced its inference API around three structural ideas: bespoke LPU silicon that delivers throughput GPU-based competitors cannot match, TPU-pioneer founder credibility (Jonathan Ross designed the original Google TPU) that justifies the silicon investment, and — through 2026-07-21 — transparent per-use agentic tool pricing (web search at $5–$8/1K, website visits at $1/1K, code execution at $0.18/hour) that made Groq one of the most legible commercial inference platforms in the market. That legibility has since eroded on two fronts rather than one. Marketing had explicitly framed this as “linear and predictable” pricing, and the rate card made that claim accurate per unit consumed — but the 2026-07-21 catalog contraction first showed it said nothing about which units would still be purchasable next week, and the flagship proof of the silicon moat itself has since moved: Llama 3.1 8B Instant, previously the platform’s headline $0.05/$0.08 example at 800+ TPS, went Enterprise-only “Contact Sales” on 2026-08-26 (alongside Llama 3.3 70B Versatile) and the docs now show it running 560 T/SEC, below the 840 TPS the retired marketing page had claimed. And on 2026-08-11 the marketing page carrying every one of these numbers — including the “linear and predictable” copy itself — was removed and now redirects to a homepage pitching “Premier NeoCloud for fast inference” and fresh funding instead of per-token rates. GPT OSS 20B at $0.075/$0.30, 1,000 T/SEC, is now the number that best carries the original thesis.
For AI engineering teams running latency-critical voice agents, real-time interactive UIs, and high-throughput batch workloads, Groq still delivers throughput and per-token economics that GPU-based platforms cannot match on the models that remain self-serve. Most of the remaining gaps — no fine-tuning SKU, ASR limited to Whisper, a restrictive free tier, non-stacking discounts — are TAM-expansion problems rather than structural pricing flaws. The bigger exception is catalog durability and discoverability: four LLMs are still self-serve priced (GPT OSS 20B/120B, plus Preview-tier Safety GPT OSS 20B, Qwen 3.6-27B, and — newly priced 2026-08-27 — Qwen 3.8-27B), the surviving docs catalog does label its Preview Models table distinctly, but two former self-serve models now require an Enterprise sales conversation and there is no longer any public marketing page stating a Groq rate at all. The exposure a Groq buyer actually carries is no longer just “the model leaving” — it is increasingly “the model, or the entire self-serve tier, requiring a sales call to reach.”
Compare with peers via the blueprint corpus, or model your own spend with the Groq pricing calculator.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
Qwen 3.8-27B Gains a Published Preview-Tier Price
Qwen/Qwen3.8-27B joined the GroqCloud docs' Preview Models pricing table at $0.80 input / $4.00 output per 1M tokens (450 T/SEC, 131,042-token context window, 16,384 max completion tokens, 20 MB max file size). The model ID had appeared on the Rate Limits docs page since at least 2026-08-26 (30 RPM / 8K TPM / 2M TPD on the Developer plan) but carried no price and was absent from the models catalog table entirely — this capture is the first to show a published rate. It was the only change versus the prior day's capture: the homepage, every other model price, and the Rate Limits page all held exactly.
Llama 3.1 8B Instant and Llama 3.3 70B Versatile Moved to Enterprise-Only Pricing
The two Llama models on Groq's rate card — previously self-serve at $0.05/$0.08 and $0.59/$0.79 per 1M tokens — are now marked "Enterprise" in the GroqCloud docs model catalog, with "Contact Sales" printed in place of both the price and rate-limit columns. Confirmed by a second surface: both models are also absent from the Free Plan Limits and Developer Plan Limits tables on the docs' Rate Limits page, where they previously had public RPM/TPM caps. GPT OSS 20B ($0.075/$0.30) is now the cheapest self-serve model on the card. Every other production price (GPT OSS 120B, Whisper, Orpheus, Qwen 3.6 27B) held unchanged. The docs catalog also grew two new Preview moderation models (Llama Prompt Guard 2 22M at $0.03/$0.03, Prompt Guard 2 86M at $0.04/$0.04 per 1M tokens) and formalized "Groq Compound" / "Compound Mini" as named Production Systems with no published per-token price.
Public Pricing Page Removed — Rates Survive Only in Developer Docs
groq.com/pricing/ now 308-redirects to Groq's homepage, which carries no pricing content, no Pricing nav link, and instead promotes a $650M raise and a "Premier NeoCloud for fast inference" repositioning. Per-token, per-hour, and per-character rates did not move — the entire published rate card (Llama 3.1 8B $0.05/$0.08, Llama 3.3 70B $0.59/$0.79, GPT OSS 120B $0.15/$0.60, GPT OSS 20B/Safeguard $0.075/$0.30, Qwen 3.6 27B $0.60/$3.00, Whisper $0.111/$0.04 per hour, Orpheus $22/$40 per 1M characters) is identical to 2026-07-21 — but it now lives exclusively in the GroqCloud developer-docs model catalog (console.groq.com/docs/models), not on any public marketing page.
Catalog Contraction: Two Models and the Browser Automation Tool Withdrawn
One week after expanding the rate card, Groq contracted it without repricing. Llama 4 Scout (17Bx16E, $0.11/$0.34) and Qwen3 32B ($0.29/$0.59) left both the pricing page and the docs model catalog, the Browser Automation built-in tool ($0.08/hour) was pulled from Built-In Tools (Compound) seven days after launch, Qwen3-VL 32B was dropped from the enterprise-only list, and Minimax M2.5 was replaced by M2.7. Every retained price held exactly, so the cost impact lands as forced substitution rather than as a rate move — the nearest remaining Qwen, Qwen 3.6 27B at $0.60/$3.00, costs 2.1x more on input and 5.1x more on output than the delisted Qwen3 32B.
Text-to-Speech SKU + Browser Automation Tool + New Models
Groq added a text-to-speech product line billed per 1M characters (Canopy Labs Orpheus V1 English at $22.00/1M chars, Orpheus Arabic Saudi at $40.00/1M chars), a new Browser Automation built-in tool at $0.08/hour, and several new models (GPT OSS Safeguard 20B $0.075/$0.30, Qwen 3.6 27B $0.60/$3.00, Kimi K2 $1.00/$3.00 with $0.50 cached input) plus enterprise-only Minimax M2.5 and Qwen3-VL 32B. Core per-token rates for Llama, GPT OSS, and Qwen3 held stable.
Built-In Tools Pricing: Web Search, Visits, Code Execution
Groq published per-use pricing for built-in agentic tools: web search at $5–$8 per 1,000 requests, website visits at $1 per 1,000 requests, code execution at $0.18/hour. Made Groq one of the few platforms with transparent line-item billing on agentic tool usage.
Saudi PIF Investment at $6.9B Valuation
Groq raised a Saudi PIF-led round at a $6.9B post-money valuation, with the funding earmarked for international data center expansion and additional LPU production. The Saudi investment was paired with a strategic partnership to build Groq's largest data center in Saudi Arabia.
Batch API + Cached Input Discounts
Groq launched a Batch API at 50% discount for asynchronous workloads and added cached input token discounts at 50% off standard rates. Brought Groq into parity with Fireworks, OpenAI, and Anthropic on standard inference discount mechanics.
Whisper Large v3 + Turbo Variants at Per-Hour Pricing
Groq added Whisper Large v3 ($0.111/hour transcribed) and Whisper Large v3 Turbo ($0.04/hour transcribed) for audio transcription. Per-hour billing with 10-second minimum per request established the audio SKU structure that remains canonical.
Series D ($640M) at $2.8B Valuation
Groq raised a $640M Series D led by BlackRock at a $2.8B post-money valuation. Tiger Global, Samsung Catalyst, KDDI, and others participated. The round funded LPU silicon production scaling and the GroqCloud Enterprise tier launch.
Llama 3 Day-One Support at Industry-Leading TPS
Groq launched Llama 3 support on day one with throughput exceeding 800 tokens/second for the 8B variant — the fastest published inference rate for any commercial Llama 3 endpoint at launch. Per-token pricing held competitive with serverless GPU competitors.
GroqCloud Public Launch
Groq launched GroqCloud — its inference API — with Llama 2 and Mistral models at competitive per-token rates. The launch positioning emphasized speed-of-inference (hundreds of TPS) rather than absolute cheapness, marketing LPU silicon as a structural advantage over GPU-based inference middleware.
Groq Founded
Jonathan Ross (original engineer behind Google's TPU) founded Groq to build bespoke inference silicon. The company spent 2016–2022 designing the LPU (Language Processing Unit) architecture before launching GroqCloud — the commercial inference API — in early 2023.
- · Groq's Llama 3.1 8B Instant at 840 tokens/second is one of the fastest published throughput rates for any 8B-class model in commercial inference — the LPU silicon architecture is designed specifically for inference rather than training, which is what makes the speed and pricing combination possible.
- · Groq was founded in 2016 by Jonathan Ross, the original engineer behind Google's TPU (Tensor Processing Unit) — making it the rare commercial AI inference platform built on bespoke silicon designed by the same engineer who pioneered modern ML accelerators.
- · Groq's LPU (Language Processing Unit) deliberately avoids the GPU model: each chip has deterministic execution, no HBM (uses on-die SRAM), and no GPU-style branch prediction — the trade-off is lower per-chip memory but vastly higher single-stream throughput.
Questions & answers
- How much does Groq cost per month?
- Groq has no monthly subscription fee — you pay only for the tokens you consume on serverless inference, plus per-hour audio transcription, per-use built-in tools, and per-session agentic infrastructure. A small chat application using GPT OSS 20B (the cheapest self-serve model as of 2026-08-26) at 30M input + 10M output tokens would cost roughly $5.25/month — Groq's per-token rates are among the lowest in the market for small-model inference. Note that Llama 3.1 8B Instant, which previously anchored this estimate at a lower rate, moved to Enterprise-only "Contact Sales" pricing on 2026-08-26 and is no longer self-serve.
- What are Groq's per-token rates for popular models?
- As of 2026-08-11, Groq's dedicated pricing page (groq.com/pricing/) has been removed and redirects to the homepage — rates are now published only in the GroqCloud developer docs (console.groq.com/docs/models). As of 2026-08-26, Llama 3.1 8B Instant and Llama 3.3 70B Versatile — Groq's two Llama models — are marked "Enterprise" with "Contact Sales" in place of a published rate, and both are absent from the Free and Developer rate-limit tables. The self-serve catalog per 1M tokens (input/output) is now: GPT OSS 20B and Safety GPT OSS 20B (formerly "GPT OSS Safeguard 20B") $0.075/$0.30 at 1,000 T/SEC; GPT OSS 120B $0.15/$0.60 at 500 T/SEC; Qwen 3.6 27B $0.60/$3.00 at 500 T/SEC; Qwen 3.8-27B $0.80/$4.00 at 450 T/SEC (newly priced as of 2026-08-27); and two Preview moderation models, Llama Prompt Guard 2 22M and Prompt Guard 2 86M, at $0.03/$0.03 and $0.04/$0.04. Kimi K2 Instruct 0905 was quoted on the marketing page's prompt-caching table at $1.00 uncached / $0.50 cached input and $3.00 output; it has never appeared in the docs model catalog, and with the marketing page gone that rate can no longer be verified against any live Groq source — confirm availability directly in the console before relying on it. Check availability as well as rate before committing a workload: on 2026-07-21 Groq removed Llama 4 Scout (17Bx16E) and Qwen3 32B from both the rate card and the docs catalog without changing any retained price, and Groq's docs warn that models in the Preview tier "may be discontinued at short notice" — a tier which as of 2026-08-27 currently includes Safety GPT OSS 20B, Qwen 3.6 27B, Qwen 3.8-27B (newly priced 2026-08-27), both Orpheus text-to-speech voices, and the two Prompt Guard models.
- Does Groq have a free tier?
- Yes — Groq's free tier offers rate-capped access to most models for experimentation, though as of 2026-08-26 that no longer includes Llama 3.1 8B Instant or Llama 3.3 70B Versatile, which moved to Enterprise-only "Contact Sales" pricing and dropped off the published Free Plan Limits table entirely. Throughput limits apply on the remaining models (low requests-per-minute and tokens-per-minute caps), with sustained production usage requiring a paid plan. Free tier credit-card-on-file is not required for evaluation.
- How does Groq's LPU differ from GPU-based inference?
- Groq's LPU (Language Processing Unit) is bespoke silicon designed exclusively for inference. Unlike GPUs which use HBM (high-bandwidth memory) and branch prediction, the LPU uses on-die SRAM and deterministic execution. The trade-off is lower per-chip memory capacity but vastly higher single-stream throughput — enabling 800+ tokens/second on Llama 3 8B and 500+ TPS on 120B-class models.
- What does Groq charge for audio transcription and text-to-speech?
- Transcription is billed per hour: Whisper Large v3 (higher accuracy) at $0.111 per hour and Whisper Large v3 Turbo (faster, slightly lower accuracy) at $0.04 per hour, both with a 10-second minimum per request and a 2.8× spread reflecting the accuracy-versus-speed trade-off. Text-to-speech arrived in July 2026 and is billed per 1 million characters instead: Canopy Labs Orpheus V1 English at $22.00 per 1M characters and Orpheus Arabic (Saudi) at $40.00 per 1M characters — pairing with Whisper to give Groq a full speech-in, speech-out audio round-trip on the LPU.
- How are Groq's built-in tools priced?
- Unknown as of 2026-08-11. Groq had published per-use pricing for agentic tools in two families — Compound tools (basic search, advanced search, visit website, code execution) and GPT-OSS tools (browser search, visit website, Python code execution) — last confirmed 2026-07-21 at $5 per 1,000 basic-search requests, $8 per 1,000 advanced-search requests, $1 per 1,000 website visits, and $0.18/hour of code-execution compute. Groq's marketing pricing page (which carried these rates) has since been removed and now redirects to the homepage, and the GroqCloud docs for each tool explicitly say "see the Pricing page for more information" rather than printing a number — so there is currently no live public source for these rates; confirm directly with Groq before budgeting tool usage. As of 2026-08-26 the docs model catalog also lists "Groq Compound" and "Compound Mini" (`groq/compound`, `groq/compound-mini`) as named Production Systems at ~450 T/SEC, but the price column for both shows "-" rather than a rate. Cached input still gets a 50% discount on the three caching-supported GPT-OSS models per the docs; the Batch API still gets a 50% discount for asynchronous workloads, and the two do not stack.