Cerebras Pricing Calculator
Updated August 2026Usage-based per-token inference API ($5-credit Free Trial, $10 self-serve Developer, and Enterprise tiers) plus fixed-price Cerebras Code coding subscriptions; hardware systems on custom enterprise contracts
Cerebras pricing: Cerebras pricing 2026: per-token inference (GPT-OSS-120B $0.35/$0.75, Gemma 4 31B $0.99/$1.49), $5 free-trial credits, $10 Developer, $50/$200 Code plans. Use the free calculator below to enter your usage and get an instant all-in monthly estimate — including overages — and see which plan is cheapest for your needs.
How Cerebras prices: Usage-based per-token inference API ($5-credit Free Trial, $10 self-serve Developer, and Enterprise tiers) plus fixed-price Cerebras Code coding subscriptions; hardware systems on custom enterprise contracts.
How Cerebras prices
Usage-based per-token inference API ($5-credit Free Trial, $10 self-serve Developer, and Enterprise tiers) plus fixed-price Cerebras Code coding subscriptions; hardware systems on custom enterprise contracts
Best fit — auto-selected: Enterprise / CS-3 / CS-4
$0.00/mo— the cheapest plan that fits; adjust usage and it re-picks
Your usage
💡 Prompt + context tokens sent to the model. 300M ≈ a small production chatbot (~10M input tokens/day).
💡 Generated response tokens — the pricier side. Output is typically 20-40% of input volume.
Pick the same model in both rate selectors. GPT-OSS-120B is the only production model; Gemma 4 31B is Preview (evaluation only). ZAI-GLM-4.7 was removed from the public rate card as of 2026-08-26, consistent with its disclosed 2026-08-17 deprecation date — it now lives only on Dedicated Endpoints custom pricing.
Output tokens cost more than input on both public rate-card models.
Compare Plans at Your Usage
Monthly Estimate
Enterprise / CS-3 / CS-4
$0.00
per month
Cost Breakdown
Annual
$0.00
Why Enterprise / CS-3 / CS-4 — $0.00/mo
- Enterprise / CS-3 / CS-4 at $0.00/mo.
- How we got there: 300 M tokens (Input tokens / mo) × 0 $/M tokens (Model — input rate) = 105 $.
- How we got there: 90 M tokens (Output tokens / mo) × 1 $/M tokens (Model — output rate) = 68 $.
Budget range
Plan for usage swings, not just today's estimate.
Conservative
$0.00
45 M tokens/mo
Expected
$0.00
90 M tokens/mo
Aggressive
$0.00
180 M tokens/mo
Need a calculator like this on your pricing page?
Embed interactive pricing calculators on your website to help customers understand costs and boost conversions.
Get StartedCerebras plans at a glance
| Plan | Price | Best for |
|---|---|---|
| Free Trial | Free | Developers getting started — $5 in free credits, Discord support |
| Developer (pay-per-token) | Usage-based | Power users and production — self-serve, 10x higher rate limits |
| Cerebras Code Pro | $50/mo | Indie devs, agentic coding — fixed monthly (sold out) |
| Cerebras Code Max | $200/mo | Full-time dev, IDE + multi-agent — fixed monthly (sold out) |
| Enterprise / CS-3 / CS-4 | Usage-based | Custom weights, guaranteed uptime, hardware — contact sales |
Prices shown are list rates as of August 2026. Enter your usage above for an all-in estimate including overages and add-ons.
About this Cerebras calculator
This calculator estimates your Cerebras cost from publicly available pricing. Actual costs may vary with your specific agreement, volume discounts, and usage patterns — always verify on the provider's official pricing page for the most current rates.
Cerebras pricing — frequently asked questions
How much does Cerebras Inference cost per million tokens? ▼
The public Cerebras rate card lists two models as of August 2026: GPT-OSS-120B at $0.35 input/$0.75 output per million tokens (production) and Google Deepmind Gemma 4 31B at $0.99 input/$1.49 output (Preview). ZAI-GLM-4.7, previously listed at $2.25 input/$2.75 output, was removed from the public card on its disclosed deprecation date of August 17, 2026. Preview models are intended for evaluation only, so GPT-OSS-120B is the only public model sanctioned for production. Other models such as Llama 3.3 70B, Qwen3-32B, and now the Z.AI GLM family are available via Dedicated Endpoints on custom pricing rather than the public rate card.
Does Cerebras offer a free tier for the inference API? ▼
No longer. As of July 21, 2026 Cerebras replaced its open free tier with a Free Trial that grants $5 in one-time credits after you create an account, with access to all Cerebras-powered models and Discord support. The docs add two conditions the pricing page leaves out: the credits are granted only after you add a verified payment method (adding one is free), and they expire 30 days after they are granted. That $5 is worth roughly 14 million input tokens on GPT-OSS-120B, and it does not renew. Past the trial, the self-serve Developer tier starts at just $10 and offers 10x higher rate limits and higher-priority processing; the Enterprise tier (contact sales) adds custom weights, dedicated queue priority, and guaranteed uptime.
How fast is Cerebras inference compared to GPU-based providers? ▼
Cerebras Inference delivers 1,000–2,100 tokens per second on Llama 3.1 70B-class models, compared to 40–80 tokens/second on GPU-based providers like Together AI or Fireworks AI. The speed advantage comes from the on-chip SRAM of the WSE eliminating GPU memory bandwidth bottlenecks.
What models are available on Cerebras Inference? ▼
As of August 2026, the public Cerebras per-token rate card lists just GPT-OSS-120B (production) and Google Deepmind Gemma 4 31B (Preview) — ZAI-GLM-4.7 was removed from the card on its disclosed August 17, 2026 deprecation date. A much wider catalog — including Llama 3.3 70B, Llama 4 Maverick and Scout, Qwen3-32B and Qwen3-235B, Qwen3-Coder, Mistral, DeepSeek, Kimi K2.x, and Z.AI's GLM 4.X/5.X families (including the former GLM-4.7) — is available through Dedicated Endpoints on reserved-capacity custom pricing. The catalog focuses on open-source models that benefit most from Cerebras's speed advantage.
Is this Cerebras pricing calculator free and accurate? ▼
Yes — it's completely free with no signup. It uses the latest publicly available Cerebras pricing (verified August 2026). Actual costs may vary with volume discounts or enterprise terms — always confirm on Cerebras's official pricing page.