Ask
Launch

Cerebras Code subscriptions launch; rate card narrows

Cerebras pricing

Cerebras introduces fixed-price Cerebras Code coding subscriptions alongside a tightened public inference rate card.

From Cerebras's pricing timeline
Public rate card narrows; Cerebras Code subscriptions launch

By mid-2026 the public per-token rate card lists just two models — GPT-OSS-120B ($0.35/$0.75, production) and ZAI-GLM-4.7 ($2.25/$2.75, labeled a Preview/evaluation model). Llama and Qwen3 families moved to Dedicated Endpoints on reserved-capacity custom pricing. Access is now tiered (Free, a self-serve Developer tier from $10, and Enterprise), and Cerebras introduced fixed-price Cerebras Code coding plans — Pro at $50/month (24M tokens/day) and Max at $200/month (120M tokens/day), both sold out at launch.

About Cerebras
cerebras.ai ↗

Cerebras operates a per-token inference API (Cerebras Inference) powered by its proprietary Wafer Scale Engine (WSE) chips — the only major LLM inference platform that runs entirely on non-Nvidia hardware at inference scale.

Free tier
No
Commits
Available
Transparency
public

Cerebras pricing history

  1. Aug 2026
    Cerebras announces CS-4, a 4th-generation wafer-scale system
  2. Aug 2026
    ZAI-GLM-4.7 deprecated and removed from the public rate card
  3. Jul 2026
    Free tier becomes a $5 credit trial; Gemma 4 31B joins the rate card
  4. May 2026
    Public rate card narrows; Cerebras Code subscriptions launch
  5. Jul 2025
    Qwen-3-32B and ZAI-GLM-4.x Models Added
Full Cerebras timeline

More Cerebras activity

All pricing activity