Ask
Deprecation

ZAI-GLM-4.7 removed from Cerebras's public per-token rate card

Cerebras pricing

Cerebras removed ZAI-GLM-4.7 ($2.25/$2.75 per 1M tokens) from its public Developer Tier Pricing rate card and docs Model Catalog on its previously disclosed Aug 17, 2026 deprecation date, leaving two public models.

Before

Three-model public rate card: GPT OSS 120B $0.35/$0.75 (Production), Google Deepmind Gemma 4 31B $0.99/$1.49 (Preview), ZAI GLM 4.7 $2.25/$2.75 (Preview, footnoted for deprecation on Aug 17, 2026).

After

Two-model public rate card: GPT OSS 120B $0.35/$0.75 (Production) and Google Deepmind Gemma 4 31B $0.99/$1.49 (Preview). ZAI GLM 4.7 is gone from both the pricing page rate table and the docs Model Catalog; it remains reachable only via Dedicated Endpoints (Z.AI GLM 4.X/5.X) on custom pricing.

Cerebras’s public “Developer Tier Pricing” rate card narrowed from three models to two. ZAI-GLM-4.7, which had carried a footnoted deprecation date of August 17, 2026 since at least the prior capture on July 21, 2026, is no longer listed on the pricing page’s rate table or in the inference docs’ Model Catalog as of an August 26, 2026 capture — confirmed by inspecting the live page HTML, where the rate table’s scroll-left/scroll-right controls are disabled (no hidden third row) and the docs sidebar under Models > Model Catalog now lists only two entries (Gemma 4 31B and OpenAI GPT OSS).

No per-token price moved on either remaining public model. ZAI-GLM-4.7 is not gone from Cerebras entirely — the Z.AI GLM 4.X and GLM 5.X families remain available through Dedicated Endpoints on custom reserved-capacity pricing, and the fixed-price Cerebras Code coding subscriptions (Pro $50/mo, Max $200/mo) still route to GLM 4.7 independent of the public per-token card. This is the second model Cerebras has fully sunset from its public rate card, after ZAI-GLM-4.6 in January 2026 — both times on a disclosed, followed-through deprecation date rather than an abrupt removal.

From Cerebras's pricing timeline
Cerebras announces CS-4, a 4th-generation wafer-scale system

Cerebras announced CS-4, built from three new Wafer Scale Engine 3 Turbo processors on a redesigned modular rack ("Nexus Platform"). Cerebras claims up to 30x faster inference than GPU systems, up to 10x more throughput per watt and up to 2x faster performance than CS-3, and wafer-to-wafer interconnect latency as low as 2 microseconds. No pricing was disclosed — CS-4 remains an enterprise-contract, "contact us" product like CS-3 — with first shipments described as beginning "this quarter" (Q3 2026).

About Cerebras
cerebras.ai ↗

Cerebras operates a per-token inference API (Cerebras Inference) powered by its proprietary Wafer Scale Engine (WSE) chips — the only major LLM inference platform that runs entirely on non-Nvidia hardware at inference scale.

Free tier
No
Commits
Available
Transparency
public

Cerebras pricing history

  1. Aug 2026
    Cerebras announces CS-4, a 4th-generation wafer-scale system
  2. Aug 2026
    ZAI-GLM-4.7 deprecated and removed from the public rate card
  3. Jul 2026
    Free tier becomes a $5 credit trial; Gemma 4 31B joins the rate card
  4. May 2026
    Public rate card narrows; Cerebras Code subscriptions launch
  5. Jul 2025
    Qwen-3-32B and ZAI-GLM-4.x Models Added
Full Cerebras timeline

More Cerebras activity

All pricing activity