ZAI-GLM-4.7 removed from Cerebras's public per-token rate card
Cerebras removed ZAI-GLM-4.7 ($2.25/$2.75 per 1M tokens) from its public Developer Tier Pricing rate card and docs Model Catalog on its previously disclosed Aug 17, 2026 deprecation date, leaving two public models.
Three-model public rate card: GPT OSS 120B $0.35/$0.75 (Production), Google Deepmind Gemma 4 31B $0.99/$1.49 (Preview), ZAI GLM 4.7 $2.25/$2.75 (Preview, footnoted for deprecation on Aug 17, 2026).
Two-model public rate card: GPT OSS 120B $0.35/$0.75 (Production) and Google Deepmind Gemma 4 31B $0.99/$1.49 (Preview). ZAI GLM 4.7 is gone from both the pricing page rate table and the docs Model Catalog; it remains reachable only via Dedicated Endpoints (Z.AI GLM 4.X/5.X) on custom pricing.
Cerebras’s public “Developer Tier Pricing” rate card narrowed from three models to two. ZAI-GLM-4.7, which had carried a footnoted deprecation date of August 17, 2026 since at least the prior capture on July 21, 2026, is no longer listed on the pricing page’s rate table or in the inference docs’ Model Catalog as of an August 26, 2026 capture — confirmed by inspecting the live page HTML, where the rate table’s scroll-left/scroll-right controls are disabled (no hidden third row) and the docs sidebar under Models > Model Catalog now lists only two entries (Gemma 4 31B and OpenAI GPT OSS).
No per-token price moved on either remaining public model. ZAI-GLM-4.7 is not gone from Cerebras entirely — the Z.AI GLM 4.X and GLM 5.X families remain available through Dedicated Endpoints on custom reserved-capacity pricing, and the fixed-price Cerebras Code coding subscriptions (Pro $50/mo, Max $200/mo) still route to GLM 4.7 independent of the public per-token card. This is the second model Cerebras has fully sunset from its public rate card, after ZAI-GLM-4.6 in January 2026 — both times on a disclosed, followed-through deprecation date rather than an abrupt removal.
Cerebras announced CS-4, built from three new Wafer Scale Engine 3 Turbo processors on a redesigned modular rack ("Nexus Platform"). Cerebras claims up to 30x faster inference than GPU systems, up to 10x more throughput per watt and up to 2x faster performance than CS-3, and wafer-to-wafer interconnect latency as low as 2 microseconds. No pricing was disclosed — CS-4 remains an enterprise-contract, "contact us" product like CS-3 — with first shipments described as beginning "this quarter" (Q3 2026).