Ask
Launch

Cerebras announces CS-4, its fourth-generation wafer-scale system

Cerebras pricing

Cerebras unveiled CS-4, built on three Wafer Scale Engine 3 Turbo processors on a new modular Nexus rack, claiming up to 30x faster inference than GPUs and up to 2x faster than CS-3; pricing is undisclosed and sold via enterprise contract.

Before

CS-3 (WSE-3, announced March 2024) was Cerebras's newest enterprise compute system, sold via direct enterprise contract with no public pricing.

After

CS-4 (announced Aug 18, 2026) becomes Cerebras's newest system — three WSE-3 Turbo processors on a redesigned modular 'Nexus Platform' rack, claimed up to 30x faster than GPU systems and up to 2x faster than CS-3, with up to 10x more throughput per watt. No public pricing disclosed; first shipments described as beginning Q3 2026.

Cerebras announced CS-4 on August 18, 2026, its fourth-generation compute system and the newest addition to its enterprise hardware line alongside CS-3. The system pairs three new Wafer Scale Engine 3 Turbo processors with a completely redesigned rack architecture Cerebras calls the “Nexus Platform” — a modular design the company says cuts component count by roughly half and reduces wafer-to-wafer interconnect latency to as low as 2 microseconds, enabling more than 1,000 tokens/second on models exceeding 10 trillion parameters.

Cerebras claims CS-4 delivers up to 30x faster inference than production GPU systems and up to 10x more throughput per watt than CS-3, with up to 2x faster raw performance than its predecessor. As with CS-3, no public pricing is disclosed — CS-4 is sold exclusively through direct enterprise sales engagement, with Cerebras stating first shipments begin “this quarter” (Q3 2026). The announcement does not affect pricing or packaging on Cerebras’s separate, self-serve Inference API product line (Free Trial, Developer, Enterprise tiers).

From Cerebras's pricing timeline
Cerebras announces CS-4, a 4th-generation wafer-scale system

Cerebras announced CS-4, built from three new Wafer Scale Engine 3 Turbo processors on a redesigned modular rack ("Nexus Platform"). Cerebras claims up to 30x faster inference than GPU systems, up to 10x more throughput per watt and up to 2x faster performance than CS-3, and wafer-to-wafer interconnect latency as low as 2 microseconds. No pricing was disclosed — CS-4 remains an enterprise-contract, "contact us" product like CS-3 — with first shipments described as beginning "this quarter" (Q3 2026).

About Cerebras
cerebras.ai ↗

Cerebras operates a per-token inference API (Cerebras Inference) powered by its proprietary Wafer Scale Engine (WSE) chips — the only major LLM inference platform that runs entirely on non-Nvidia hardware at inference scale.

Free tier
No
Commits
Available
Transparency
public

Cerebras pricing history

  1. Aug 2026
    Cerebras announces CS-4, a 4th-generation wafer-scale system
  2. Aug 2026
    ZAI-GLM-4.7 deprecated and removed from the public rate card
  3. Jul 2026
    Free tier becomes a $5 credit trial; Gemma 4 31B joins the rate card
  4. May 2026
    Public rate card narrows; Cerebras Code subscriptions launch
  5. Jul 2025
    Qwen-3-32B and ZAI-GLM-4.x Models Added
Full Cerebras timeline

More Cerebras activity

All pricing activity