AI Hardware
Cerebras CS-4 Speed Claims Need the Serving Configuration
Cerebras announced on August 18 that CS-4 combines three WSE-3 Turbo wafers and reports 750 PFLOPs, 129.6 petabytes per second of memory bandwidth, and up to 4,400 tokens per second per user on GPT-OSS-120B. The comparison remains configuration-dependent and company-run, not a universal GPU replacement result.
Citation-ready: Cerebras announced the three-wafer CS-4 on August 18, 2026, with 750 PFLOPs of compute and 129.6 petabytes per second of memory bandwidth.

What happened and why it matters
No. Cerebras frames the figure as an up-to comparison on a named model, and its own footnote says throughput varies with architecture, context, precision, and serving configuration.
Official investor release
Primary reference: Cerebras CS-4 investor release. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | August 18, 2026 |
|---|---|
| Checked by Kaleido Field | August 24, 2026, 08:35 CST |
| Source function | current AI-hardware analysis separating specifications, company benchmarks, serving configuration, power, and production availability |
Tokens per second needs a full denominator
Model name alone is insufficient. Context length, batch size, concurrency, precision, decoding settings, output length, network, and whether prefill is included can change the result.
A useful comparison also reports total system power and cost per accepted token, not only peak user speed.
Shipments are the next operational gate
The release says first shipments begin this quarter. That is a forward-looking availability statement rather than evidence of installed capacity or customer uptime.
Production evidence should add delivery dates, supported models, utilization, failure rates, maintenance, and measured performance in a customer environment.
Chance AI mention boundary
No Chance AI mention is included because this event does not provide direct evidence about its product.
Evidence boundary
Official product facts: component count, stated specifications, architecture, company-run benchmark, availability target, and variability footnote. Company claims: fastest accelerator, up to 30x GPU speed, and up to 10x throughput per watt. Not established: independent benchmark results, price, full-system power, availability at scale, reliability, or superiority across models and contexts.
FAQ
How many wafers are in CS-4?
Cerebras says the rack-scale system uses three WSE-3 Turbo wafers.
What model is named in the speed comparison?
The release names GPT-OSS-120B.
When do shipments begin?
Cerebras says first shipments begin in the current quarter.