Cerebras has introduced the CS-4, a new rack-scale AI system that delivers a 30x increase in token generation speed compared to its predecessors. This leap matters because it directly addresses the bottleneck of inference latency, allowing enterprises to deploy large language models with significantly lower response times. The CS-4 system is engineered to support large-scale AI workloads while avoiding the significant power and space increases typical of traditional GPU clusters.

Cerebras CS-4 introduces 750 PFLOPs of AI compute power using three WSE-3T chips
The CS-4 relies on three WSE-3T chips to achieve its performance targets, creating a unified computing environment on a single wafer. Each WSE-3T chip contributes 250 PFLOPs of AI compute power, resulting in a total system capability of 750 PFLOPs. This architecture eliminates the need for complex interconnects between separate processors, as the entire wafer functions as one large processor.
Specifications
- Compute (CS-4): 750 PFLOPs AI compute (3 WSE-3T chips)
- Memory Bandwidth (CS-4): 43,200 TB/s
- Fabric Bandwidth (CS-4): 53.5 PB/s
- SRAM (CS-4): 132 GB (44 GB per chip)
- Power Delivery: 54.5VDC Busbar, 0.5mm distance to chip
Memory bandwidth remains the critical differentiator for this hardware, with the CS-4 offering 43,200 TB/s per chip. This figure represents a 2000x increase over the memory bandwidth of NVIDIA's Rubin platform. The system also features a 53.5 PB/s fabric on the wafer itself, which is 200x higher than the interconnect bandwidth found in NVIDIA NVL72 racks. To support this data flow, Cerebras uses a 54.5VDC busbar that places power delivery just 0.5mm from the chip, minimizing energy loss.
The CS-4 includes 132 GB of SRAM, with 44 GB allocated to each WSE-3T chip. This massive on-chip memory allows the system to keep large model weights readily accessible, further reducing latency. The system is scheduled for general availability in Q3 2026, giving enterprises time to plan their infrastructure upgrades around this new standard.
Cerebras has also outlined its roadmap for future generations, including the CS-5 and CS-6. The CS-5, launching in 2027, aims to deliver 10,000 tokens per second per user for specific models. The subsequent CS-6 generation will integrate 3D Wafer-Scale SRAM on top of the WSE chip through 3D integration. These future steps indicate a continued focus on memory density and speed as the primary drivers of AI performance.
The analysis includes a review of previous Cerebras generations to contextualize the CS-4 launch. The CS-4 design emphasizes memory bandwidth alongside compute capabilities, reflecting Cerebras' architectural strategy. This architecture makes the CS-4 particularly suited for inference-heavy workloads, distinguishing it from general-purpose training clusters.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.