Cerebras CS-4: The Rack-Scale AI Accelerator Built for High-Speed Inference

The demand for faster AI inference is forcing data center operators to rethink how AI systems are designed.

As increasingly large AI models move from training into production, token generation speed, power efficiency, memory bandwidth, networking, cooling and data center deployment time are becoming just as important as raw compute performance.

Cerebras is addressing this challenge with CS-4, its next-generation rack-scale AI accelerator platform.

Designed as the successor to CS-3, CS-4 combines three newly announced Wafer Scale Engine 3 Turbo (WSE-3T) processors with a redesigned rack architecture covering compute, power, cooling and I/O.

The result is a system designed specifically around the requirements of modern AI inference—including extremely large models.

According to the specifications provided for the platform, CS-4 can deliver up to 30× higher token generation speed than production GPU systems, while providing up to 10× the throughput per watt of CS-3.

But the most interesting part of CS-4 isn’t simply the processor.

It is the architecture surrounding it.

What Is Cerebras CS-4?

Cerebras CS-4 is the company’s fourth-generation rack-scale AI accelerator platform, succeeding the CS-3.

At its core are three WSE-3T processors, which are designed to work together as part of a system-level architecture rather than operating as conventional standalone accelerator cards.

This distinction is important.

Traditional AI infrastructure typically consists of servers containing GPUs or other accelerators, with separate systems responsible for power delivery, networking and cooling.

Cerebras takes a more integrated approach.

With CS-4, the company is redesigning:

  • Compute
  • Power delivery
  • Direct liquid cooling
  • High-speed I/O
  • Interconnects
  • Rack architecture

as a coordinated system.

The goal is to remove bottlenecks that can appear when compute, power and networking are designed independently.

Cerebras CS-4 Key Specifications

FeatureCerebras CS-4
Generation4th-generation
Processor3 × WSE-3T
ArchitectureRack-scale AI accelerator
Primary focusAI inference
Inference speedUp to 30× higher token generation
Throughput efficiencyUp to 10× per watt vs. CS-3
Large-model performance1,000+ tokens/sec on 10T+ models
Wafer-to-wafer latencyAs low as 2 µs
I/O bandwidthUp to 2×
Manufacturing complexityUp to 50% fewer components
Power delivery distance~0.5 mm vs. ~50 mm conventional design
NetworkingRoCE v2 RDMA over Ethernet
AvailabilityFirst shipments beginning Q3 2026

WSE-3T: The Processor Behind Cerebras CS-4

The foundation of CS-4 is the Wafer Scale Engine 3 Turbo, or WSE-3T.

Rather than building an accelerator around relatively small silicon dies and connecting many separate accelerator packages together, Cerebras’ approach uses wafer-scale computing.

CS-4 integrates three WSE-3T processors into its rack-scale architecture.

This allows Cerebras to focus on minimizing communication overhead between processing resources.

For extremely large AI models, communication between compute elements can become a major bottleneck.

The CS-4 architecture is designed to address this through extremely low-latency wafer-to-wafer communication.

Cerebras specifies latency as low as:

2 microseconds

That becomes particularly important when running very large models.

Cerebras CS-4 Can Exceed 1,000 Tokens per Second on 10T+ Models

One of the most striking performance claims associated with CS-4 is its ability to generate more than 1,000 tokens per second on models exceeding 10 trillion parameters.

That is particularly relevant to the emerging generation of frontier AI models.

As models become larger, inference becomes increasingly difficult to scale efficiently.

Simply adding more accelerators doesn’t automatically solve the problem.

The system must also efficiently move data between those accelerators.

Cerebras’ wafer-scale architecture is designed to reduce that communication overhead.

The company says CS-4’s ultra-low wafer-to-wafer latency enables extremely high token-generation rates even for models exceeding 10 trillion parameters.

For enterprise AI deployments, this could be particularly significant for workloads where response latency and throughput directly affect operating costs and user experience.

Up to 30× Faster Inference Than Production GPU Systems

Cerebras positions CS-4 primarily around inference performance.

According to the specifications provided, the platform can deliver:

Up to 30× higher token generation speed

compared with production GPU systems across both smaller models and massive frontier architectures.

The distinction between training and inference is important.

Training teaches an AI model how to perform a task.

Inference is what happens when users actually interact with that trained model.

Every chatbot response, AI coding request, document analysis task and inference API request consumes inference capacity.

As AI applications become more widespread, infrastructure providers need to generate more tokens while keeping latency and power consumption under control.

That makes inference-focused hardware increasingly important.

CS-4 Uses a Disaggregated Inference Architecture

One of the more interesting aspects of the CS-4 design is its separation of prefill and decode workloads.

AI inference can broadly be divided into two stages:

Prefill

The system processes the incoming prompt and prepares the model state needed to generate the response.

Decode

The system generates the response token by token.

These stages have different computational characteristics.

Cerebras’ CS-4 architecture allows them to be handled by different types of hardware.

GPUs and ASICs Can Handle Prefill

The CS-4 architecture allows the prefill stage to be handled by GPUs or other accelerators.

The supplied specifications specifically identify platforms such as:

  • AWS Trainium
  • AMD Helios

as potential compute engines for this portion of the workload.

CS-4 can then focus on the decode phase.

CS-4 Focuses on Decode

The decode engine is responsible for streaming generated tokens.

This is where extremely low latency becomes particularly valuable.

Instead of forcing a single type of accelerator to handle every part of the inference process, Cerebras is proposing a heterogeneous architecture in which different systems can perform the tasks they are best suited for.

That is an important shift from the traditional approach of simply deploying increasingly large GPU clusters.

Cerebras Nexus Rack-Scale Architecture

The hardware surrounding the WSE-3T is arguably just as important as the processor itself.

Cerebras calls the architecture Nexus.

The system is organized around three modular pillars:

1. Compute

2. Power

3. I/O

Instead of treating the rack as one monolithic system, these components are designed as modular assemblies.

That provides several potential advantages.

50% Fewer Components

Cerebras says the Nexus architecture can reduce the number of components required by approximately 50%.

That isn’t merely a manufacturing statistic.

Fewer components can potentially simplify:

  • Manufacturing
  • Assembly
  • Maintenance
  • Deployment
  • Upgrades
  • Supply-chain management

Cerebras also says the architecture can reduce deployment timelines from days to hours.

For large AI data centers, that could become an important operational advantage.

The “Wafer-Scale Backpack”

One of the more unusual elements of the CS-4 architecture is the pluggable Wafer-Scale Backpack.

This rear-mounted 3D compute assembly connects vertically to the power array.

Instead of separating major infrastructure components, the design integrates several functions around the wafer.

These include:

  • Power conversion
  • Direct liquid cooling
  • High-speed I/O
  • Control electronics

The result is a more tightly integrated compute assembly.

Cerebras also says the manufacturing process uses 60% more automation, which is intended to improve manufacturing efficiency and consistency.

CS-4 Brings Power Conversion Much Closer to the Silicon

Power delivery becomes increasingly challenging as AI processors consume more electricity.

Traditional systems can have significant physical distance between the power conversion circuitry and the silicon.

Cerebras is taking a very different approach with CS-4.

The company moves power conversion approximately:

100× closer to the silicon

The supplied specifications put the distance at roughly:

~0.5 mm

compared with approximately:

~50 mm

in conventional GPU board designs.

The idea is straightforward:

The shorter the electrical path, the less opportunity there is for power to be lost between the power-delivery system and the processor.

According to the provided specifications, this design enables CS-4 to deliver approximately 2× more power to the WSE-3T, allowing the processor to sustain higher clock frequencies.

Direct Liquid Cooling for High-Density AI Compute

Power density and cooling are closely connected.

As AI processors become more powerful, removing heat becomes one of the major challenges for data center operators.

CS-4 incorporates direct liquid cooling into its integrated compute architecture.

Rather than treating cooling as an external infrastructure layer, Cerebras integrates cooling directly into the wafer-scale compute assembly.

This system-level approach is particularly relevant for high-density AI racks where conventional air cooling becomes increasingly difficult.

For US data center operators dealing with power and thermal constraints, this is an important part of the CS-4 design.

Modular I/O and Direct Wafer Links

The third major component of the Nexus architecture is I/O.

CS-4 is designed to:

  • Double I/O bandwidth
  • Reduce communication latency
  • Support heterogeneous AI clusters
  • Provide direct wafer-to-wafer connectivity

The platform also supports RoCE v2 RDMA over Ethernet.

RoCE, or RDMA over Converged Ethernet, allows high-performance data transfers over Ethernet networks while reducing communication overhead.

That becomes particularly useful when CS-4 is integrated into larger heterogeneous infrastructure.

Direct Wafer Links Remove the Need for Some Switching

CS-4 also introduces Direct Wafer Links.

These links are designed to provide high-speed connectivity across and within racks without relying entirely on conventional switching architectures.

The potential benefit is reduced communication latency.

For very large AI models, communication can become as important as raw compute.

If processors spend too much time waiting for data to arrive, adding more compute doesn’t necessarily produce proportional performance improvements.

Direct connectivity is therefore a central part of Cerebras’ approach to scaling AI inference.

CS-4 vs. Traditional GPU Infrastructure

The most interesting way to understand CS-4 is not simply as another AI accelerator.

It represents a different philosophy for building AI infrastructure.

AreaTraditional GPU ApproachCerebras CS-4
ComputeMultiple GPU accelerators3 × WSE-3T
ScalingCluster GPUsRack-scale architecture
InterconnectNetworking/switch fabricDirect wafer links + networking
PowerSeparate infrastructureClosely integrated power delivery
CoolingRack/server-level systemsDirect liquid cooling
I/OConventional accelerator networkingModular high-bandwidth I/O
InferenceGeneral-purpose GPU workloadsStrong focus on high-speed decode
ArchitectureRelatively distributedHighly integrated

This doesn’t mean CS-4 automatically replaces GPU infrastructure.

Instead, Cerebras is targeting a different point in the AI infrastructure stack.

Why Cerebras CS-4 Matters for AI Data Centers

The AI infrastructure market is entering a phase where performance per watt and performance per rack are becoming increasingly important.

A data center doesn’t have unlimited:

  • Electricity
  • Cooling capacity
  • Floor space
  • Network bandwidth
  • Capital expenditure

This creates a problem for AI infrastructure providers.

Adding more accelerators increases computational capacity, but it also increases power consumption, cooling requirements and networking complexity.

Cerebras’ CS-4 approach attempts to address several of those problems simultaneously.

The company’s stated figures include:

30× higher token-generation speed

10× throughput per watt compared with CS-3

1,000+ tokens/sec on 10T+ models

50% fewer components

2× I/O bandwidth

These numbers should be viewed as Cerebras’ stated performance and architectural claims rather than universal benchmarks applicable to every workload.

Actual performance will depend on the model, software stack, workload and deployment configuration.

When Will Cerebras CS-4 Be Available?

According to the provided information, first CS-4 shipments are scheduled to begin in Q3 2026.

That places the platform’s initial availability in the second half of 2026.

Its eventual impact will depend on how the hardware performs in production deployments and how easily it integrates with existing AI infrastructure.

The Bigger Picture: The Race Beyond GPUs

The AI accelerator market is no longer simply a competition between different GPU generations.

Companies are increasingly experimenting with:

  • Wafer-scale processors
  • Custom AI ASICs
  • Specialized inference accelerators
  • Disaggregated inference
  • Rack-scale systems
  • Liquid-cooled infrastructure
  • Ethernet-based AI networking

Cerebras CS-4 fits directly into this broader transition.

Instead of asking only:

“How powerful is the processor?”

AI infrastructure designers increasingly have to ask:

“How efficiently can the entire system generate useful AI output?”

That includes compute, memory, networking, power and cooling.

Cerebras’ approach is to design those components together.

Final Thoughts

Cerebras CS-4 is an ambitious attempt to rethink AI inference infrastructure at the rack level rather than treating the accelerator as an isolated component.

Its combination of three WSE-3T processors, direct liquid cooling, high-density power delivery, modular I/O and direct wafer links is designed around one central objective: generating AI tokens extremely quickly while using data center resources efficiently.

The platform’s headline claims—up to 30× faster token generation, 10× higher throughput per watt than CS-3, and more than 1,000 tokens per second on models exceeding 10 trillion parameters—make it an interesting system to watch as AI inference workloads continue to scale.

The bigger question is whether this rack-scale approach can translate its architectural advantages into consistent advantages across real-world enterprise and frontier AI workloads.

With first shipments expected to begin in Q3 2026, CS-4 could become an important test of whether wafer-scale computing can carve out a meaningful role alongside GPUs and custom AI accelerators in the next generation of AI data centers.

What is Cerebras CS-4?

Cerebras CS-4 is a fourth-generation rack-scale AI accelerator platform designed primarily for high-speed AI inference.

What processor does CS-4 use?

CS-4 is built around three Wafer Scale Engine 3 Turbo (WSE-3T) processors.

How fast is Cerebras CS-4?

Cerebras’ stated specifications indicate up to 30× higher token-generation speed than production GPU systems, depending on the workload and comparison configuration.

How many tokens per second can CS-4 generate?

The platform is specified to deliver more than 1,000 tokens per second on models exceeding 10 trillion parameters.

What is the CS-4 power-efficiency improvement?

Cerebras specifies up to 10× more throughput per watt compared with CS-3.

Does CS-4 use liquid cooling?

Yes. Direct liquid cooling is integrated into the CS-4 compute architecture.

When will Cerebras CS-4 ship?

First shipments are expected to begin in Q3 2026, according to the information provided.

Leave a Comment