Cerebras Systems has launched a new server platform aimed at accelerating inference tasks for AI chatbots. The CS-4 rack is centered on three of the company's large, dinner-plate-sized chips and leverages the firm’s Nexus server architecture, which organizes the hardware into pluggable modules that house the chips.
Company executives say the physical size of the chips reduces the need to move data between chips, avoiding both the energy cost and the latency associated with inter-chip data transfers. The CS-4 rack includes a chip branded WSE-3 Turbo and updated networking hardware designed to improve data movement among the three chips in the chassis.
The new system is slated to be available in the third quarter. Cerebras indicated the chips are produced using TSMC’s 5-nanometer manufacturing process. In addition to raw performance, the vendor emphasized ease of deployment: the CS-4 is engineered with 50% fewer components compared with prior systems, a simplification that the company’s Chief Technology Officer, Sean Lie, told a media briefing in San Francisco will accelerate data center construction.
Looking further ahead, Cerebras plans to ship another generation of both chip and server hardware in 2027. The company’s engineering roadmap is focused on raising the amount of data its future chips and systems can process, and leadership provided specific performance goals tied to that work.
At the briefing, CEO Andrew Feldman outlined the company’s performance targets and capacity goals. Feldman said the business expects to deliver 600 megawatts’ worth of computing power by the end of 2027. He also stated, "We’re going to get four times as fast between now and the end of the year, end of 2027, and we’re going to get 20 times more throughput," referring to planned improvements to speed and data-handling capacity.
Cerebras positions its hardware in the portion of AI workloads known as inference, the stage of computation that generates answers in chatbots such as Anthropic’s Claude. The company competes with larger GPU vendors in this market segment, noting the performance advantages the wafer-scale chip approach can deliver.
On the financial front, the company reported an adjusted loss of $6.9 million on sales of $180.1 million last week. The launch of the CS-4 and the roadmap toward higher throughput and capacity are presented alongside these recent results as the company aims to scale both performance and deployed power.
What the CS-4 includes
- Three wafer-scale chips in a single server rack.
- WSE-3 Turbo chip paired with new networking components to speed inter-chip data movement.
- Nexus server architecture with pluggable modules and 50% fewer components for simpler setup.
- Chips fabricated on TSMC's 5-nanometer process.
The company says the combination of larger chips, reduced component count and updated networking will translate into better performance for inference workloads. Executives discussed the product at a media briefing in San Francisco and set out a multi-year cadence for subsequent generations of chips and systems.
While Cerebras has set explicit targets for speed increases, throughput gains and cumulative deployed computing capacity, the company’s recent adjusted loss on revenue is part of the current financial backdrop against which the new hardware will be marketed to cloud and enterprise customers.