Cerebras used Hot Chips 2026 to lay out how it plans to scale its wafer-scale AI accelerators beyond individual systems and into a rack-scale architecture. ServeTheHome, reporting live from the presentation, says the company presented its recently announced WSE-3 Turbo accelerators and CS-4 racks to the conference’s technical audience. Tom’s Hardware also reports that Cerebras discussed the CS-4 rack-scale accelerator, its Nexus rack design, and the refreshed wafer-scale engines inside it. The current system discussed is CS-4, a rack-scale system that the sources describe as incorporating three refreshed wafer-scale engines. ServeTheHome describes the chips as WSE-3 Turbo or WSE-3T engines; Tom’s Hardware refers to them as WS-3T wafers. Both accounts point to the same architectural move: Cerebras is packaging multiple wafer-scale processors into a single rack-scale system rather than treating each wafer-scale engine as a standalone unit. ServeTheHome says Cerebras framed CS-4 as its first dedicated rack-scale system operating within a single scale-up domain. The outlet reports that Cerebras sees the Nexus platform behind CS-4 as the hardware backbone for its systems over the next several years. That matters because Cerebras’ existing advantage is built around very large processors with on-chip SRAM, and the rack design is the mechanism for scaling that approach into larger deployments. The performance claims should be read as company claims reported from the presentation. According to ServeTheHome, Cerebras said CS-4 offers twice the tokens and 10 times the tokens per watt versus CS-3. The outlet also reports that the WSE-3T relies entirely on on-chip SRAM and is specified at 43,000 TB/s of memory bandwidth. Tom’s Hardware’s headline says the Nexus architecture triples rack-scale performance, but the provided text does not give the same metric detail, so the exact performance comparison remains source-attributed rather than independently reconciled. Nexus is also a packaging and serviceability story. ServeTheHome reports that power enters at the front of the Nexus rack while compute is installed at the rear through pluggable “backpacks.” Those backpacks provide I/O connectivity as well as power and cooling attachment, according to ServeTheHome, which also reports that the CS-4 backpack offers twice the power and cooling of the CS-3 design with 50% fewer components. Tom’s Hardware describes the same modules as self-contained units that incorporate power delivery, scale-up networking, and liquid cooling, and says future wafer-scale engines built for the architecture could be swapped in without replacing the full rack. Cerebras is positioning that design against GPU rack complexity. Tom’s Hardware reports that Cerebras contrasted Nexus with Nvidia’s Rubin NVL72 rack-scale domain, citing 5,000 cables used to connect that system’s NVLink scale-up fabric. ServeTheHome says Cerebras also discussed power delivery differences versus GPUs, including a vertical design intended to keep DC-DC converters and WSEs close to the busbar to reduce resistive losses. The longer-range roadmap is more ambitious. Tom’s Hardware reports that Cerebras disclosed a CS-6 system, two generations out, whose wafer-scale engine would use 3D-stacked DRAM on top of a logic and SRAM wafer. The outlet frames that as Cerebras’ answer to a scaling problem: GPU vendors can add high-bandwidth memory capacity near the processor, while a wafer-scale design already consumes a full 300mm wafer and must find other ways to add memory resources. There is also a hybrid-compute angle. ServeTheHome says Cerebras is already working to mesh WSEs with GPU racks, including pairing WSEs with AMD’s forthcoming Instinct MI455X accelerators, building on the WSE’s high decode performance. That detail is single-source in the provided cluster, but it shows how Cerebras is pitching wafer-scale systems not only as GPU alternatives, but as possible complements inside larger AI infrastructure designs. Who benefits: Cerebras benefits if Nexus makes its wafer-scale systems more modular and easier to scale. Operators serving low-latency, high-throughput inference workloads are the audience most directly addressed by the architecture described in the reports. Who's exposed: GPU rack vendors are the comparison target in both accounts, especially around power delivery and cabling. The exposure is not a proven market shift yet; it is the competitive frame Cerebras presented at Hot Chips.