Cerebras has introduced CS-4, a rack-scale AI inference system built around its WSE-Turbo wafer-scale processor and a new Nexus platform architecture. The Register reports that the company unveiled the next-generation Wafer Scale Engine and Nexus rack systems on Tuesday, while Cerebras’ CS-4 page describes the product as the first implementation of its Nexus Platform Architecture. The core pitch is higher inference throughput from the same wafer-scale approach Cerebras has used to differentiate itself from graphics processing unit-based systems. Cerebras says CS-4 delivers up to 10x more throughput per watt than CS-3 and generates tokens up to 30x faster than production GPU systems. The company also says it can deliver more than 1,000 tokens per second on models exceeding 10 trillion parameters, helped by wafer-to-wafer interconnect latency as low as two microseconds. Those performance figures should be read as claims, not neutral benchmarks. Cerebras’ own page says comparisons are based on third-party benchmarking or internal testing and may vary by workload, configuration, date, and model. The Register also cautions that Cerebras’ headline compute number depends heavily on sparsity, which it says generally does not help large language model inference in the same way it can help other workloads. The chip at the center of the launch is WSE-3T, with the “T” standing for Turbo, according to The Register. The outlet reports that the part doubles compute, memory fabric, and I/O bandwidth versus WSE-3, while using the same process technology, wafer area, transistor count, core count, and SRAM capacity. In The Register’s reading, that means WSE-3T is not a new piece of silicon so much as the existing wafer-scale engine operated more aggressively. The Register says each WSE-3T is rated at 250 petaFLOPS of AI compute, 44 GB of SRAM, 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity. It estimates that Cerebras may be running the silicon at 2.8 GHz, up from 1.4 GHz in the previous generation, though that clock estimate is The Register’s analysis rather than a directly stated company figure. Power delivery is the other major engineering claim. Cerebras says CS-4 places power delivery 0.5 millimeters from the processor, compared with roughly 50 millimeters on conventional GPU boards, which it says enables twice as much power to be delivered to the WSE-3T. The company says its Wafer-Scale Backpack combines the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics in a compact assembly with 50% fewer components, reducing deployment time from days to hours. Cerebras is also separating rack infrastructure from compute modules. The company says the PowerRack can be installed and facility-qualified before compute arrives, after which compute backpacks slide into place and connect to power, cooling, and data. Its programmable I/O subsystem is claimed to double I/O bandwidth, reduce latency, and link wafers within and across racks without a switch. The Register adds one important architectural detail for inference buyers: Cerebras is not trying to run the whole inference pipeline only on its own accelerators. The outlet reports that Cerebras has partnered with Amazon Web Services and AMD to offload prompt-processing work to AWS Trainium XPUs and AMD Instinct GPUs, leaving Cerebras’ chips to operate primarily as decode accelerators for inference. Who benefits: Cerebras benefits if operators value low-latency decode performance for very large models. AWS and AMD may also benefit if the disaggregated approach described by The Register brings more prompt-processing work onto Trainium and Instinct systems. Who's exposed: GPU-based inference systems are the comparison target in Cerebras’ positioning. Buyers are exposed to benchmark risk if they treat sparse peak figures or vendor-reported speedups as equivalent to workload-specific production performance.