Nvidia used Hot Chips 2026 to lay out Vera, its next-generation server CPU for AI systems, according to ServeTheHome’s live coverage. The chip is based on Nvidia’s Olympus CPU architecture and is described as an in-house Arm server CPU designed for the company’s upcoming Vera Rubin AI systems. A separate SiliconANGLE report, carried by Techmeme, says Nvidia’s dedicated AI inference accelerator — named in that headline as “Groq 3 LPX” — has entered full production. The same summary says Nebius has signed on as the first customer, and that SpaceXAI will adopt Vera CPUs. Those production and customer details are not corroborated elsewhere in the provided material. There is one naming caveat. Techmeme’s SiliconANGLE summary refers to “Groq 3 LPX,” while ServeTheHome’s report refers to “Grorq LPX3” and “Groq LPX3” racks that use Groq dedicated hardware for decode acceleration. The cluster does not resolve whether those labels refer to the same product or reflect a typo or naming mismatch, so the safest reading is to attribute each name to the outlet using it. On the CPU side, ServeTheHome says Vera has 88 CPU cores and eight 128-bit LPDDR5X memory controllers. The chip includes NVLink-C2C, which ServeTheHome says is intended to pair with Nvidia Rubin GPUs, connect to other NVLink-C2C implementations, or link to another Vera chip for a two-socket CPU-only setup. ServeTheHome frames Vera as the successor to Nvidia’s Grace CPUs, which supported the Grace Hopper and Grace Blackwell generations. The report says Nvidia positioned Vera as part of a broader co-design effort across hardware and software for AI workflows, spanning Vera Rubin systems, NVL72, LPX3, Bluefield racks, and other components. The performance argument Nvidia presented, according to ServeTheHome, centers on the trade-off between token throughput and interactivity. Higher batching can raise token generation, but it can also slow response time for users. Nvidia’s claim is that Vera Rubin pushes that frontier outward, with up to 30 times the total throughput of Grace Blackwell at higher interactivity levels, depending on the point on the curve. ServeTheHome also reports Nvidia claims for Groq/LPX3 decode acceleration. In Nvidia’s framing, the LPX3 racks are meant to extend performance for customers that need more than Vera Rubin’s CPUs and GPUs alone can provide, particularly at high interactivity rates. The reported claim is that long-context decode is four times faster than current public services. The customer story is still early from the evidence provided here. SiliconANGLE’s reported Nebius and SpaceXAI details would make the announcement more than an architecture disclosure, but those claims remain single-source in this cluster. The technical story is clearer: Nvidia is presenting Vera as a CPU tightly coupled to its AI systems strategy, not as a general-purpose server chip in isolation. Who benefits: Nvidia benefits if customers adopt more of its in-house CPU and inference stack around Vera Rubin. If SiliconANGLE’s report is accurate, Nebius and SpaceXAI are among the early named adopters or customers. Who's exposed: Existing Grace and Grace Blackwell users may face a new upgrade path, but the provided sources do not quantify timing, pricing, or migration costs. The cluster does not establish specific competitive exposure for other chipmakers.