OpenAI used the Hot Chips conference on Tuesday to give a more detailed look at Jalapeño, its inference-focused AI hardware system, according to TechCrunch. The report says OpenAI shared the first batch of benchmark results for Jalapeño, tested on Semianalysis’s InferenceX benchmark. The core claim is directional but significant: TechCrunch reports that Jalapeño delivered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors on that benchmark. The comparison was notably against an Nvidia Blackwell system, according to the report. OpenAI framed the result as an efficiency and latency story. TechCrunch quotes Richard Ho, OpenAI’s head of hardware, saying the results show a substantial performance advance over the current state of the art, with Jalapeño able to serve more AI work per unit of power while returning responses more quickly. The company is positioning the system for high-volume inference, where serving many users at low latency is the central workload. The timing matters. Ho estimated that Jalapeño would reach deployment at the end of 2026 in “very small volumes,” with more meaningful deployment coming in 2027, TechCrunch reports. That means the benchmark is being compared with currently available systems, while the competitive landscape may look different by the time Jalapeño is deployed at scale. TechCrunch says Jalapeño was first announced last October and was developed by OpenAI in close collaboration with Broadcom. The report also says OpenAI’s own models assisted in the development process. OpenAI plans to make Jalapeño a multigenerational platform, aligning AI products, models, chips and memory more closely over time. The technical emphasis is on reducing friction in inference rather than training. TechCrunch reports that OpenAI designed Jalapeño to address the prefill and communication phases of inference, which the company says can become bottlenecks. In OpenAI’s description, the system minimizes data movement and communication delays by keeping model state, including the key-value cache used while generating responses, local while activating the appropriate mix of compute, memory and networking for each inference phase. For now, the public evidence in this cluster is a single detailed TechCrunch report based on Hot Chips disclosures, benchmark claims and OpenAI comments. The claims are credible enough to follow, but the absence of independent corroboration or quoted numeric benchmark scores in the provided material makes this a developing hardware story rather than a settled performance verdict. Who benefits: OpenAI benefits if the system can improve latency and throughput per watt for its own AI services. Broadcom also stands to benefit from its reported role as OpenAI’s close development partner on Jalapeño. Who's exposed: Nvidia is the obvious comparison point because TechCrunch says Jalapeño’s benchmark was measured against a Blackwell system. The exposure is not immediate, however, because OpenAI’s significant deployment is not reported until 2027.