On August 25, at the Hot Chips conference, OpenAI published the first performance benchmarks for Jalapeño, its inference chip built with Broadcom — and the numbers beat systems running Nvidia's Blackwell. It's the first time the company has shown concrete data on the custom hardware it's been building since last year.
What Changed
According to OpenAI, tested on InferenceX — a public benchmark maintained by SemiAnalysis, an independent semiconductor analysis firm — Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Nvidia Blackwell (GB200/GB300) systems, with 1.7 to 3.6 times lower end-to-end latency. On highly interactive workloads — the low-latency, chatty traffic ChatGPT itself generates — the gain reached 2.1 to 4.1 times. The tests used three public models: GPT-OSS 120B (OpenAI's own), DeepSeek R1 670B, and Kimi K2.5 1T — meaning OpenAI benchmarked the chip against direct competitors' models too. Context matters here: the numbers are self-reported by OpenAI, following SemiAnalysis's public methodology, but still lack independent third-party verification. Jalapeño, unveiled in June after roughly nine months of development with Broadcom and Celestica, is only targeting low-volume production by late 2026, and hasn't been tested against Nvidia's upcoming Vera Rubin generation.
Why It Matters
The timing matters: Nvidia is about to report quarterly earnings, and the market is already reading this as a threat to the company's margins, per CNBC. But the more relevant data point for anyone buying AI as a service is different: even with a custom chip beating Nvidia on an initial benchmark, OpenAI says it will keep buying hardware from its rival. This isn't a vendor break — it's diversification. Companies that build their own hardware to cut inference cost (the most expensive input in running models at scale) gain negotiating leverage and reduce single-vendor dependency, without necessarily abandoning that vendor.
The Impact for Brazil
For Brazilian companies that consume AI via API — the vast majority outside of Big Tech — this kind of news doesn't change tomorrow's price, but it signals direction: major labs are investing heavily to cut their own inference cost, which has historically preceded more competitive pricing passed on to end customers (as we've seen with AT&T's model routing and DeepSeek's price repricing). The practical point of caution is different: benchmark numbers released by the vendor itself, even following a public methodology, deserve the same skepticism as any technical marketing material — it's worth waiting for independent validation before using this kind of announcement as a purchasing argument.
Entercast's Take
This is another chapter in a theme we've been tracking here: AI infrastructure — not just the model — has become the battleground that determines cost and vendor dependency. We've covered Anthropic taking on its own grid costs, Google Cloud formalizing identity and governance for agents, and AT&T cutting cost through model routing. Now it's the silicon layer: whoever depends on AI in production should watch this move not for the benchmark race itself, but for what it signals about where inference cost is headed over the next 12 to 18 months — and about the risk of concentration among a handful of chip vendors, a theme we've also covered here through the geopolitical fight over semiconductor supply chains.