OpenAI's First Chip Beats Nvidia a Day Before Earnings
OpenAI's Jalapeño chip beat Nvidia Blackwell by up to 1.9x on power efficiency, independently verified, one day before Nvidia's earnings.
OpenAI picked the timing for this one deliberately. On Tuesday, August 25, at the Hot Chips conference in Silicon Valley, one of Nvidia's largest customers stood up and presented independently verified benchmark data showing its own custom chip beating Nvidia's current flagship hardware, exactly one day before Nvidia reported quarterly earnings that Wall Street would scrutinize for any sign of cracks in its dominance. The chip is called Jalapeño. The numbers behind it are real, measured, and confirmed by an outside party, not marketing copy.
What OpenAI Actually Showed
Richard Ho, OpenAI's head of hardware, presented the results in terms that left little room for hedging: "The bottom line is that the results show a very, very significant performance advance over state of the art." The claim rests on SemiAnalysis's InferenceX benchmark, a public testing framework that measures the full pipeline of serving an AI request, not just raw chip throughput in isolation, covering power consumption, latency, and total workload capacity together. Critically, SemiAnalysis did not simply take OpenAI's word for it. The firm's own engineers visited OpenAI's labs and ran the workloads themselves, a genuine independent verification rather than a vendor self-report dressed up as a study.
Tested across three publicly available models, GPT-OSS 120B, DeepSeek R1 670B, and the significantly larger Kimi K2.5 1T, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput than the Nvidia Blackwell-generation systems it was measured against, alongside 1.7 to 3.6 times lower end-to-end latency. On the specific kind of highly interactive, conversational workload that ChatGPT itself generates constantly, the gap widened further, to 2.1 to 4.1 times faster. Against an Nvidia GB200 system running GPT-OSS 120B specifically, OpenAI reported roughly 1.9 times higher peak throughput per kilowatt, 85,448 versus 44,960, with latency dropping from 1.80 seconds to 1.03 seconds.
The Comparison That Deserves Scrutiny
Here is where the story gets more complicated than the headline numbers suggest, and where the more careful outlets covering this story earned their credibility by saying so directly. The comparison measures Jalapeño against Nvidia's GB200 and GB300 Blackwell-generation racks, hardware that launched in 2024 and 2025, not against Nvidia's upcoming Vera Rubin platform, which represents the company's actual next-generation answer to exactly this kind of competitive pressure. Enterprise DNA's own analysis was blunt about the limitation: "These results don't mean Nvidia has lost its lead," it noted, while still crediting OpenAI's chip effort as "looking like a credible performance option, rather than simply insurance against relying on one supplier."
A separate analysis from 24/7 Wall St added a further technical caveat worth understanding: OpenAI's tests excluded speculative decoding, a technique that can meaningfully improve inference performance, and the company did not disclose full system-level power figures, meaning the comparison measures throughput more precisely than it measures true performance-per-watt in every configuration. The same analysis noted that AMD's and Nvidia's most recent hardware racks still deliver 1.46 to 2 times more raw compute and up to 12 percent more memory than OpenAI's rack design, at roughly 85 percent of its memory bandwidth, a genuine tradeoff rather than a clean, unambiguous win in every dimension. OpenAI also disclosed separately that internal testing on its own frontier models showed an even wider advantage than the public benchmark figures, a claim that, unlike the SemiAnalysis-verified numbers, no outside party has confirmed.
An Unusually Fast Path From Idea to Silicon
The engineering timeline behind Jalapeño is genuinely remarkable on its own terms, independent of how the final benchmarks stack up. OpenAI unveiled its chip partnership with Broadcom in June, but design work reportedly began in the middle of 2024, meaning the project went from initial team formation to actual manufacturing tape-out in roughly 16 months. SemiAnalysis's own technical newsletter noted that first-generation custom chips are typically not competitive against established incumbents, an industry pattern OpenAI's own results appear to have bucked. The chip was co-developed with Broadcom on the underlying silicon and networking architecture, and with systems integrator Celestica on physical assembly, part of a broader October 2025 agreement to jointly develop 10 gigawatts of custom AI accelerator capacity.
OpenAI's own explanation of the architecture emphasizes flexibility over narrow specialization: the design lets model state, including the key-value cache used while generating a response, be explicitly placed and kept local while the system dynamically activates the right combination of compute, memory, and networking for each specific phase of inference. That is a meaningfully different design philosophy than building a chip narrowly optimized for one workload type, and it is part of why OpenAI claims the chip performs well broadly rather than only in cherry-picked scenarios.
Why OpenAI Still Needs Nvidia, and Says So Plainly
The most important sentence in OpenAI's own announcement is not the benchmark numbers. It is the company's explicit commitment to keep buying from the competitor it just outperformed. "Meeting growing demand for AI will require more compute from every available source," OpenAI wrote in its own post. "We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads." That is not diplomatic hedging so much as a straightforward acknowledgment of scale reality: Jalapeño is only planned for deployment in "very small volumes" by the end of 2026, with meaningful scale-up not arriving until 2027, and a second generation already in development behind it. Nvidia's global supply chain and manufacturing capacity dwarf anything a first-generation custom chip program can match in the near term, regardless of how favorably its benchmarks compare on a conference stage.
This restraint is worth reading against the backdrop of Nvidia's own current position in the market, one where even the dominant chipmaker has recently had to pass memory-driven cost increases on to its largest customers rather than absorb them entirely. A credible, independently verified alternative chip does not need to replace Nvidia outright to matter. It simply needs to give a customer like OpenAI genuine negotiating leverage and a real fallback option, which is a meaningfully different competitive position than having no alternative at all.
The Timing Nobody at OpenAI Needed to Explain
Nvidia reported its fiscal second-quarter results after market close on August 26, the day immediately following OpenAI's Hot Chips presentation, in what CNBC's own live coverage framed as one of the most closely watched earnings reports of the year given persistent investor questions about the sustainability of current AI infrastructure spending. Nvidia's stock had closed at $213.05 on August 25, up 2.19 percent that session, meaning the market absorbed OpenAI's benchmark reveal calmly rather than treating it as an immediate threat. CNBC quoted analysts describing Nvidia's near-monopoly position as now facing genuine "threat" from the accumulating wave of custom silicon efforts across the industry, not just OpenAI's, but Google's, Amazon's, and Meta's parallel chip programs as well.
That framing connects directly to a broader pattern already reshaping how the largest AI companies think about chip supply. Google's own recent warrant arrangement with Marvell, tied to custom silicon purchases through 2033, reflects the same underlying calculation driving OpenAI's Jalapeño program: even companies that continue buying enormous volumes of Nvidia hardware are simultaneously investing heavily in reducing how dependent they are on any single supplier for the industry's single largest recurring cost.
What This Means for OpenAI's Own Financial Story
OpenAI's own infrastructure financing situation makes the stakes here more concrete than an abstract technology competition. Nvidia itself recently had to scale back a proposed data-center financing guarantee for OpenAI's Ohio campus after Nvidia's own investors balked at the scale of exposure involved, a reminder that even Nvidia's willingness to underwrite OpenAI's infrastructure ambitions has real limits. A credible in-house chip that cuts inference costs, the single largest ongoing operational expense in running a service like ChatGPT at global scale, gives OpenAI a genuine lever for improving its own margins independent of whatever financing arrangements it can or cannot secure from outside partners. Inference is not a one-time training cost. It runs every single time someone opens the app and sends a message, which means even a modest efficiency gain compounds into meaningful savings across billions of daily requests once deployment actually scales in 2027.
What Actually Changes, and What Doesn't
Nvidia's dominance in AI training remains entirely untouched by this specific announcement, since Jalapeño is purpose-built for inference only, the process of running an already-trained model rather than building one from scratch. What changes is more subtle and, in some ways, more consequential: one of Nvidia's largest customers has now demonstrated, with independent verification, that it can design and manufacture chips genuinely competitive with Nvidia's current commercial hardware on the metric that determines whether AI products can actually be profitable at scale. Whether that translates into real pricing leverage over Nvidia, a template other major AI labs accelerate their own custom silicon efforts around, or simply a well-timed public relations moment ahead of a competitor's earnings call, will become clearer once Jalapeño actually ships in meaningful volume in 2027, running alongside the very Nvidia hardware OpenAI has explicitly promised to keep buying regardless of how these benchmarks turned out.
Written by
Mr. Aayush Bhatt
Software Engineer interested in how models work and where they fail.