Applied AI28/08/20265 min read

OpenAI's Jalapeño Chip Outperforms Nvidia Blackwell

OpenAI has just unveiled Jalapeño, its first in-house AI inference chip, and the early results throw down the gauntlet to Nvidia. An ASIC that promises up to 3.6x lower latency than Blackwell systems and nearly twice the performance per watt. The figures come from an independent benchmark. The chip is fast, the numbers make that plain. What matters is what it means for you, the one paying per token.

Mascota Marketing Ultra

TL;DR: No-Fluff Summary

  • Jalapeño is an ASIC: a chip designed by OpenAI and Broadcom exclusively for AI inference, not a general-purpose GPU.
  • Numbers vs. Nvidia: between 1.5x and 1.9x more performance per watt, and up to 3.6x lower latency than Blackwell systems (GB200/GB300).
  • API impact: in the medium term, faster responses and downward pressure on prices. Limited rollout by late 2026.
  • Availability in the UK/EU: unknown. The official source does not confirm regions.
Bottom line: Jalapeño isn't a product you'll buy, but if the numbers hold up in production, your API bill has a real shot at coming down.

Availability in the UK/EU: Unconfirmed

The source does not confirm availability in the UK/EU. Phase: general availability (GA).

What Is Jalapeño and Why OpenAI Is Building Its Own Silicon

Jalapeño is an application-specific integrated circuit (ASIC) designed by OpenAI in partnership with Broadcom, built exclusively for inference, the process by which a trained model generates a response to your query. It doesn't train models. It's not a general-purpose GPU. It's a piece of silicon that does ONE thing and does it very fast.

The difference from Nvidia's GPUs is conceptual. A GPU is a Swiss Army knife: it handles training, inference, and gaming. An ASIC is a scalpel. It cuts sharper, but only cuts. That's what allows it to squeeze more tokens per watt than any general-purpose GPU. Pure specialization.

Why now? Because OpenAI spends billions a year on compute and relies almost entirely on Nvidia to get it. As Bloomberg reported, the move is essentially a diversification strategy, a way of not putting all your eggs in one basket. Sensible. But it goes further: OpenAI wants to own the full stack, from model to metal.

Jalapeño vs. Nvidia Blackwell: The Numbers

OpenAI has published the first Jalapeño results on InferenceX, SemiAnalysis's public benchmark. Models evaluated: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. The declared figures are not trivial.

Jalapeño chip on a test bench with three benchmark gauges showing efficiency, latency and throughput advantages over a Blackwell GPU system in the background
  • Performance per watt: between 1.5x and 1.9x more AI work per watt versus Nvidia Blackwell systems (GB200 and GB300).
  • Latency: 1.7x to 3.6x lower end-to-end.
  • Interactive workloads: 2.1x to 4.1x higher throughput. This is what you feel when using ChatGPT or calling the API.
  • Against Vera Rubin: matches or beats Nvidia's next-generation accelerator (not yet on the market) in output tokens per megawatt.

The chip has a rated power draw of 700W. During testing it stayed at 550W or below.

One figure that says a lot about this moment: Jalapeño's full development, from design to production, took nine months. OpenAI used its own AI models to accelerate the process. The chip helping to design the chip. This is starting to look like something real.

What Changes for OpenAI API Users

If you're running OpenAI's API in production, the first thing you'll notice is latency. A chip that responds 1.7x to 3.6x faster translates into snappier agents and users who stop waiting. In real time, that matters. A lot.

The second-order effect is more interesting: downward pressure on costs. If OpenAI cuts its per-token inference spend, it has room to lower prices. Will it? That's another story. But competition with Anthropic, Google, and DeepSeek suggests that savings will, at least in part, make their way to the user. If you're already weighing when to use ChatGPT or Claude, each provider's proprietary silicon is another factor that's going to matter.

The timeline: limited rollout to data centers in late 2026, gradual expansion from 2027 onward. OpenAI will keep buying Nvidia accelerators. Jalapeño complements, it doesn't replace.

Availability in the UK / EU: unknown

OpenAI's official source does not confirm which regions have access to Jalapeño or whether European API users already benefit from this hardware. Global phase: general availability (GA).

The Bigger Play: Vertical Integration in AI

Jalapeño's numbers are impressive, but the real story is the strategic move behind them. Google has run its own TPUs (tensor processing units) for years, proprietary chips that give it independence from Nvidia and control over pricing and priorities. Amazon has Trainium and Inferentia. Meta is working on its own.

The mascot carefully mounts a small efficient chip in the foreground while a towering wall of overheating GPU server racks fills the dark background

OpenAI was the missing piece. And the heaviest one: one of the world's largest GPU consumers now building its own. The price gap between AI providers is going to depend more and more on who holds the most efficient silicon. In my experience, when a provider controls the entire chain, the first to benefit are high-volume customers. Smaller users take longer to see the discounts.

Watch for a side risk. If every major provider builds chips optimized for THEIR models, portability between APIs gets complicated. Today you can switch from OpenAI to Anthropic with relative ease. Tomorrow, if each model's performance depends on the proprietary silicon running it, migration carries an implicit cost that won't appear in any pricing table. If you want to understand how these infrastructure advances end up shaping real workflows, the AI automation guide at Marketing Ultra puts this transition in context.

Jalapeño won't change your day-to-day tomorrow. It's invisible infrastructure. But if the numbers hold up in real production, the war for your token is getting serious. Watch OpenAI's API prices over the next 12 months. And if they drop, you'll know what's underneath: a chip named after a hot pepper, ready for a fight.

Leave a comment

Your email will not be published. We review comments before showing them.