Onset

OpenAI’s Chip Answers Faster With Less Power Than Nvidia’s

OpenAI’s Chip Answers Faster With Less Power Than Nvidia’s

Working Jalapeño samples led Nvidia systems in selected answer-generation tests, and OpenAI plans to start installing the chips by year-end.

A posed indoor photograph of Sam Altman, at left in a gray sweater, and another man in a dark suit, apparently Richard, jointly holding a large circular semiconductor wafer mounted on a display labele
via @wallstengine

Latest update

SemiAnalysis puts Jalapeño up to 2× ahead of Rubin

SemiAnalysis now says Jalapeño delivered up to twice the performance per watt of Nvidia’s July Vera Rubin NVL72 result. Jalapeño ran without speculative decoding, while Rubin used it.

That strengthens the current-generation comparison, though it remains SemiAnalysis’s analysis of published results rather than an independent head-to-head production test.

View 3 earlier updates

OpenAI pairs Jalapeño with Cerebras

OpenAI core-products lead Thibault Sottiaux now says Jalapeño is meant to raise the speed and capacity available broadly, while Cerebras continues powering the fastest service for demanding customers. He expects tomorrow’s Fast tier to feel like today’s Ultrafast, but gave no date, price or performance commitment.

Cerebras said Jalapeño and its next-generation system are expected to operate together in production in 2027. That clarifies that OpenAI’s own chip is not replacing its specialist inference supplier. The company is building separate layers of computing for everyday scale and maximum speed, backed by an existing agreement to add 750 megawatts of Cerebras capacity through 2028.

Jalapeño narrowly tops Rubin on power efficiency

SemiAnalysis now puts Jalapeño at about 1.35 million output tokens per second per utility megawatt on DeepSeek R1, just above Nvidia’s Vera Rubin peak near 1.34 million.

SemiAnalysis_ @SemiAnalysis_

OpenAI's new Jalapeño chip announced at Hot Chips beats Vera Rubin's July results on Output Throughput per MW. (1/7)🧵

That closes one gap in the original comparison: Jalapeño’s efficiency lead is no longer confined to Nvidia’s older GB200 and GB300 systems. But it remains a comparison of OpenAI’s engineering-chip results with Rubin’s published July figures, not independent production testing, and the narrow margin could change as both systems mature.

OpenAI ties Jalapeño to ChatGPT and Codex

OpenAI now says Jalapeño will bring faster ChatGPT responses, more responsive Codex sessions and agents, and more reliable access as demand grows. That supplies the user-facing promise missing from its initial results, though the benefits remain prospective until deployment begins.

The company also says a second-generation chip is deep in development and a third is taking shape, turning Jalapeño from a one-off alternative to Nvidia hardware into the start of a continuing chip program.

Full story

Sam Altman gripped one side of a dinner-plate-sized silicon wafer while another OpenAI executive held the other. Beneath it, a small plaque read “Jalapeño Intelligence Processor.”

For a company whose rise rested on Microsoft’s data-center capacity and Nvidia’s chips, the object marked a new kind of control. OpenAI now has working engineering chips of its own generating model answers.

Jalapeño is built for inference, the work a trained model does when it turns a prompt into an answer. OpenAI designed the chip and its software together, aiming to produce more answers from every available megawatt.

What OpenAI measured

OpenAI tested Jalapeño with three public models: GPT-OSS, DeepSeek R1 and Kimi K2.5. Each run fed the model 8,000 input tokens and requested 1,000 output tokens, the small pieces of text models read and write.

Across those tests, Jalapeño processed 1.5 to 1.9 times as much text per second for each kilowatt at the chips’ most efficient settings. Requests finished 1.7 to 3.6 times faster, while fast interactive settings produced 2.1 to 4.1 times as much output per kilowatt as Nvidia’s GB200 or GB300 systems.

Jalapeño and Nvidia GB200 power efficiency across GPT-OSS answer speeds
via @alcreon

The 104.3× figure attached to one result comes from a particular operating point: both chips were held to GB300’s previous-best DeepSeek R1 answer speed. It is not the peak-to-peak comparison; at each chip’s most efficient point in that test, OpenAI reported a 1.7× Jalapeño lead.

Whether independent, production-like testing against Nvidia’s newer Vera Rubin system preserves that lead remains unknown: OpenAI supplied every published number, SemiAnalysis observed but did not run the full suite, and no result exists yet for AgentX, the long-context, multi-turn test designed to resemble agent use.

Power is the constraint

A data center cannot add another rack when it has no electricity left to feed it. More output per kilowatt lets OpenAI serve more answers inside the same power supply, capacity without first securing another megawatt.

Designing both Jalapeño and its serving software gives OpenAI control over both sides of that equation. It also provides another source of computing capacity and more choice among suppliers.

The chip does not replace the processors used to train models. OpenAI says Nvidia remains important for training and other work that Jalapeño does not cover.

OpenAI plans to start installing Jalapeño by year-end. It announced no faster or cheaper ChatGPT or Codex service. For now, the benefit sits inside OpenAI’s data centers: more answers from the electricity already reaching them.