메뉴
HN
Hacker News • 31일 전

OpenAI '할라페뇨' 칩, 엔비디아 블랙웰 성능 능가

IMP
9/10
핵심 요약

OpenAI가 브로드컴과 함께 약 16개월 만에 개발한 자체 추론 ASIC '할라페뇨(Jalapeño)'가 Hot Chips에서 공개되었습니다. 특정 모델에 특화되지 않은 범용 추론 칩임에도, MTP 없이 엔비디아·AMD·구글 칩을 테스트한 모든 오픈소스 모델에서 perf/W(전력 대비 성능)와 토큰 처리량을 능가했습니다. 1세대 칩으로는 이례적인 경쟁력으로, AI를 활용한 칩 설계 가속화가 실제임을 보여주는 사례입니다.

번역된 본문

OpenAI '할라페뇨': 엔비디아 블랙웰보다 뛰어난 성능

OpenAI의 자체 설계 ASIC를 루빈(Rubin)과 비교하고, 할라페뇨의 TCO, MW당 처리량, 그리고 매운(?) 세부 정보들을 다룹니다.

브라이언 샨, 마이런 시, 조던 나노스 외 3인 / 2026년 8월 25일 / 유료 콘텐츠

OpenAI는 지난 몇 년간 Hot Chips에서 막 공개된 추론 칩 '할라페뇨(Jalapeño)'를 조용히 개발해왔습니다. 테이프아웃 성공 소문이 한동안 돌았지만, 이제 실제 세부 사양이 공개되었습니다. OpenAI는 우리를 초청해 칩을 직접 보여주고, 연구실에서 실제 여부를 확인하게 했으며, 자체 벤치마크 스위트 'InferenceX'로 성능을 측정하게 했습니다.

6월, OpenAI는 브로드컴과 파트너십으로 이 칩 프로그램을 공개했는데, 백지 상태에서 LLM 추론 전용으로 설계되었습니다. 설계는 2024년 중반에 시작되어 초기 팀 채용부터 제조 테이프아웃까지 약 16개월 만에 완료된, 매우 빠른 ASIC 개발 주기였습니다.

일반적으로 1세대 칩은 경쟁력이 없지만, OpenAI는 업계 최고 수준을 보여주며 테스트 가능했던 모든 엔비디아, AMD, 구글 칩을 여러 주요 오픈소스 모델에서 능가하며 이런 통념을 깼습니다. OpenAI는 이를 극단적인 하드웨어-소프트웨어 동시 설계(co-design)로 달성했습니다. 흥미롭게도 모델 추론의 특정 부분에 과도하게 특화한 것이 아니라, 모든 시나리오에서 고성능을 내는 범용 칩에 집중했습니다.

이 글에서는 할라페뇨의 아키텍처 세부 사항, 소프트웨어 세부 사항, InferenceX 성능 결과를 다룹니다.

범용 추론 칩

많은 이들이 OpenAI의 칩이 OpenAI 모델에 특화되었다고 말하지만 그것은 틀렸습니다. OpenAI는 AI 추론을 위한 범용 칩을 만들었습니다. 개발 일정이 놀라울 정도로 빨랐는데, 이는 AI 활용이 칩 설계를 가속화한다는 주장이 실제임을 보여줍니다. 빠른 일정에도 불구하고, OpenAI는 막대한 자금을 투입했고 실용적인 설계 결정을 내렸으며 팀이 뛰어났기에 이는 놀라운 일이 아닙니다.

스펙만 봐도 즉시 경쟁자입니다. HBM4 사용 덕분에 엔비디아와 AMD의 플래그십 GPU와 비교 가능한 수준입니다.

이 칩에 대한 언론 보도 상당수는 OpenAI가 다른 칩과 달리 자사 모델에 최화될 것이라는 몇 마디 던진 발언을 따라갔지만, 이는 틀렸습니다. 할라페뇨는 모든 종류의 모델과 워크로드를 실행할 수 있는 범용 추론 칩이며, OpenAI 엔지니어들과 함께 연구실에서 실행한 당사 벤치마크 InferenceX에서도 작동했습니다. 농담으로 OpenAI는 Codex 프롬프트만으로 이 칩에 포팅된 둠(Doom)을 실행하는 모습까지 보여줬습니다.

다음은 당사의 대표 성과인 perf/W 결과로, All-in 유틸리티 MW당 토큰 처리량을 나타냅니다. 할라페뇨는 다른 모든 칩을 압도합니다. 이 모든 것이 Multi Token Prediction(MTP) 없이 달성된 반면, 차트의 다른 칩들은 각 SKU의 최고 성능 구성으로 모두 MTP를 사용한 결과입니다.

할라페뇨는 곡선상 특정 지점에 맞춰 튜닝되지 않았음에도 거의 모든 시나리오에서 perf/W 기준 블랙웰을 능가합니다. 저지연 시나리오뿐 아니라 고처리량 시나리오에서도 뛰어납니다.

더 공정한 비교는 Single Token Prediction 결과로, 모든 경쟁사를 압도합니다. 저동시성 시나리오에서 할라페뇨는 놀라운 반응성을 보여주며, DeepSeek R1 모델에서 동시성 1 기준 사용자당 초당 700토큰을 넘습니다. 놀랍게도 이는 모두 단일 토큰 예측(STP)만으로, 추측 디코딩(speculative decoding)이나 프리필-디코드 분리(prefill-decode disaggregation) 없이 달성되었습니다.

DeepSeek R1 외에도 Kimi-K2.5, GPT-OSS 등 다른 모델도 확인했으며, GPT-OSS는 약 초당 1,400토큰/사용자로 실행되었습니다. 모든 모델에서 할라페뇨의 GSM8k 평가 결과는 엔비디아 칩과 동등한 수준임을 확인했습니다.

몇 가지 유의사항이 있습니다. 첫째, 모든 수치는 OpenAI가 제공한 것입니다. 당사는 연구실에서 InferenceX 실행을 직접 검증했지만, 전체 InferenceX 벤치마크 스위트를 실행하지는 않았으며 AgentX 결과도 보지 못했습니다. AgentX는 데이터셋의 긴 컨텍스트와 다중 턴 특성 때문에 칩 성능 비교를 위해 선호하는 스위트입니다.

원문 보기
원문 보기 (영어)
OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets Bryan Shan , Myron Xie , Jordan Nanos , and 3 others Aug 25, 2026 ∙ Paid 104 10 Share OpenAI has spent the past couple years quietly building “Jalapeño,” an inference chip just announced at Hot Chips. Rumors of a successful tapeout had been swirling for a while. But now we have details. OpenAI invited us to look at their chip, go to their labs to check out how real it is, and benchmark it with our InferenceX suite. In June, OpenAI unveiled the chip program in partnership with Broadcom, built from a blank slate exclusively for LLM inference. Design work began in the middle of 2024 , going from initial team hiring to manufacturing tape-out in ~16 months, an extremely fast ASIC development cycle. In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models. OpenAI does this with extreme hardware software codesign. Surprisingly, OpenAI is not over specialization on any specific part of model inference, but instead by focusing on being a general chip that delivers high performance in all scenarios. In this article, we will go into architectural details, software details and performance results for Jalapeño on InferenceX. A generalized inference chip Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong, OpenAI made a generalized chip for AI inference. The timelines are insane. It shows that claims that use of AI is being used to accelerate chip design are real. Regardless of the quick timelines,Open AI spent a bunch of money, made pragmatic design decisions and their team is cracked, so this comes as no surprise. Just looking at the specs, it is an immediate contender: And the use of HBM4 makes it stand out as comparable to flagship GPUs from NVIDIA and AMD: A lot of the media coverage of this chip has followed a few throwaway comments from OpenAI that claim the chip will be optimized for their models in a way that other chips are not. This is wrong. Jalapeño is a generalized inference chip capable of running all sorts of models, and all sorts of workloads, including our benchmark InferenceX, where we ran the benchmark with OpenAI engineers in the lab. As a joke, OpenAI even showed us it running Doom, which was ported to their chip with just Codex prompts. The following is our headline perf/W result, looking at token throughput per All-in utility MW. Jalapeño smokes every other chip . All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP. Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve. It excels not only in low-latency scenarios but also in high-throughput scenarios. A more apples to apples comparison is against Single Token Prediction results, it knocks every competitor out of the water. At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. Incredibly, this is all achieved with single-token prediction (STP), no speculative decoding and no prefill-decode disaggregation. In addition to DeepSeek R1, we also got to see some other models, including Kimi-K2.5 and GPT-OSS which ran at approximately 1,400 tok/sec/user. For all models, we confirmed that Jalapeño’s GSM8k evals attained results on par with Nvidia chips. Some caveats on this. First, all numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results. AgentX is our preferred suite for comparing chip performance due to the datasets’ long context and multi-turn characteristics that reflect the cache behavior of realistic production workflows. Frameworks that perform well on 8k1k may perform worse on AgentX as real production loads stress components like routers, prefix cache mechanisms, cache management, offload infrastructure, etc. These are not tested by single turn 8k1k. Read more about this in out AgentX article. AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? Cam Quilici , Bryan Shan , and 5 others · Aug 24 Read full story Second, we believe that comparison to Blackwell is somewhat incomplete and unfair. Jalapeño is really competing against chips like Rubin that also use HBM4. Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño. Thus, performance should really be compared against Rubin, not Blackwell, and in some sense we expect a custom chip like Jalapeño to outperform Blackwell. Vera Rubin NVL72 delivers 5.4x the perf/MW of GB200 NVL72 as we described in our article analyzing the NVIDIA performance claims in their launch with CoreWeave last month . We will compare Jalapeño to Vera Rubin’s July performance figures later below. Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis Alec Ibarra , Bryan Shan , and 6 others · Jul 23 Read full story Third, the models being tested are not on the open frontier. NVIDIA and AMD have published results on larger models such as DeepSeek V4 Pro and Kimi K3, using AgentX. The larger the model and the more recent the release, the more complicated it is to bring up on a new chip. With that said the models OpenAI has working on Jalapeno aren’t exactly small either. Performance Analysis OpenAI designs for perf/W. The reason is simple: OpenAI is currently limited by datacenter power, not by budget or floorspace, and thus tokens per MW is paramount. At Computex 2026, Jensen said that perf/W, reliability and long lifetime are the core features of future GPUs. To quote: “If you have 1 gigawatt of power, then throughput per watt is revenue”. He also mentioned that choosing the wrong architecture just because the chips are cheaper doesn’t make sense. This was emphasized by Nvidia during the Vera talk at Hot Chips 2026 while showing the same revenue graph: “The data center is power limited today.” Power matters and drives revenue. Operators cannot simply obtain more MW because adding GPUs and adding grid capacity happen on very different timescales. Datacenter power envelopes have constraints such as their utility interconnection, infrastructure, cooling capacity, and UPS/backup-generation design. Grid delays repeatedly outpace hardware and construction timelines, driving the need for BtM (behind-the-meter) power capacity: gas turbines and on-site generators built and located at the data center itself. This capacity sits behind the utility’s meter rather than being drawn from the public grid. It lets an operator power a facility without waiting on grid interconnection and utility upgrades, which is exactly why xAI’s Colossus 2 relies so heavily on BtM while its actual grid connection lags far behind. Find out more in our Energy model . As we wrote in an X post, tok/s/MW reduces to tokens per joule since a watt is a joule per second. This makes tok/s/MW representative of a system’s efficiency and ability to convert energy into tokens. On this front, even when compared with Rubin, Jalapeño wins. OpenAI’s Jalapeño has STP output token throughput per MW surpassing Vera Rubin’s MTP results that NVIDIA and CoreWeave published in July. It also far exceeds GB200’s 2025 MTP results. As mentioned in our Vera Rubin article, VR was compared to 2025 GB200 results because that was a similar stage of early bring-up, and comparing to GB200 in 2025 holds software maturity constant. Following this logic, we compare Vera Rubin’s latest July 2026 results, GB200 2025 results, and today’s Jalapeño results. This is a very va