메뉴
BL
The Decoder • 31일 전

오픈AI 자체 칩 '할라페뇨', 엔비디아 블랙웰·루빈 추론 벤치마크 앞선다

IMP
8/10
핵심 요약

오픈AI가 핫칩스 컨퍼런스에서 공개한 자체 추론 칩 '할라페뇨'가 와트당 처리량과 토큰 지연시간에서 엔비디아 블랙웰은 물론 차세대 루빈까지 능가하는 것으로 보고됐습니다. 브로드컴과 9개월 만에 개발한 1세대 칩으로, 엔비디아의 'CUDA 해자'가 무너질 수 있다는 분석이 나옵니다. 다만 아직 엔지니어링 샘플 단계란 점은 유의할 필요 있습니다.

번역된 본문

오픈AI의 첫 자체 칩 '할라페뇨', 엔비디아 블랙웰과 루빈을 추론 벤치마크에서 능가한 것으로 보도됨 (Matthias Bastian, 2026년 8월 25일)

오픈AI가 핫칩스(Hot Chips) 컨퍼런스에서 자체 개발한 추론 칩의 첫 벤치마크 결과를 공개했다. '할라페뇨(Jalapeño)'는 와트당 처리량과 토큰 지연시간에서 엔비디아의 블랙웰(Blackwell)과 루빈(Rubin)을 모두 능가하는 것으로 알려졌다.

이 칩은 추론 전용으로, AI 모델을 실행하지만 학습은 하지 않는다. 또한 오픈AI 자체 모델에 특화된 것도 아니며, 범용 LLM 추론 가속기다.

오픈AI에 따르면 할라페뇨는 테스트된 세 모델 모두에서 최대 처리량 기준 와트당 AI 작업량을 1.51.9배 더 처리하며, 상업적으로 이용 가능한 최고 시스템 대비 종단 간(end-to-end) 지연시간은 1.73.6배 낮다. 대화형 워크로드에서는 성능이 2.1~4.1배 높다고 회사는 밝혔다.

이 결과는 SemiAnalysis의 공개 벤치마크 'InferenceX'를 이용한 테스트에서 나왔다. 수치는 오픈AI가 제공했으며, SemiAnalysis가 실험실에서 일부 실행을 현장 검증했다. 테스트 모델은 GPT-OSS 120B, Deepseek R1 670B, Kimi K2.5 1T다. GPT-OSS에서는 사용자당 초당 약 1,400토큰을 기록했고, Deepseek R1에서는 단일 동시 요청 기준 초당 700토큰을 넘겼다.

할라페뇨는 멀티 토큰 예측(multi-token prediction)이나 추론 디코딩(speculative decoding) 같은 기법을 쓰지 않고도 이런 수치를 냈다. 일부 비교 대상 시스템은 이런 최적화를 활용했으므로, 할라페뇨에는 아직 개선 여지가 남아 있다.

와트당 성능 비교에서 "할라페뇨는 다른 모든 칩을 압도한다"라고 SemiAnalysis는 평가했다. SemiAnalysis의 딜런 패텔(Dylan Patel) CEO는 "보통 1세대 칩은 경쟁력이 없지만, 오픈AI는 엔비디아 블랙웰은 물론 루빈까지 이기고 있다"라고 덧붙였다.

SemiAnalysis는 더 공정한 비교 대상은 블랙웰이 아니라 둘 다 HBM4 메모리를 쓰는 엔비디아의 최신 베라 루빈(Vera Rubin) 플랫폼이라고 지적한다. 여기서도 할라페뇨는 엔비디아 가속기가 아직 도입하지 않은 멀티 토큰 예측 최적화를 루빈이 쓰고 있음에도 메가와트당 더 많은 출력 토큰을 뽑아낸다. 토큰당 총소유비용(TCO)에서는 양측이 대등한 수준이다.

단, 주의할 점도 있다. 엔비디아와 AMD는 이미 Deepseek V4 Pro, Kimi K3 같은 더 큰 모델로 결과를 발표했지만 이 모델들은 아직 할라페뇨에서 테스트되지 않았다. 또 루빈 시스템은 이미 고객에게 출하 중인 반면, 할라페뇨는 엔지니어링 샘플 단계를 벗어나지 못한 것으로 알려졌다.

9개월 만에 개발, 자체 모델 활용도

오픈AI는 브로드컴(Broadcom)과 함께 할라페뇨를 개발했다. 설계 작업은 2024년 중반에 시작됐고 최종 설계는 2025년 11월에 파운드리로 넘어갔다. 전체 주기는 약 16개월이었지만, 오픈AI는 첫 칩 설계부터 완성된 설계도가 공장으로 향하기까지 9개월만에 끝냈다고 말한다. 개발 과정에서 자체 AI 모델을 사용했다고 오픈AI는 밝혔다. 구세대 모델은 칩 설계를 도왔고, 신세대 모델은 프로그래밍과 최적화를 가속했다.

SemiAnalysis는 이를 통해 엔비디아의 'CUDA 해자(moat)'가 더 이상 유효하지 않을 수 있다는 시그널로 본다. "오픈AI가 자체 실리콘에서 얼마나 빠르게 새 모델을 올리는지를 보면 CUDA 해자는 사실상 끝났을 수 있다"라고 해당 업체는 작성했다.

오픈AI의 세라 프라이어(Sarah Friar) CFO는 이 칩이 데이터센터, 칩, 모델, 개발자 플랫폼, 제품, 기기가 하나의 통합 시스템으로 작동하는 더 포괄적인 컴퓤팅 전략에 부합한다고 말했다. 그녀는 할라페뇨가 엔비디아, AMD, AWS, 세레브라스(Cerebras), 코어위브(CoreWeave) 등과의 기존 파트너십을 대체하는 것이 아니라 보완한다고 주장했다.

오픈AI는 이들 기업 여러 곳과 깊은 관계를 맺고 있다. 엔비디아, AMD, AWS는 모두 투자자 또는 컴퓤팅 파트너이며, 엔비디아는 그중 최대 규모다. 이들은 각자 자체 AI 칩을 개발 중이어서 관계는 협력이자 경쟁이기도 하다. 다만 이 모든 기업들은 세상에 컴퓨팅은 결코 충분하지 않다고 계속 말하는데, 이는 자사 비즈니스 모델에 편리하게 부합하는 주장이기도 하다.

과장 없는 AI 뉴스 – 인간이 엄선합니다

원문 보기
원문 보기 (영어)
OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 25, 2026 OpenAI Ask about this article… Search OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. "Jalapeño" reportedly outperforms both Nvidia's Blackwell and Rubin in throughput per watt and token latency. The chip handles inference only, meaning it runs AI models but doesn't train them. Jalapeño isn't tuned to OpenAI's own models either. It's a general-purpose LLM inference accelerator. OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher. Ad The results come from tests using SemiAnalysis's public InferenceX benchmark . OpenAI provided the numbers. SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño hit about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request. Ad Jalapeño posted these numbers without using techniques like multi-token prediction or speculative decoding , while some of the comparison systems did rely on those optimizations, so there's still room to improve. In its headline performance-per-watt comparison, "Jalapeño smokes every other chip," SemiAnalysis writes . SemiAnalysis CEO Dylan Patel added , "Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin." Ad SemiAnalysis points out that the fairer comparison isn't Blackwell but Nvidia's newer Vera Rubin platform , since both use HBM4 memory. Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization that Jalapeño hasn't adopted yet. On total cost of ownership per token, the two come out roughly even. There are caveats, though. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. And while Rubin systems are already shipping to customers, Jalapeño reportedly hasn't moved beyond engineering samples. Ad OpenAI built the chip in nine months, partly using its own models OpenAI developed Jalapeño with Broadcom. Design work kicked off in mid-2024, and the final design went to fabrication in November 2025. The full cycle took about 16 months, but OpenAI says only nine months passed between the first chip design and the finished blueprint heading to the factory. The company used its own AI models during development, according to OpenAI . Older model generations helped with chip design, while newer ones sped up programming and optimization. Ad SemiAnalysis sees this as a sign that Nvidia's much-discussed "CUDA moat" may not hold anymore. "The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon," the firm wrote. OpenAI CFO Sarah Friar says the chip fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system. She claims Jalapeño complements OpenAI's existing partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and others rather than replacing them. OpenAI has deep ties with several of these companies. Nvidia , AMD , and AWS are all investors or compute partners, with Nvidia being one of the largest. Each of them is also building its own AI chips, which makes the relationship both cooperative and competitive. That said, all of these companies keep saying the world can't have enough compute , a claim that conveniently supports their own business models. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI | SemiAnalysis