메뉴
BL
TechCrunch AI • 31일 전

오픈AI '할라페뇨' 칩, 벤치마크서 최신 추론 칩 능가

IMP
8/10
핵심 요약

오픈AI가 Hot Chips 컨퍼런스에서 브로드컴과 공동 개발한 자체 AI 추론 칩 '할라페뇨(Jalapeño)'의 첫 벤치마크 결과를 공개했다. SemiAnalysis의 InferenceX 벤치마크에서 엔비디아 블랙웰 대비 사용자당 토큰 수와 킬로와트당 처리량을 모두 앞섰으며, 2026년 말 소량 배포 후 2027년 본격 배포될 예정이다.

번역된 본문

오픈AI는 화요일 Hot Chips 컨퍼런스에서 '할라페뇨(Jalapeño)'에 대한 보다 상세한 정보와 새 시스템의 첫 벤치마크 결과를 공개했다. SemiAnalysis의 InferenceX 벤치마크 테스트에서 할라페뇨는 현재 판매 중인 최고 수준의 추론 프로세서들보다 사용자당 더 많은 토큰을 처리했고, 킬로와트당 처리량도 더 높았다.

오픈AI의 하드웨어 총괄 리처드 호(Richard Ho)는 보도자료 브리핑에서 "결론적으로, 이 결과는 현존 최고 수준 대비 매우, 매우 큰 성능 향상을 보여준다"며 "할라페뇨는 단위 전력당 더 많은 AI 작업을 처리하면서 동시에 더 빠르게 응답을 반환한다. 많은 고객에게 서비스를 제공하는 데 매우 효율적이면서 동시에 매우 낮은 지연시간도 달성할 수 있다"고 말했다.

주목할 점은 이 비교가 엔비디아 블랙웰 시스템을 기준으로 한 것이지만, 할라페뇨가 전면 배치되는 시점에는 경쟁 제품도 상당히 발전했을 수 있다는 것이다. 호는 할라페뇨가 2026년 말 '아주 적은 물량'으로 배치되기 시작해 2027년에 본격적인 배치가 이뤄질 것으로 추정했다.

작년 10월 처음 발표된 할라페뇨는 오픈AI가 브로드컴과 긴밀히 협력해 개발했으며, 개발 과정에는 오픈AI 자체 모델도 활용됐다. 회사는 할라페뇨를 여러 세대에 걸친 플랫폼으로 만들 계획이며, 이를 통해 AI 제품, 모델, 칩, 메모리가 모두 유기적으로 함께 개발될 수 있다.

이러한 풀스택 방식 덕분에 오픈AI는 추론 처리 중 마찰을 일으키는 특정 단계들을 집중적으로 개선할 수 있었다. 특히 할라페뇨는 처리 과정에서 병목으로 작용하는 프리필(prefill)과 통신 단계의 지연을 최소화하도록 설계됐다.

오픈AI는 결과를 발표한 블로그 포스트에서 "우리는 데이터 이동과 통신 지연을 최소화하도록 할라페뇨를 설계했다"며 "이는 응답 생성 시 사용되는 KV 캐시를 포함한 모델 상태를 명시적으로 로컬에 배치·유지하면서, 시스템이 각 추론 단계마다 컴퓨팅·메모리·네트워킹의 적절한 조합을 활성화할 수 있음을 의미한다"고 설명했다.

원문 보기
원문 보기 (영어)
At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors. “The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency.” Notably, that comparison is against an Nvidia Blackwell system — but by the time Jalapeño reaches full deployment, the competition may have advanced significantly. Ho estimated that Jalapeño would deploy at the end of 2026 “in very small volumes,” with more significant deployment coming in 2027. First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI's own models assisting in the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory all developed in concert. Because of that full-stack approach, OpenAI was able to address specific phases in the inference process that often cause friction during inference processing. In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says often act as bottlenecks. “We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post presenting the results. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.” Topics AI , Broadcom , OpenAI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco In less than 48 hours, your chance to save up to $300 on your tickets will end! REGISTER NOW Most Popular Two years after launch, Walmart's Flipkart is closing in on India's quick-commerce leaders Jagmeet Singh Inherent, founded by DeepMind alumni, says its AI ‘teammate' just outperformed Anthropic and OpenAI at replicating research Anna Heim Michael Polansky is training an AI model on skin that’s still alive Connie Loizos How AI accounting startup Rillet raised $100M and became a unicorn in 48 hours Dominic-Madori Davis Tesla’s solar roof is dead — here’s what went wrong Tim De Chant Oura faces lawsuit accusing it of misleading consumers about sleep-tracking accuracy Aisha Malik Home batteries are suddenly cheap and everywhere. Here’s why. Tim De Chant