OpenAI가 브로드컴과 함께 대규모 언어 모델(LLM) 추론에 최적화된 첫 번째 자체 칩인 '할라피뇨(Jalapeño)'를 공개했습니다. 이 칩은 데이터 이동을 최소화하고 이론적 최고 성능에 가까운 활용도를 끌어내도록 처음부터 새롭게 설계되었으며, 초기 테스트에서 현재 최고 수준의 제품보다 월등히 높은 전력 효율을 보여줍니다. 이는 OpenAI가 인프라부터 모델, 칩까지 아우르는 풀스택 전략을 완성하여 기가와트 규모의 데이터센터에 배포하려는 중요한 첫 단계입니다.
번역된 본문
2026년 6월 24일
기업
OpenAI와 브로드컴, LLM 최적화 추론 칩 공개
로딩 중... 공유
초기 테스트 결과, 1세대 가속기는 현재 최고 수준(SOTA) 제품보다 상당히 뛰어난 성능 대비 전력 효율을 제공할 것으로 보입니다.
산업계의 현재 및 미래 LLM을 위해 처음부터 새롭게 설계되었습니다.
OpenAI의 모델에 힘입어 설계부터 생산까지 단 9개월 만에 개발되었습니다.
제품, 모델을 넘어 이제 칩에 이르기까지 OpenAI의 풀스택 플랫폼을 확장합니다.
여러 세대에 걸쳐 데이터센터 파트너들과 함께 기가와트(GW) 규모로 배포될 예정입니다.
OpenAI와 브로드컴(NASDAQ: AVGO)은 오늘 OpenAI의 첫 번째 지능형 처리자(Intelligence Processor)인 '할라피뇨(Jalapeño)'를 공개했습니다. 이는 OpenAI가 그리는 LLM 추론의 미래를 중심으로 설계된 가속기이자, 두 회사가 선진 AI를 더 빠르고, 신뢰할 수 있으며, 더 많은 사람들이 접근할 수 있도록 만들기 위해 공동으로 구축하는 다세대 컴퓨팅 플랫폼의 첫 번째 AI 가속기입니다.
브로드컴의 확탄(Hock Tan) 사장 겸 CEO와 찰리 카와스(Charlie Kawwas) 사장이 OpenAI의 샘 알트만(Sam Altman) CEO와 그렉 브록먼(Greg Brockman) 사장에게 할라피뇨를 전달했습니다. 이는 OpenAI가 자사 모델 및 제품의 기반이 되는 풀스택을 직접 구축하려는 전략에 있어 중요한 이정표입니다. OpenAI는 모델, 커널, 서빙 시스템 및 제품 요구사항에 대한 로드맵을 바탕으로 LLM 기본 원리에 대한 깊은 이해를 바탕으로 이 칩을 처음부터 새롭게 설계했습니다. 파트너사인 브로드컴과 셀레스티카(Celestica)는 칩 구현, 보드 및 랙 시스템 통합, 고성능 네트워킹, 확장 가능한 생산 시스템을 통해 해당 플랫폼의 산업화를 지원했습니다.
할라피뇨는 산업계 전반의 현재 및 미래 AI 모델에 대한 추론 요구사항에 대한 OpenAI의 통찰력을 바탕으로 모든 LLM과 유연하게 작동하도록 설계되었습니다. 할라피뇨 칩의 엔지니어링 샘플은 현재 생산 목표 주파수 및 전력으로 실험실에서 GPT-5.3-Codex-Spark를 포함한 머신러닝 워크로드를 구동하고 있습니다. OpenAI는 여전히 최종 성능을 측정하고 있지만, 초기 테스트 결과에 따르면 할라피뇨는 현재 최고 수준 제품들보다 상당히 우수한 전력 대비 성능 비율(per watt)을 제공할 것으로 예상됩니다. 성능에 대한 자세한 기술 보고서는 향후 몇 달 내에 공개될 예정입니다.
이 아키텍처는 데이터 이동을 줄이고 컴퓨팅, 메모리 및 네트워킹 리소스의 균형을 맞춰 실제 활용도를 이론적 최고 성능에 훨씬 더 가깝게 끌어올립니다. 브로드컴의 실리콘 구현 및 토마호크(Tomahawk) 네트워킹 실리콘을 포함한 네트워킹 기술은 이 플랫폼을 대량 생산 수준으로 끌어올리는 데 도움을 줍니다.
OpenAI의 공동 창립자이자 사장인 그렉 브록먼은 "세계는 컴퓨팅이 주도하는 경제로 이동하고 있다"며, "할라피뇨는 컴퓨팅 자원을 더욱 풍부하게 만들어 사람과 기업을 위해 더 빠르고 신뢰할 수 있으며 저렴한 AI를 제공하고, 더 중요한 문제를 해결할 수 있도록 돕는 장기적인 풀스택 인프라 전략의 일환이다. 스택의 더 많은 부분을 직접 설계함으로써 더 높은 효율로 더 많은 지능을 서비스하고, 고급 AI를 더 많은 사람들이 사용할 수 있도록 계속 밀어붙일 수 있다"고 말했습니다.
OpenAI의 하드웨어 프로그램을 이끌고 있는 리처드 호(Richard Ho)는 "할라피뇨는 OpenAI 연구원들과의 긴밀한 협력을 통해 얻은 세부적인 통찰력을 활용하여 LLM 추론을 위해 처음부터 새롭게 설계되었다"며, "최첨단 AI 모델에 가장 중요한 커널, 메모리 이동, 네트워킹 및 서빙 패턴에 맞게 아키텍처를 최적화했다. 초기 테스트를 기반으로 볼 때, 할라피뇨는 하드웨어의 이론적 한계에 가까운 수준에서 우리의 가장 중요한 워크로드를 효율적으로 실행할 것"이라고 밝혔습니다.
브로드컴의 황탄 사장 겸 CEO는 "OpenAI와의 협력은 향후 10년간의 AI에 필요한 물리적 인프라를 확장하려는 근본적인 약속을 나타낸다"며, "이것은 다세대 로드맵의 시작에 불과하다. OpenAI와 직접 업계 최고의 실리콘을 공동 개발함으로써 우리는 2026년부터 마이크로소프트 및 기타 파트너들과 함께 기가와트 규모의 데이터센터 배포를 가능하게 할 것"이라고 전했습니다.
LLM을 위한 최고의 추론 플랫폼을 목표로 설계
할라피뇨는 기존 AI 워크로드에서 변형된 범용 가속기가 아닌, 현대적인 LLM 추론을 위해 백지상태에서 새롭게 설계된 결과물입니다.
June 24, 2026 Company OpenAI and Broadcom unveil LLM-optimized inference chip Loading… Share Early testing shows that the first-generation accelerator will deliver performance per watt substantially better than current state-of-the-art Built from the ground up for current and future LLMs across the industry Developed from design to production in nine months, accelerated by OpenAI’s models Expands OpenAI’s full-stack platform, from products to models and now to chips To be deployed at gigawatt scale with data center partners, over multiple generations OpenAI and Broadcom (NASDAQ: AVGO) today unveiled Jalapeño, OpenAI’s first Intelligence Processor: an accelerator architected around OpenAI’s vision for the future of LLM inference, and the first AI accelerator in a multi-generation compute platform the companies are building together to make advanced AI faster, more reliable, and more accessible to more people. Jalapeño was delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom President and CEO Hock Tan and President Charlie Kawwas, marking an important step in OpenAI’s strategy to build the full stack behind its models and products. OpenAI designed the chip from scratch around its deep understanding of LLM fundamentals, informed by its roadmap of models, kernels, serving systems, and product needs, with partners Broadcom and Celestica, helping industrialize the platform through chip implementation, board, rack system integration, high-performance networking, and scalable production systems. Jalapeño is designed with flexibility to work with all LLMs guided by OpenAI’s insights into the inference needs of current and future AI models across the industry. Engineering samples of the Jalapeño chip are running ML workloads in the lab at production target frequency and power, including GPT‑5.3‑Codex‑Spark. While OpenAI is still measuring final performance, early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art. A detailed technical report on performance will be presented in the coming months. The architecture reduces data movement and balances compute, memory, and networking resources to achieve realized utilization much closer to theoretical peak performance. Broadcom’s silicon implementation and networking technologies, including Tomahawk networking silicon, help bring the platform to large-scale production. “The world is moving to a compute-powered economy,” said Greg Brockman, President and Co-Founder of OpenAI. “Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems. By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access.” “Jalapeño was designed from the ground up for LLM inference using detailed insights from our close collaboration with OpenAI researchers,” said Richard Ho, who leads OpenAI’s hardware program. “We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.” “Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI,” said Hock Tan, President and CEO, Broadcom. “This is just the beginning of a multi-generation roadmap. By co-developing our industry-leading silicon directly with OpenAI, we are enabling the deployment of gigawatt scale data centers with Microsoft and other partners beginning in 2026.” Designed to be the best inference platform for LLMs Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads. It is informed by the systems OpenAI runs every day across ChatGPT, Codex, the API, and future agentic products, while also being designed for current and future LLMs across the industry. The goal is to combine the power and throughput of today’s leading AI accelerators with latency closer to the fastest specialized inference systems, making Jalapeño well suited for interactive LLM products at scale. That is the full-stack advantage. OpenAI is not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience. Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable for users. Jalapeño strengthens the flywheel behind OpenAI’s progress. Better infrastructure drives compute efficiency. Greater compute efficiency enables better training and serving, ultimately powering more capable AI models. Better models become better products for people, developers, and businesses. Better products drive more usage, more customers, and more revenue, which lets OpenAI reinvest in the next generation of infrastructure. Over time, that cycle helps make intelligence more capable, more reliable, and less expensive for everyone. Nine-month tape-out, accelerated by OpenAI models Jalapeño was co-developed from initial design to manufacturing tape-out in just nine months, and the custom AI accelerator program represents what we believe to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. That speed reflects deep software-hardware co-development with OpenAI’s engineering teams, Broadcom’s silicon implementation expertise, and the use of OpenAI models to accelerate parts of the design and optimization process. The same models served to users are helping improve the infrastructure used to run future models. If AI can help engineers design better chips faster, it can lower the cost of compute across the industry and help democratize access to advanced AI. Building a multi-generation platform with partners Jalapeño is the first step in a multi-generation compute platform designed for initial deployment by the end of 2026 and expanding in the years ahead, combining OpenAI-designed accelerators with Broadcom silicon implementation, networking, and connectivity technologies; and Celestica’s board, rack, and system expertise. Making advanced AI more broadly available The point of this work is simple: inference is where AI reaches people. Every improvement in cost, speed, and reliability can show up as a faster ChatGPT answer, a Codex task that can take more steps with less waiting, an API product that is cheaper to build, or more dependable access when demand is high. Democratizing AI means making advanced models available, dependable, and affordable enough for more people to use every day. Jalapeño helps OpenAI turn more of its infrastructure into useful intelligence for students, developers, small businesses, researchers, enterprises, and anyone trying to learn, create, or solve hard problems. 2026 Author OpenAI Keep reading View all Daybreak: Tools for securing every organization in the world Security Jun 22, 2026 Samsung Electronics brings ChatGPT and Codex to employees Company Jun 21, 2026 OpenAI to acquire Ona Company Jun 11, 2026