메뉴
BL
TechCrunch AI • 27일 전

엔비디아의 AI 우위, GPU를 넘어서다

IMP
7/10
핵심 요약

엔비디아의 경쟁 우위가 GPU 자체를 넘어 데이터 오케스트레이션과 시스템 전반으로 확장되고 있습니다. 기가와트급 AI 인프라에서 베라 루빈(Vera Rubin) 아키텍처처럼 GPU 주변의 CPU·스토리지·네트워킹을 효율화하는 전문 하드웨어가 새로운 경쟁 무대가 되고 있으며, 이는 데이터 이동 최소화에 집중한 OpenAI의 할라페뇨 칩과 같은 접근과 같은 맥락입니다.

번역된 본문

이번 주까지 엔비디아에 대한 지배적 서사는 대략 이랬다: AI 붐 초기 몇 년 동안 엔비디아는 최첨단 GPU의 유일한 공급원이었고, 업계가 확장되면서 막대한 수익을 올렸다. 최근 몇 년간 아마존과 구글 같은 하이퍼스케일러들이 자체 칩을 개발하기 시작하면서 엔비디아는 더 이상 유일한 선택지가 아니게 됐고, 많은 투자자들은 그 우위가 과연 얼마나 지속될지 의문을 갖게 됐다. 이는 설득력 있는 이야기이고 대체로 사실이다. 2023년 초부터 2025년 중반까지 시가총액이 10배 성장한 엔비디아 주가는 GPU 경쟁에 대한 우려로 지난 1년간 완만한 흐름을 보였다.

수요일 발표된 실적 이후 새로운 서사가 형성되고 있으며, 투자자들은 엔비디아의 우위가 GPU를 훨씬 넘어선다는 것을 깨닫기 시작했다. AI 컴퓨팅이 기가와트 규모로 커지면서 오케스트레이션(총괄 관리)은 점점 더 복잡한 과제가 되었다. 놀랍지 않게도 엔비디아는 이를 처리하는 데 필요한 최첨단 하드웨어의 상당 부분을 구축했으며, GPU 자체에 대한 경쟁이 심해지는 가운데서도 GPU를 둘러싼 시스템에서 큰 우위를 점하고 있다. 컴퓨팅이 커모디티(일반 상품)화되었다는 이야기가 많지만, 메가스케일 데이터센터를 최대 효율로 운영하는 것은 여전히 지극히 어려운 일이며, 배포 규모가 커지고 빨라질수록 이 과제는 더욱 커질 뿐이다.

래크 단위로 보기 엔비디아가 실제로 판매하는 것의 세부 사항만 봐도 이를 알 수 있다. 회사는 현재 루빈(Rubin) GPU를 베라(Vera) CPU, Groq 3 LPX 추론 가속기, 그리고 스토리지와 네트워킹을 위한 유사한 랙 등 다른 유닛들과 결합하는 베라 루빈(Vera Rubin) 아키텍처를 출시하고 있다. 지난 주 나는 이 시스템들이 실제로 무엇을 하는지 엔비디아 관계자들과 이야기를 나눴는데, 그 결과는 놀라웠다. 루빈 GPU 자체와 마찬가지로 이들은 극도로 특화된 시스템이지만, 토큰을 처리하는 대신 GPU 외부의 모든 것이 가능한 한 효율적으로 작동하도록 보장하는 역할을 한다. GPU가 엔진이라면 이들은 자동차의 나머지 부분인 셈이다.

특히 베라 CPU는 데이터 오케스트레이션 문제에 집중한다. "베라가 중요한 이유는 단일 서버나 어떤 컴퓨팅 플랫폼에도 넣을 수 있는 메모리 용량에는 한계가 있기 때문입니다"라고 엔비디아의 스토리지 기술 부사장 제이슨 하디(Jason Hardy)가 말했다. 데이터센터가 컴퓨팅 파워를 확장하면서 메모리 용량도 함께 확장됐고, 이것이 마이크론 같은 회사들이 인프라 붐의 2차 수혜를 입은 이유다. 하지만 그 데이터를 적시에 GPU에 전달하는 것은 결코 간단하지 않으며, 기업들이 와트당 토큰 수를 점점 더 낮추려 하면서 이런 트래픽 제어의 중요성을 실감하고 있다.

"베라 CPU가 가속을 허용하는 이런 작업에서 최대 3배의 성능 향상을 확인했습니다"라고 하디는 말했다. "이제 병목 없이 모든 성능을 끌어낼 수 있어 플래시 스토리지를 최대한 활용할 수 있습니다."

엔비디아 외부에서도 같은 문제의 다른 버전을 볼 수 있다. OpenAI가 할라페뇨(Jalapeño) 칩을 개발할 때 주요 초점은 이동해야 하는 데이터량을 최소화함으로써 이런 문제를 완전히 피하는 것이었다. "우리는 데이터 이동과 통신 지연을 최소화하도록 할라페뇨를 설계했습니다"라고 회사는 이번 달 블로그 포스트에서 밝혔다. "넓은 도메인 덕분에 전체 워크로드가 하나의 연결된 시스템 내에 머물러 데이터 이동이 최소화되고, 요청 전체가 처음부터 끝까지 빠르고 효율적으로 처리될 수 있습니다." 이는 워크로드를 하나의 통합 칩 안에서 수행함으로써 데이터 이동을 아예 피하는 다른 접근이지만, 전체 논리는 같다. 단순히 더 많은 프로세서 사이클이 아니라 더 스마트한 트래픽 제어로 효율을 높이는 것이다. 이는 결국 기업들이 경쟁할 완전히 새로운 인프라 계층을 열어준다.

이런 데이터 오케스트레이션에 대한 새로운 초점이 엔비디아에게 자동으로 승리를 안겨주는 것은 아니다. 회사는 GPU에서 그래왔던 것처럼 경쟁 칩 제조사들 및 하이퍼스케일러들과 경쟁해야 할 것이다.

원문 보기
원문 보기 (영어)
Before this week, the dominant story about Nvidia went something like this: For the first few years of the AI boom, Nvidia was the only source for state-of-the-art GPUs, which became immensely profitable as the industry scaled out. In the last few years, hyperscalers like Amazon and Google have started building their own chips, and Nvidia is no longer the only game in town, leading many investors to wonder how durable its advantage really is. It’s a compelling story, and mostly true. After growing its market cap 10x between the start of 2023 and mid-2025, Nvidia shares have been on a more modest trajectory for the past year, driven by concerns about GPU competition. A new narrative has taken shape since the company’s earnings on Wednesday and investors are starting to realize that Nvidia’s advantage goes far beyond GPUs. As AI’s compute grows into the gigawatt scale, orchestration has become an increasingly complex task. Not surprisingly, Nvidia has built much of the state-of-the-art hardware needed to handle it, giving the company a huge advantage in the systems that surround the GPU even as it sees increased competition on the GPUs themselves. For all the talk of compute as a commodity , it’s still incredibly difficult to operate a megascale data center at peak efficiency — and as deployments get bigger and faster, that challenge is only growing. Rack by Rack You can see some of this just by looking at the details of what Nvidia is actually selling. The company is currently rolling out its Vera Rubin architecture, which pairs the Rubin GPU with a collection of other units, including the Vera CPU, the Groq 3 LPX inference accelerator and similar racks for storage and networking. Over the past week, I’ve been talking to folks at Nvidia about what those systems actually do, and the results have been surprising. Like the Rubin GPU itself, they’re extremely specialized systems, but instead of churning through tokens, they’re making sure everything outside the GPU works as efficiently as possible. If the GPU is the engine, these are the rest of the car. The Vera CPU in particular is focused on the problem of orchestrating data. “Vera is important because there's only so much memory that you can put in a single server or any sort of compute platform,” Jason Hardy, Nvidia’s VP of storage technology, told me. As data centers have scaled up computing power, memory capacity has scaled up too, which is why companies like Micron have gotten rich in the second wave of the infrastructure boom. But getting that data to the GPU at the right time isn’t straightforward — and as companies look to drive tokens-per-watt lower and lower, they’re realizing how important that kind of traffic direction is. “We saw upwards of 3x improvement in these operations, where the Vera CPU is allowing for acceleration,” Hardy said. “So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking.” You can see versions of the same problem outside of Nvidia. When OpenAI developed its Jalapeño chip, a major focus was avoiding these challenges entirely by minimizing the amount of data that needs to be moved around. “We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post earlier this month . “Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end.” It’s a different approach, avoiding data movement entirely by conducting a workload within one integrated chip. But the overall logic is the same, increasing efficiency with smarter traffic control instead of just more processor cycles. That in turn opens up a whole new layer of infrastructure for companies to compete over. This new focus on data orchestration isn’t automatically a win for Nvidia. The company will have to compete with rival chipmakers and hyperscalers just as it has with GPUs. But the competition has moved to a new layer, where building a rival GPU matters less than being able to make the entire system work efficiently. And at least in the early stages, Nvidia looks to have a commanding lead. Topics AI , GPU , nvidia , TC When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Don't miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era? REGISTER NOW Most Popular Hugging Face is selling a cute $399 open source duck robot, Microduck Rebecca Bellan Nvidia closes in on Hugging Face acquisition Connie Loizos Fitbit founders launch Luffu Link, an LTE health and safety band Aisha Malik Hugging Face reportedly in talks to be acquired for $13B Rebecca Bellan Who's behind the new ‘stealth model’ Ox Alpha? Anthony Ha Flock CEO calls for ‘compromise’ as surveillance company faces growing backlash Anthony Ha Two years after launch, Walmart's Flipkart is closing in on India's quick-commerce leaders Jagmeet Singh