메뉴
HN
Hacker News 50일 전

xAI, 최첨단 연구소보다 데이터센터 기업에 가까워지는 이유

IMP
8/10
핵심 요약

최근 전 세계적인 AI 컴퓨팅 인프라 부족 현상 속에서, xAI가 Anthropic과 Google에 대규모 데이터센터 용량을 임대하는 거래를 성사시켰습니다. 스페이스X와의 합병을 앞둔 xAI는 이를 통해 막대한 수익을 창출하며 사실상 데이터센터 부동산 임대 기업과 같은 형태로 변모하고 있습니다. 이는 심각한 서버 부족을 겪던 Anthropic과 같은 AI 기업들에게는 단비와 같으며, 일론 머스크의 압도적인 인프라 구축 속도가 시장에서 강력한 무기로 작용하고 있음을 보여줍니다.

번역된 본문

지난 몇 주 동안의 가장 놀라운 변화는 xAI가 Anthropic과 Google과 새로운 파트너십을 맺고 이들에게 막대한 컴퓨팅 용량을 제공하기 시작했다는 점입니다. 지난 2월 두 회사가 합병하면서 xAI는 이제 SpaceX의 일부가 되었으므로, 이러한 계약으로 발생하는 수익은 곧 상장(IPO)을 앞둔 법인으로 직접 흘러들어가게 됩니다. SpaceX의 다가오는 상장을 고려할 때 재무 엔지니어링의 가능성이 많이 제기되었지만, 저는 이것이 단순한 회계 기술 그 이상이라고 생각합니다.

Anthropic의 심각한 위기 Claude 제품을 자주 사용해 보셨다면 아시겠지만, Anthropic은 심각한 용량 문제를 겪고 있었습니다. 특히 유럽 시간으로 이른 오후 이후와 미국 시간으로 오전에 사용자들의 요구가 몰리면서 시스템에 큰 부담이 있었습니다. 이는 유럽과 아메리카 대륙의 사용자들이 동시에 업무를 보며 컴퓨팅 자원을 두고 경쟁하는 시간대입니다. 저는 이러한 컴퓨팅 위기에 대해 여러 번 글을 쓴 적이 있습니다. 이 문제로 인해 Anthropic은 구독 서비스에 새로운 피크 시간대 제한을 도입할 수밖에 없었습니다. 사용량이 많은 시간대에 사용자의 이용 한도를 더 많이 차감하여, 비교적 여유가 있는 한가한 시간대로 수요를 분산시키려는 목적이었습니다. 하지만 Anthropic의 수요가 폭발적으로 증가하는 상황에서는 수요를 분산하는 데에도 한계가 있습니다. 결국 어느 시점부터는 사용자에게 더 엄격한 제한을 가할 수밖에 없는데, Google과 OpenAI가 바로 뒤에서 고객을 노리고 있는 상황에서 이는 결코 바람직하지 않습니다.

xAI의 구출 투수 등판? 5월 초, xAI는 Anthropic과의 파트너십을 발표하며 멤피스에 있는 자사의 (구형) Colossus 1 데이터센터에 대한 액세스 권한을 제공했습니다. 덕분에 Anthropic은 구독 서비스의 사용 제한 조치를 철회할 수 있었고, 서비스의 전반적인 안정성은 여전히 개선의 여지가 있지만 최소한 현재로서는 피크 시간대의 병목 현상은 완화되었습니다. 이 계약에 따른 비용은 어마어마합니다. 300MW(약 22만 개의 GPU)의 용량에 대해 월 12억 5,000만 달러를 지불합니다. 지난주에는 Google이 유사한 파트너십을 발표했는데, 11만 개의 GPU에 대해 월 9억 2,000만 달러를 지불하기로 했습니다. 두 계약 모두 초기 고정 기간 이후 90일 전 통보 시 해지할 수 있는 조항이 포함되어 있다는 점을 참고해야 합니다. 표면적으로만 보더라도 이는 xAI에게 터무니없이 수익성이 좋은 거래입니다. 운영 비용(Opex)과 감가상각은 포함되지 않았지만, 이 계약이 18개월 동안 유지된다고 가정하면 xAI는 투자한 자본적 지출(Capex)을 전액 회수하고도 수십만 메가와트의 GPU를 추가로 보유하게 됩니다. 거대한 컴퓨팅 부족 현상이 중기적으로 계속될 가능성이 높은 만큼, 구형인 H100 GPU조차도 18개월 후에도 여전히 매우 유용할 것입니다.

우려되는 점들 이 거래에는 분명 몇 가지 레드 플래그(위험 신호)가 존재합니다. 첫째, 일론 먀스크와 OpenAI는 현재 치열한 법적 분쟁 중에 있으며, 이 Anthropic과의 계약이 상업적 현실보다는 OpenAI에 압박을 가하기 위한 목적일 수 있습니다. 또한 Google은 SpaceX의 주요 주주이므로 상장 시 기업 가치를 높일 충분한 동기가 있습니다. 이러한 관점들에 상당 부분(잠재적으로는 많이!) 사실이 있을 것이라고 확신하지만, 수많은 GPU가 극심하게 부족하다는 점을 간과해서는 안 됩니다. 이 데이터센터 자본 지출 호황의 숨겨진 이야기 중 하나는 모든 데이터센터 건설이 심각하게 지연되고 있다는 것입니다. 건축 규제에 대해 자유로운 것으로 유명한 아랍에미리트(UAE)에 건설 중인 OpenAI의 핵심 데이터센터 '스타게이트 UAE'조차 현재 이란과의 분쟁으로 인해 직접적인 위협을 받고 있으며, 실제로 이란 드론이 UAE의 다른 데이터센터를 타격한 바 있습니다. 이와 비교할 때, SpaceX와 xAI는 데이터센터를 제때 건설하는 데 있어 놀라운 능력을 발휘합니다. 원래의 Colossus 1 데이터센터는 단 122일 만에 건설되었습니다. 머스크의 제국은 거대한 인프라 프로젝트를 신속하게 계획, 건설 및 실행하는 방법을 진정으로 이해하고 있으며 이는 엄청난 이점입니다. 대형 클라우드 기업들도 분명 이를 수행할 경험과 기술력이 있지만, 훨씬 덜 긴급하게 구축되었으며 일반적으로 프로젝트를 완료하는 데 수년이 걸립니다.

원문 보기
원문 보기 (영어)
An unexpected development over the past few weeks is xAI's new partnerships with Anthropic and Google, providing them with a huge amount of capacity. It's worth remembering that xAI is now part of SpaceX, after the two merged back in February - so the revenue from these deals flows straight into the entity about to go public. While much has been made of the potential financial engineering given SpaceX's upcoming IPO, I think there's a bit more to this than just pure accounting tricks. Anthropic was in a serious bind If you use Claude products much, you'll be (very, probably) aware that Anthropic has had serious capacity problems, especially early afternoon onwards in Europe and in the mornings in the US (this is when demand seems to be highest as both European users and the Americas are both at work, fighting for capacity). I've written about this compute crunch before a few times - the coming crunch , whether it's here yet , and what comes next . This resulted in Anthropic having to introduce new peak hour restrictions on their subscriptions, with usage between 5am–11am PT / 1pm–7pm GMT using more of your usage limit - with the aim of smoothing demand between peak hours and off peak hours where they had more capacity available. However, there is only so much demand shifting you can do when demand is growing as fast as Anthropic's. At some point you end up having to ration users further, which definitely is far from ideal when you have both Google and OpenAI breathing down your neck for customers. xAI to the rescue? At the start of May, xAI announced a partnership with Anthropic to provide access to their (older) Colossus 1 datacentre in Memphis. This allowed Anthropic to reverse the usage limit restrictions on their subscriptions, and in general while stability of Anthropic services still leaves a lot to be desired, the peak time crunch has abated (for now, at least). The fees involved are enormous, ramping to $1.25bn/month for 300MW of capacity - approximately 220k GPUs. Last week, Google announced a similar partnership - $920mn/month for 110k GPUs [1] . It's important to note that both agreements have cancellation clauses - allowing either party to cancel with 90 days' notice after an initial lock-in period. If you take this on face value, this is a ludicrously profitable deal for xAI: While this doesn't include opex [2] and depreciation, if the deals continue for 18 months, xAI recoups all the capex they spent and still has many hundreds of MW of GPUs available. With the giant compute shortages likely to persist into the medium term, even older H100s are likely to be extremely useful even 18 months out. The case against It's important to note there are certainly some red flags with the deal. Firstly, Elon Musk and OpenAI were/are locked in a bitter legal battle, and the Anthropic deal could be motivated to add pressure to OpenAI more than commercial reality. And Google is a major shareholder in SpaceX, so they certainly have incentive to juice the valuation of the IPO. While I'm sure there is some degree (potentially a lot!) of truth in these viewpoints, it's important to note that huge volumes of GPUs are in enormously short supply. One of the untold stories of this capex boom in datacentres is just how behind all of them are. Even OpenAI's flagship Stargate UAE datacentre - being built in a jurisdiction that is renowned for a laissez-faire attitude to building regulations - is now under direct threat from the current Iran conflict, with Iranian drones having already hit other UAE datacentres . In comparison, SpaceX/xAI are incredible at building datacentres on time. The original Colossus 1 datacentre was built in 122 days. Musk's empire does have a huge advantage in really understanding how to plan, build and execute enormous infrastructure projects quickly. While the hyperscalers no doubt have the experience to do this, they were built with far less urgency - with typical project execution taking many years. Given the capex only really started to ramp up in the last couple of years, many of these projects are still years away. This gives xAI a serious competitive advantage that shouldn't in my opinion just be hand waved away. But what about Grok? There is no doubt this leaves Grok in an odd spot, with a lot of the datacentre capacity that was destined for Grok training and inference now being leased to a direct competitor. While it's foolish to write off any model provider, it certainly looks like a serious retreat from Grok vying to be a frontier class lab. But, perhaps, they over-specified their datacentre capacity - there is no doubt that inference demand for Grok models is likely to be seriously behind projections, leaving a bunch of spare capacity which might as well be monetised while the training lottery continues? It's hard to say and the xAI & Cursor deal muddies the water even further. As such, I think all three things are true to some degree. There's no doubt some level of financial engineering going on. There's also an enormous compute shortage. And it seems to me SpaceX/xAI does have a real competitive advantage in datacentre buildout. It's just the magnitude of how true each of these are is going to define the success or failure of the biggest IPO in North American history. Either way, the more I look at it, the more xAI is starting to resemble a datacentre REIT with a frontier lab attached, rather than the other way around. I suspect that these are likely to be GB200s given the pricing, vs the mostly H100/H200 for Anthropic, but this is speculation on my part. ↩︎ Power is the obvious big opex line, but at this scale it's almost a rounding error. 300MW running flat out is roughly 300,000 kW × 8,760 hours, or about 2.6 billion kWh a year. Tennessee has some of the cheapest industrial electricity in the US at around 6 cents/kWh , so buying it off the grid would cost somewhere around $160mn a year. Colossus actually runs largely on its own on-site gas turbines, which comes out even cheaper: at a simple-cycle heat rate of ~10,000 Btu/kWh and Henry Hub gas at ~$3.50/MMBtu , the fuel bill is only around $90mn a year. Either way, set against the ~$15bn a year Anthropic is paying for that 300MW, power is no more than about 1% of revenue. The deal value utterly dwarfs the running costs. ↩︎ If you found this useful, I send a newsletter every month with all my posts. No spam and no ads. Subscribe