메뉴
BL
Wired AI • 58일 전

링크드인, 올해 AI 데이터센터 추가 확장 보류

IMP
7/10
핵심 요약

링크드인은 최근 6개월간 기존 GPU의 효율성을 두 배로 끌어올린 성과를 바탕으로, 올해 데이터센터 및 GPU 투자를 동결하기로 결정했습니다. 이는 막대한 자본을 쏟아부으며 AI 인프라 확장에 경쟁하는 다른 빅테크 기업들과 대비되는 신중한 접근법으로, AI 투자가 단순한 실험 단계를 넘어 비용 효율을 따지는 '생산 관리' 단계로 진입했음을 시사합니다.

번역된 본문

다른 많은 빅테크 플랫폼들과 달리, 링크드인(LinkedIn)은 올해 회계연도에 AI 데이터센터 확장에 공격적인 지출을 하지 않기로 결정했습니다. 이 직장인 소셜 네트워크의 경영진은 WIRED에 GPU(Graphics Processing Units)에 대한 투자를 현행 수준으로 유지하며, 컴퓨팅 및 스토리지 규모 또한 동결할 계획이라고 밝혔습니다. 이러한 지출 계획은 지난달에 시작되어 내년 6월에 끝나는 링크드인의 회계연도에 적용됩니다. 이 회사는 지난 6개월 동안 기존 GPU를 두 배 더 효율적으로 사용할 수 있는 방법을 찾았기 때문에 AI 하드웨어에 막대한 비용을 지출하는 것을 피할 수 있었다고 말합니다. AI의 하드웨어 수요가 빠르게 변화하고 있어 링크드인의 계획이 차질을 빚을 가능성도 여전히 있지만, 경영진은 회사가 이미 메모리 칩 가격 급등을 고려했다고 밝혔습니다.

링크드인의 엔지니어링 최고 기술 책임자(CTO)인 에런 버거(Erran Berger)는 "더 많은 컴퓨팅 파워를 요구하는 기능들을 제품화하면서도, 컴퓨팅 인프라 규모를 기존 수준에 머물거나 최대한 가깝게 유지하는 것을 목표로 삼았습니다. 오늘날의 상황에서 이는 매우 대담한 선언입니다."라고 말했습니다. 버거와 인프라 담당 최고 기술 책임자인 라구 하이레마갈루르(Raghu Hiremagalur)는 지출에 신중을 기해야 하며, 이러한 새로운 제약 조건이 엔지니어링 팀이 링크드인이 출시할 예정인 수많은 새로운 생성형 AI 기능을 개발할 때 더 창의적으로 만들 것이라고 밝혔습니다. 버거는 이러한 효율성 개선이 시간이 지남에 따라 복합적으로 쌓여, 결국 예산을 다시 늘릴 때 데이터센터 확장의 효과를 극대화할 수 있을 것이라고 믿는다고 덧붙였습니다.

하이레마갈루르는 "우리와 같은 규모의 회사가 1년 내내 스토리지나 컴퓨팅을 추가 증설 없이 이를 해내겠다고 말하는 것은 결코 쉬운 일이 아님을 강조하고 싶습니다. 이를 위해 엄청난 노력이 필요했습니다."라고 말했습니다. OpenAI, 메타(Meta), 구글(Google)과 같은 기업들은 최신 컴퓨터 칩으로 가득 찬 대규모 데이터센터를 건설, 설비 갖추기, 운영하기 위해 동원할 수 있는 모든 자금을 긁어모으고 예상치 못한 파트너십을 맺고 있습니다. 인력 및 부품 부족으로 인해 많은 프로젝트가 지연되었고, 수많은 기업이 일부 AI 도구의 고객 사용량을 제한해야만 했습니다. 하지만 이처럼 맹목적인 AI 투자가 지속 가능한지에 대한 의문도 커지고 있습니다. 13억 명 이상의 사용자를 보유한 링크드인은 건설 붐에 역행하며 공개적으로 지출 우려를 제기한 것으로 보이는 가장 큰 기업일 것입니다.

서버 제조업체인 HP의 이사회 멤버이자 프린시펄 벤처 파트너스(Principal Venture Partners)의 매니징 파트너인 윤송이(Songyee Yoon)는 "이는 업계에 고무적인 일입니다."라며, "이는 AI가 단순한 실험 단계에서 생산 관리로 넘어가기 시작했음을 시사합니다. 승리하는 기업은 단순히 인프라에 가장 많은 돈을 지출하는 곳이 아닐 것입니다."라고 말했습니다.

소유와 운영의 변화 2016년 마이크로소프트(Microsoft)가 링크드인을 인수한 지 몇 년 후, 이 회사는 모회사의 Azure 클라우드 서비스로 이전을 시도했지만 거대한 소셜 네트워크를 범용 데이터센터에 구겨 넣는 것은 경제성이 맞지 않았습니다. 하이레마갈루르는 "마이크로소프트 Azure는 미친 듯이 성장했고 고객의 수요 수준은 하늘을 찌르고 있었으며, 동시에 링크드인 측면에서도 급증하는 성장세를 보였습니다."라고 말했습니다. 2022년 링크드인은 오리건주, 텍사스주, 버지니아주에 위치한 자체 데이터센터에 전적으로 의존하기로 했습니다. 이러한 인프라의 자체 소유는 링크드인이 기술의 모든 세부 사항을 상당 부분 통제할 수 있게 해주었으며, 새로운 시대의 현실에 대응할 수 있는 발판을 마련했습니다. 거의 같은 시기에 링크드인은 사용자가 메시지를 작성하고, 일자리를 찾고, 후보자를 채용하는 것을 도울 수 있는 AI 기반 비서 개발을 시작했습니다. 이러한 노력은 비용이 많이 들었습니다.

하이레마갈루르는 "사이트로 들어오는 모든 쿼리의 비용이 시간이 지나면서 증가했습니다."라고 말하며, 링크드인이 저장하는 데이터의 양은 매년 두 배씩 증가하고 있다고 덧붙였습니다. "이는 지속 가능한 상황이 아닙니다." 링크드인은 모델을 학습시키는 단계부터 사용자 쿼리에 응답하기 위해 이를 서비스하는 단계까지, AI 파이프라인의 모든 단계에서 데이터센터 사용량을 최적화하는 방향으로 전환했습니다. 하이레마갈루르의 팀은 이를 위한 기술을 개발했습니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Unlike many other big tech platforms, LinkedIn has decided it won’t spend aggressively on expanding its AI data centers this fiscal year. Executives at the professional social network tell WIRED that it plans to keep its investment in GPUs steady, and its compute and storage footprint is also remaining flat. The spending calculations apply to LinkedIn’s fiscal year that began last month and ends next June. The company says it was able to avoid spending big on AI hardware because it found ways to use its existing GPUs twice as efficiently over the past six months. LinkedIn’s plan could still unravel because the hardware demands of AI are shifting rapidly, but executives say the company has already taken into account surging prices for memory chips . “One of the goals we've set is to try to basically keep our compute footprint flat or as close to flat as possible while shipping more compute-hungry things to production,” says Erran Berger, LinkedIn’s chief technology officer for engineering. “That’s a pretty bold statement to make in today's world.” Berger and Raghu Hiremagalur, LinkedIn’s chief technology officer for infrastructure, say they want to be prudent about spending and that the new constraints will motivate engineering teams to get more creative when developing the many new generative AI features LinkedIn is planning to launch. Berger says he believes the efficiency gains could compound over time, enabling LinkedIn to get more out of data center expansions when it eventually increases its budgets again. “I really want to double underscore that for a company of our scale, to say a full year we're going to do this with no incremental storage and compute is no small feat, but it's taken a ton of work to get there,” Hiremagalur says. Companies such as OpenAI, Meta, and Google are scrounging up all the money they can find and coupling up in unexpected partnerships to construct, furnish, and operate massive data centers filled with the newest computer chips. Labor and parts shortages have held up many projects, and many businesses have had to limit customer usage of some AI tools. But there are also growing questions about whether the relentless investment in AI is sustainable. LinkedIn, with more than 1.3 billion users, is perhaps the largest business yet to publicly address spending concerns by bucking the building boom. “It is encouraging for the industry,” says Songyee Yoon, managing partner of Principal Venture Partners and a board member at the server maker HP. “It suggests AI is beginning to move from experimentation into production discipline. The companies that win will not simply be the ones that spend the most on infrastructure.” Owning It A few years after Microsoft acquired LinkedIn in 2016, the company tried moving to its parent company’s Azure cloud service, but it didn’t make economic sense to squeeze the giant social network into general-purpose data centers. “Microsoft Azure was growing like crazy, the level of customer demand was through the roof, and at the same time we saw skyrocketing growth on the LinkedIn side,” Hiremagalur says. In 2022, LinkedIn went all-in on its own data centers in Oregon, Texas, and Virginia. The ownership gave LinkedIn significant control over every detail of its technology, setting itself up well to meet the realities of a new era. Around the same time, LinkedIn began developing AI-based assistants that could help users write messages, find jobs, and recruit candidates. The endeavor wasn’t cheap. “Every query that's coming to our site has increased in cost over time,” Hiremagalur says, adding that the amount of data LinkedIn stored was doubling annually. “That is not a sustainable place to be.” LinkedIn moved to optimize its data center usage at every stage of the AI pipeline, from training models to serving them in response to user queries. Hiremagalur’s team developed measurement tools to understand how much compute and storage individual teams were using. It then set up a system to allocate projects to the computers in its data centers more efficiently so that they sat idle less often. “Our allocation efficiency and utilization of GPUs on the training side is the best that I have seen,” Hiremagalur says, describing usage at north of 95 percent. LinkedIn also used techniques such as distillation to train smaller AI models from larger ones. For its job recommendation tools, a single model learned from two larger models to both identify relevant openings and predict which users were most likely to click on them. While the smaller model is more affordable to operate, Berger says it doesn't sacrifice on quality. “People are finding and discovering jobs that they were not successfully finding before, because the model is doing a really good job of understanding” their desires, he says. The model used to select which posts to show users on their newsfeeds also was expensive to operate initially, but LinkedIn found a way to get it to run in “a reasonably cost-effective way,” Berger says. He says that the company made dozens of improvements, such as streamlining model training, reusing information from earlier recommendations, and better balancing workloads between CPUs and GPUs. LinkedIn even reworked some of the foundational software on Nvidia processors to make them capable of handling tasks larger than they were designed for. It also rejiggered other software to run tasks on CPUs instead of Nvidia GPUs, which are pricier, more difficult to procure, and consume greater amounts of electricity. Altogether, LinkedIn estimates its efficiency work has saved about $24 million over the past 12 months, or the equivalent of roughly 1,100 GPUs running around the clock for a year. The LinkedIn executives acknowledge that the savings don’t amount to much for a company that has $18 billion in annual sales. But Hiremagalur says “craft” matters too, and so does “the agility” created by freeing up resources. Engineers can take on their next project sooner and integrate more AI capabilities without having to increase LinkedIn’s computing footprint. Berger contends that LinkedIn has been able to deliver better job and candidate results while still keeping a lid on costs and generating a financial return for Microsoft. “We should be able to deliver better quality by deploying larger models, doing deeper inference, and doing it for cheaper if we can,” he says. Despite keeping spending flat, LinkedIn’s data centers aren’t going stale. The company has committed to purchasing new servers that will enable it to keep upgrading its machines as they age out or break down over the coming months. But the company was able to lock in some savings by buying ahead. “The cost of all of this hardware has just gone through the roof,” Hiremagalur says, noting some servers have jumped three-fold in price in the past few months. “It's just nuts.” LinkedIn’s efficiency drive is part of a broader trend sometimes referred to as “tokenomics”—more deeply analyzing the cost of using generative AI tools. “Enterprises are evolving from a buy-more era to a do-more era,” says Chirag Dekate, who helps businesses think through their AI cloud strategies for the consultancy Gartner. “Until now, the mantra was, buy more to save more. But buy more only increases costs.” Dekate says he has recently seen smaller businesses that don’t have as much control over their infrastructure as LinkedIn find their own ways to shed costs. They have purged unused software and shed employees, bought data center space from so-called “neoclouds” that can be more affordable than traditional providers, and adopted the lowest-cost AI models possible for any given project. He worries somewhat that LinkedIn and others could run into a wall if they introduce too many financial constraints, because the need for greater amounts of compute and storage are inevitable.