최근 LLM(대형 언어 모델)의 한계가 부각되면서, 물리적 세계를 이해하고 시뮬레이션하는 '월드 모델(World Model)'이 AI의 차세대 프론티어로 떠오르고 있습니다. 구글 딥마인드, 세계적인 석학들이 참여한 스타트업 등에서 막대한 자금을 바탕으로 관련 연구 및 상용화를 추진 중이며, 이는 향후 로봇 공학, 3D 에셋 생성, 공간 지능 등의 분야를 혁신할 중요 기술로 평가받고 있습니다.
번역된 본문
지난 몇 년간, 우리는 지금 흔히 인공지능(AI)이라고 부르는 것에 대해 집중적인 학습 과정을 겪었습니다. 하지만 실질적으로는 이것이 대부분 대형 언어 모델(LLM)에 대한 집중 학습이었습니다. 그러나 LLM이 더 이상 막대한 기대와 대규모 자금 조달, 중요한 연구 및 제품 개발을 이끄는 유일한 AI 분야가 아닙니다. 지난 1년 동안, 우리는 '월드 모델(World Models)'이라는 카테고리에서 쏟아지는 새로운 발표를 보아왔으며, 앞으로 몇 달, 몇 년 동안 이 분야에서 더 많은 움직임이 있을 것으로 보입니다. 언어를 처리하는 것을 넘어(혹은 그와 함께), 월드 모델은 물리적 세계를 시뮬레이션하거나 최소한 그에 유용하게 근사할 수 있는 AI 시스템의 기반을 다지는 것을 목표로 합니다. 이러한 아이디어가 무엇이 다르고 중요한지 검토하기 위해, 아스(Ars)는 월드 모델 및 관련 기술을 연구하는 3명의 전문가이자 실무자들과 대화를 나누었습니다. MIT의 빈센트 지츠만(Vincent Sitzmann), 런웨이(Runway)의 아나스타시스 게르마니디스(Anastasis Germanidis), 그리고 월드 랩스(World Labs)의 벤 마일든홀(Ben Mildenhall)입니다. 이 대화를 통해 우리는 흥미로운 사실을 하나 깨달았습니다. 제품으로서의 LLM이 인터페이스(채팅)에서 시작하여 그다음에 사용 사례를 찾았던 반면, 현재 월드 모델 분야의 주요 기업들은 반대 방향으로 일하고 있습니다. 이들은 로봇 공학, 연구 및 에셋 생성 등의 특정 사용 사례와 애플리케이션에서 시작하고 있지만, 최종적으로 그 인터페이스, 시스템 및 도구가 어떤 모습일지는 아직 불분명합니다.
목차
섹션으로 이동
LLM 환상에서 벗어나는 출구
곧 보시게 되겠지만, 아키텍처와 사람들이 시간이 지남에 따라 이 기술이 발전할 것이라 기대하는 방식에 있어 LLM과 월드 모델은 많은 유사점을 가지고 있습니다. 그러나 일부 사람들에게 월드 모델은 LLM의 한계에 대한 잠재적인 해결책으로 여겨지고 있으며, 이에 대한 연구는 사실 LLM이라는 현대적 담론보다 그 이전부터 존재해 왔습니다.
올해 초, 전 메타(Meta) 최고 AI 과학자인 얀 르쿤(Yann LeCun)은 와이어드(Wired)와의 인터뷰에서 이렇게 말했습니다. "LLM의 능력을 인간 수준의 지능에 도달할 때까지 확장하겠다는 아이디어는 완전한 넌센스입니다." 르쿤은 AI 및 LLM 분야에 종사하는 일부 사람들에게는 반론을 제기하는 것처럼 보일지 몰라도, 실제로는 이 분야의 상당수를 대변하는 발언을 했습니다.
최근 월드 모델을 연구하는 새로운 기업 중 하나인 월드 랩스(World Labs)를 공동 설립한 컴퓨터 비전 선구자 페이페이 리(Fei-Fei Li)의 사례를 보십시오. 그녀는 작년 말 서브스택(Substack) 게시물에서 이렇게 썼습니다.
오늘날 대형 언어 모델(LLM)과 같은 최고의 AI 기술은 우리가 추상적인 지식에 접근하고 다루는 방식을 변화시키기 시작했습니다. 그러나 여전히 그들은 어둠 속의 언어 장인에 불과합니다. 유창하지만 경험이 없고, 지식이 있지만 실체 세계에 기반하고 있지 않습니다. 공간 지능(Spatial intelligence)은 우리가 현실 및 가상 세계를 만들고 상호 작용하는 방식을 변화시켜 스토리텔링, 창의성, 로봇 공학, 과학적 발견 등을 혁신할 것입니다. 이것이 AI의 다음 프론티어입니다.
르쿤과 리의 기업들은 이러한 아이디어를 바탕으로 구축되었으므로 그들이 이런 말을 하는 것은 놀라운 일이 아닙니다. 하지만 여전히 주로 LLM을 다루고 있는 저명한 인사들로부터도 비슷한 감상을 들을 수 있습니다. 다양한 종류의 LLM 리포지토리를 호스팅하는 플랫폼인 허깅페이스(Hugging Face)의 클렘 델랑구(Clem Delangue) CEO는 "우리는 지금 LLM 거품 안에 있다고 생각하며, 내년에는 이 LLM 거품이 꺼질 수도 있다고 생각한다"고 말했습니다. 그는 컨퍼런스 연설에서 이렇게 덧붙였습니다. "하지만 생물학, 화학, 이미지, 오디오, 비디오 등에 AI를 적용하는 측면에서 볼 때 'LLM'은 AI의 한 부분집합에 불과합니다. 우리는 이제 막 그 시작점에 서 있으며, 앞으로 몇 년 동안 훨씬 더 많은 것을 보게 될 것입니다."
자금의 몰려드는热潮
지난 몇 달 사이에만, 월드 모델은 연구 주제(물론 여전히 연구 주제이기도 하지만)에서 새로운 상업적 프로젝트와 대규모 자금 조달의 기반으로 발전했습니다. 몇 가지 주요 예시는 다음과 같습니다. 8월에 구글 딥마인드(Google DeepMind)는 비디오 생성 모델을 기반으로 실시간 상호작용성을 구축한 모델인 '지니 3(Genie 3)'을 공개했습니다. 그리고 11월에는 월드 랩스가 사용자가 내보내기(export)할 수 있는 몰입형 환경을 생성할 수 있게 해주는 모델 및 툴셋인 '마블(Marble)'을 선보였습니다.
Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav Over the past few years, many of us have gotten a crash course in what we now call artificial intelligence—but really, it has mostly been a crash course in large language models. Increasingly, however, LLMs are no longer the only category of AI drawing high expectations, massive funding rounds, and significant research and product development. Over the past year, we’ve seen a plethora of new announcements in a category labeled “world models,” and you’ll likely see more movement there in the coming months and years. Instead of or in addition to working with language, world models aim to lay the groundwork for AI systems that are capable of simulating the physical world, or at least a useful approximation of it. To examine what’s different and important about this idea, Ars spoke with three expert practitioners working on world models and related technologies: Vincent Sitzmann from MIT, Anastasis Germanidis from Runway, and Ben Mildenhall from World Labs. From these conversations, we learned that while LLMs-as-a-product started with an interface (chat) and then sought a use case, the big players in world models right now are arguably working in the other direction: They’re starting with specific use cases and applications in robotics, research, and asset generation, but it’s unclear exactly how the interfaces, systems, and tools will ultimately look. Table of Contents Jump to section The off-ramp from LLM disillusionment As you’ll soon see, there are many parallels between LLMs and world models in terms of architecture and how people expect them to improve over time. For some, though, they’re seen as a potential answer to the limitations of LLMs, even though work on them predates that contemporary narrative. “The idea that you’re going to extend the capabilities of LLMs to the point that they’re going to have human-level intelligence is complete nonsense,” former Meta chief AI scientist Yann LeCun told Wired earlier this year. LeCun has made waves with an opinion that some working in AI and LLMs see as contrarian, but he’s actually speaking for a sizable segment of the field. See also Fei-Fei Li, the computer vision pioneer who co-founded World Labs, one of the new companies working on world models. In a Substack post late last year, she wrote : Today, leading AI technology such as large language models (LLMs) have begun to transform how we access and work with abstract knowledge. Yet they remain wordsmiths in the dark; eloquent but inexperienced, knowledgeable but ungrounded. Spatial intelligence will transform how we create and interact with real and virtual worlds—revolutionizing storytelling, creativity, robotics, scientific discovery, and beyond. This is AI’s next frontier. LeCun and Li’s ventures are built on these ideas, so it’s not surprising they’d say these things. But you’ll also see similar sentiments from some prominent figures still working primarily with LLMs. “I think we’re in an LLM bubble, and I think the LLM bubble might be bursting next year,” said Clem Delangue , the CEO of Hugging Face—a platform that hosts repositories of LLMs of all stripes. “But ‘LLM’ is just a subset of AI when it comes to applying AI to biology, chemistry, image, audio, [and] video,” Delangue added, speaking at a conference. “I think we’re at the beginning of it, and we’ll see much more in the next few years.” A financial flurry Over just the past few months, world models have advanced from a research topic (which they still are, of course) to the basis for new commercial projects and huge funding rounds. A few key examples: In August, Google DeepMind unveiled Genie 3, a model that builds real-time interactivity on top of the foundation of a video generation model. In November, World Labs introduced Marble , a model and toolset that allows users to generate immersive environments that can be exported as 3D assets, based on input in the form of text, images, video, or other assets. Just a month later, video generation and filmmaking AI company Runway entered the fray with its announcement of GWM-1 , a trio of specialized world models built on Runway’s past work with video models. Even more recently, Yann LeCun started Advanced Machine Intelligence (AMI) , a company betting the farm on the notion that the real future of AI systems is in models that interact with the physical world (or at least simulate it), not just language. These and similar efforts have received substantial funding. World Labs and AMI reportedly raised around $1 billion each in February and March , respectively, and Runway also raised $315 million in February. Some of the activity around world models is at least in part aimed at ostensibly building the foundations of AGI or superintelligence, but most people working on them are talking about practical applications: training, testing, and driving robots; generating 3D assets for game development and film production; scientific simulation and modeling; and so on. It’s important to note that “world models” is an umbrella term that is often thrown around without a clear definition, though. Defining “world model” “It’s definitely an overloaded term,” Vincent Sitzmann told me in a lengthy conversation about the research and concepts underlying world models. Sitzmann is an assistant professor at MIT who has published research on neural rendering, visual computing, and robotics. He leads the Scene Representation Group within MIT’s Computer Science and Artificial Intelligence Laboratory ( CSAIL ). When asked to give a definition, he simply described a world model as any model that takes in an interaction, “and given that interaction, it enables you to simulate what would happen next in some environment.” When announcing its GWM-1 family of models in December, Runway defined a world model as “an AI system that builds an internal representation of an environment and uses it to simulate future events within that environment.” Further, “the aim of general world models is to represent and simulate a wide range of situations and interactions—like those encountered in the real world,” Runway added. I also spoke with Ben Mildenhall, co-founder of World Labs, a former Google computer vision and physics researcher and co-creator of neural radiance fields ( NeRF )—a method for constructing navigable 3D scenes from 2D images in a format that is differentiable and therefore useful in machine learning contexts. “The key things that would distinguish it from an LLM are demonstrating degrees of spatial and—maybe for lack of a better word—continuous understanding,” he said. “A very distinguishing aspect of interacting with an LLM is they are turn-based.” Meaning, users type some text, there’s a pause, and then they get a block of text back. By contrast, he sees a world model as a synchronous, real-time system. “Something that would define a world model is the degree of freedom that you have in interacting with the spatial world, where you do not have this mediated linear journey of A then B then A then B then A then B,” he said. When a user or agent is utilizing a world model, they are “actually able to interact with it like it is some sort of world and you are taking continuous actions” where “there are parallel things happening at the same time.” Mildenhall’s co-founder Fei-Fei Li has written that she believes there are three criteria that define a world model: World models “can generate worlds with perceptual, geometrical, and physical consistency,” “are multimodal by design,” and “can output the next states based on input actions.” The phrase itself is not new; it has long appeared in reinforcement learning and robotics to describe models that predict environment dynamics. What is new is the attempt to scale that idea into general-purpose, generative systems trained on massive visual and multimodal data. The huge funding r