메뉴
HN
Hacker News • 24일 전

월드랩스, 텍스트·이미지·비디오·3D 통합 월드 모델 '아틀라스' 공개

IMP
8/10
핵심 요약

World Labs가 텍스트, 이미지, 비디오, 3D를 하나의 공간적 컨텍스트로 처리하는 차세대 범용 월드 모델 '아틀라스(Atlas)'를 발표했습니다. 아틀라스는 멀티모달 자기회귀 확산 트랜스포머로, 픽셀 단위 카메라 제어가 가능한 이미지·영상 생성(최대 1분, 1440p), 실제 장면의 3D 재구성, 시공간 시뮬레이션, 텍스트 기반 이미지·360도 파노라마 생성 등 폭넓은 작업을 수행합니다. 이는 창작 도구를 넘어 로봇 학습용 Real-to-Sim 워크플로우 등 공간 지능(Spatial Intelligence) 분야의 핵심 인프라로 주목받습니다.

번역된 본문

월드 모델(World Model)은 모든 가능한 세계를 생성, 재구성, 시뮬레이션합니다. 월드 모델은 세계가 어떻게 보이고, 어떻게 행동하며, 어떻게 변화하는지 이해함으로써 창작자를 위해 상상 속 세계를 렌더링하고, 실제 세계를 고해상도로 시뮬레이션하며, 로봇이 행동을 계획하도록 돕습니다. World Labs는 공간 지능(spatial intelligence)을 추구하기 위해 이러한 범용 월드 모델을 개발하고 있습니다. 오늘 우리는 차세대 월드 모델인 아틀라스(Atlas)를 소개합니다.

아틀라스는 텍스트, 이미지, 비디오, 3D를 자연스럽게 다룰 수 있도록 처음부터 사전학습된 올라인 모델(omni model)입니다. 아틀라스는 멀티모달 자기회귀 확산 트랜스포머(multimodal autoregressive diffusion transformer)로, 모든 입력이 하나의 공유된 공간적 컨텍스트(spatial context)로 결합됩니다. 아틀라스는 이 컨텍스트를 활용해 다음에 올 내용을 생성하며, 지금까지 본 모든 것과 3D적으로 일관성을 유지하고 그 너머에 있는 것을 상상해냅니다.

아틀라스는 확장 가능하도록 설계되었습니다. 학습 컴퓨팅이 늘어날수록 성능이 향상되며, 앞으로 계속 확장해도 이 추세가 유지될 것으로 기대합니다.

아틀라스는 월드 생성, 재구성, 시뮬레이션에 걸친 폭넓은 작업을 수행할 수 있습니다:

  • 카메라 제어 생성(Camera-Controlled Generation): 아틀라스는 하나 이상의 이미지로부터 픽셀 단위로 정밀한 카메라 제어와 함께 이미지와 비디오를 생성하며, 최대 1분 길이의 1440p 비디오를 출력합니다.

  • 공간 재구성(Spatial Reconstruction): 아틀라스는 1장에서 수십 장의 입력 이미지로 실제 세계 장면을 재구성합니다. 새로운 시점의 이미지 프레임과 명시적 3D 출력을 모두 생성하며, 3D 재구성에 특화된 최신 모델들을 능가합니다.

  • 시공간 시뮬레이션(Space-Time Simulation): 아틀라스는 입력 비디오에서 공간과 시간을 모델링하여 극적인 시각 효과를 위한 비디오 리프레이밍(reframing)과 로봇 분야의 Real-to-Sim 워크플로우를 가능하게 합니다.

  • 이미지 생성(Image Generation): 아틀라스는 텍스트로부터 이미지와 360도 파노라마를 생성합니다. 복잡한 프롬프트를 따르고, 텍스트를 렌더링하며, 다양한 시각 스타일을 생성할 수 있습니다.

아틀라스는 향후 버전의 마블(Marble) 및 World Labs의 다른 제품들을 구동할 예정입니다. 아틀라스에 대한 얼리 액세스를 신청하세요.

카메라 제어 생성

아틀라스는 하나 이상의 참조 이미지를 받아 지정한 임의의 카메라 위치와 각도에서 새로운 시점을 생성합니다. 생성된 시점은 입력 이미지의 내용과 형상과 일치하며, 입력에 보이지 않는 장면 부분을 상상하며 그 너머로 부드럽게 외삽(extrapolate)합니다. 아틀라스는 다양한 유형의 장면, 시각 스타일, 카메라 움직임을 처리합니다.

픽셀 단위의 정밀한 카메라 제어

아틀라스는 정밀한 카메라 지오메트리를 기본 입력 유형으로 사용하여, 텍스트 기반의 대략적인 카메라 지시를 넘어섭니다. 이를 통해 모든 샷의 구도를 잡고 모든 움직임을 제어할 수 있습니다. 여기 예시에서 아틀라스는 단일 입력 이미지로부터 완전한 장면을 생성합니다. 입력 이미지의 내용과 폭넓은 세계 지식을 활용해 새로운 각도에서 장면이 어떻게 보일지 상상합니다. 예를 들어 로봇의 뒷면을 생성하고, 수영장 옆에 잔디밭이 있어야 한다고 추론합니다.

공간적 컨텍스트를 활용한 생성

LLM과 마찬가지로 아틀라스는 먼저 입력을 컨텍스트로 인코딩한 다음, 그 컨텍스트에 조건을 맞춰 출력을 생성합니다. 그러나 아틀라스의 독특한 점은 각 이미지가 3D 공간상의 위치에 기반(grounded)한다는 것이며, 이것이 '공간적 컨텍스트'를 형성합니다. 이 공간적 컨텍스트를 관리하면 완전히 새로운 종류의 창작 제어가 가능해집니다. 예를 들어 서로 관련 없는 두 참조 이미지를 컨텍스트에 배치하고 3D 공간에서 위치를 지정하면, 아틀라스는 두 이미지 사이를 부드럽게 보간하는 세계를 생성합니다. 이러한 예시는 모델의 세계 지식과 창의성을 보여줍니다. 서로 관련 없는 입력 이미지 사이에 문, 복도, 구석 공간 등의 전환부를 상상해 만들어냅니다.

제어 가능한 롱 비디오

아틀라스는 카메라 움직임과 공간적 컨텍스트 관리를 결합하여 정밀한 제어로 긴 비디오를 생성할 수 있습니다. 모든 장면과 모든 카메라 앵글을 직접 설계할 수 있습니다. 이는 당신을 감독의 자리에 앉히는 것입니다. 장면을 직접 연출하는 것이지, 슬롯머신 레버를 당기는 것이 아닙니다. 아래 예시에서는 소수의 참조 이미지를 사용해 1440p 해상도의 1분 길이 비디오를 생성합니다. 장면을 가로지르는 카메라 경로를 직접 설계하면, 아틀라스가 일관된 세계를 생성합니다. 비디오의 나머지 부분도

원문 보기
원문 보기 (영어)
World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave, and evolve so that we can render imagined worlds for creative users, simulate the real world in high fidelity, and help robots plan actions. At World Labs, we build these general purpose world models in pursuit of spatial intelligence. Today we are introducing Atlas, our next-generation world model. Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context. Atlas uses that context to generate what comes next, staying consistent in 3D with everything it has seen and imagining what lies beyond it. Atlas is built to scale: its performance improves with increased training compute, and we expect this trend to hold as we continue scaling. Atlas can perform a broad range of tasks spanning world generation, reconstruction, and simulation: Camera-Controlled Generation : Atlas generates images and videos from one or more images with pixel-perfect camera control, outputting up to 1 minute of video at 1440p. Spatial Reconstruction : Atlas reconstructs real world scenes from one to dozens of input images. It generates both image frames from novel views and explicit 3D outputs, outperforming state-of-the-art models specialized for 3D reconstruction. Space-Time Simulation : Atlas models space and time from input videos, reframing videos for dramatic visual effects and enabling Real-to-Sim workflows for robotics. Image Generation : Atlas generates images and 360 panoramas from text; it can follow complex prompts, render text, and generate a wide variety of visual styles. Atlas will power future versions of Marble and other products from World Labs. Request early access to Atlas Camera-Controlled Generation Atlas takes one or more reference images and generates new views at any camera position and angle you specify. Generated views match the content and geometry of the input images, smoothly extrapolating beyond them to imagine parts of the scene not visible in the inputs. Atlas handles a broad range of scene types, visual styles, and camera motions. Pixel-Perfect Camera Control Atlas uses precise camera geometry as a native input type, going beyond coarse text-based instructions for camera control. This lets you frame every shot and control every motion. In the examples here, Atlas generates a complete scene from a single input image . It uses the content of the input image along with its broad world knowledge to imagine what the scene should look like from new angles. For example, it generates the back side of the robot, and it guesses that there should be a grassy lawn next to the pool. Generating with Spatial Context Similar to an LLM, Atlas first encodes its inputs into a context, then generates outputs conditioned on the context. However Atlas is unique because each image is grounded at a 3D position in space; this forms a spatial context . Managing this spatial context unlocks entirely new kinds of creative control. For example, you can place two unrelated reference images in the context, position them in 3D space, and Atlas generates a world that smoothly interpolates between them. These examples demonstrate the model's world knowledge and creativity; it imagines doorways, hallways, nooks, and other transitions between otherwise unrelated input images. Controllable Long Videos Atlas lets you generate long videos with precise control by combining camera movement and spatial context management. You design every scene and every camera angle. This puts you in the director's chair: you are staging the scene, not pulling the lever of a slot machine. In the example below, we generate a 1 minute video at 1440p resolution using a small number of reference images. We hand-design a camera path through the scene, and Atlas generates a coherent world. The rest of the videos on this page have been compressed to optimize page performance. Spatial Reconstruction Atlas reconstructs real-world spaces from one or more input images. It doesn't require special capture equipment or hundreds of dense views to faithfully reconstruct objects and scenes. We believe Atlas is a major step forward toward solving the problem of novel view synthesis from sparse input images, a decades-old fundamental problem in 3D computer vision. Reconstructing from Multiple Images Atlas can take a variable number of input views of a scene. When parts of the world are not visible in the input views, Atlas imagines a plausible way to fill in the gaps by drawing from its rich world knowledge. But sometimes you don't want imagination; you might want an exact reconstruction of a real-world location. Passing more input images gives Atlas more context: the more it sees, the less it imagines. Atlas typically gives faithful reconstructions with as few as two or three images, outperforming state-of-the-art results by models specially trained only for 3D reconstruction. However, Atlas can also make use of over a hundred input images in its spatial context, allowing for faithful recreation of real world environments. In the first example below, Atlas generates an aerial view of the scene from just a single ground-level photo. The garden visible in the single input photo is accurately recreated in the model output, but the rest of the scene is imagined. After adding a second real-world input image of the cottage next to the garden, the model's output shows both the garden and the cottage, but it still imagines the house to the left. After adding a third input image of the main house, the entire scene is accurately depicted. In the second example, we build up Stanford's Main Quad piece by piece, beginning with the grassy main entrance and ending with the colorful mosaics decorating the facade of Memorial Church. Though Atlas only receives two to twenty-five ground-level input images, it can generate paths from aerial views flying far above the campus. Reconstructing Diverse Paths Atlas can generate many different trajectories through the same scene, giving new perspectives on the same input images. No matter how many times you change the camera path, the scene stays consistent. In the example below, we show that given a small set of input images, Atlas can generate multiple camera paths through the same scene. Different camera paths can emphasize various parts of the scene, or change moods by varying in speed, length, or complexity. Explicit 3D Outputs In the results above you have seen Atlas output 2D images and videos, which are sufficient for some applications. But workflows in robotics, gaming, design, VFX and beyond often require explicit 3D outputs. Atlas natively operates on both 2D image frames and 3D depth maps, enabling it to output worlds as point clouds or 3D Gaussian splats. From a single input image, Atlas produces a full 3D world by jointly generating new views and estimating their geometry. From a video of a real space, it predicts the depth of every frame and combines them into a 3D reconstruction. In either case, Atlas fills in regions that no camera ever saw. Point clouds estimate a scene's geometry, but 3D Gaussian splats make it usable. Atlas fills the remaining gaps and turns the point cloud into a complete splat scene that renders on-device at high resolution and framerates. This is the same representation used in Marble , enabling Atlas to integrate naturally with the rest of our products. Space-Time Simulation Atlas serves as a world simulator. It understands both the spatial structure of the world and how the world evolves over time. Combining its spatial and temporal abilities leads to new applications for VFX, robotics, and beyond. Reframing Video Atlas turns a handful of ordinary cameras into a "bullet time" multiview capture studio. With footage from as few as three cameras, Atlas can