메뉴
BL
The Decoder • 24일 전

월드랩스, 사진 몇 장만으로 3D 세계를 생성·재구축·시뮬레이션하는 통합 AI 모델 '아틀라스' 공개

IMP
8/10
핵심 요약

이비페이페이(Fei-Fei Li)가 공동 창업한 World Labs가 몇 장의 사진만으로 3D 장면을 생성, 재구축, 시뮬레이션하는 단일 올모델 'Atlas'를 발표했습니다. Atlas는 카메라 위치를 직접 제어해 1440p 최대 1분 영상을 출력하고, 포인트 클라우드와 3D 가우시안 스플랫 등 네이티브 3D 데이터를 결과물로 내놓으며, 로봇학 훈련 데이터 생성을 위한 real-to-sim 도구로도 활용될 수 있습니다. 전용 3D 모델들을 능가한다는 주장이 사실이라면 공간 지능(spacial intelligence) 분야의 판도를 바꿀 중요한 발표입니다.

번역된 본문

월드랩스(World Labs), 사진 몇 장만으로 3D 세계를 생성·재구축·시뮬레이션하는 단일 AI 모델 '아틀라스(Atlas)' 공개

AI 연구자 리페이페이(Fei-Fei Li)가 공동 창업한 World Labs가 몇 장의 이미지만으로 3D 장면을 생성, 재구축, 시뮬레이션하는 월드 모델 '아틀라스(Atlas)'를 발표했다. 이 회사는 아틀라스가 각 전용 모델들의 고유 작업에서도 그 모델들을 능가한다고 주장하며, 이 경우 많은 전용 모델들이 불필요해질 수 있다고 밝혔다.

창립 이래 World Labs는 AI가 인간처럼 3D 공간을 이해해야 한다는 '공간 지능(spatial intelligence)' 목표를 추구해왔다. 아틀라스는 이를 대규모로 실현하기 위한 회사의 첫 모델이다. 평면 이미지나 짧은 영상 클립을 만들어내는 대신, 장면이 어떤 각도에서 어떻게 보이는지, 시간에 따라 어떻게 변하는지를 파악한다.

World Labs는 아틀라스를 텍스트, 이미지, 영상, 3D 데이터로 처음부터 학습된 '올모델(omni-model)'이라고 설명한다. 모든 입력은 평면 시퀀스가 아니라 3D 공간의 특정 위치에 고정(앵커링)되어 처리된다. 회사는 이러한 공유된 공간 이해를 '공간적 컨텍스트(spatial context)'라 부르며, 모델이 새 프레임이나 시점을 생성할 때 이를 활용한다. World Labs에 따르면 이 앵커링 방식이 아틀라스를 순수 언어 모델이나 영상 모델과 구별하는 핵심이다.

리페이페이는 2025년 11월 에세이에서 이 문제를 정확히 짚은 바 있다. 그녀는 현재의 멀티모달 언어 모델과 비디오 확산 모델이 데이터를 1차원 또는 2차원 시퀀스로 분해하기 때문에 간단한 공간 작업조차 불필요하게 어렵다고 주장했다. 필요한 것은 토큰화, 컨텍스트, 메모리를 3D·4D를 인식하는 방식으로 구성하는 아키텍처이다.

1440p 최대 1분 영상

카메라 제어 생성의 경우, 아틀라스는 한 장 이상의 이미지를 받아 사용자가 자유롭게 선택한 카메라 위치와 각도에서 새로운 뷰를 생성한다. 카메라 움직임은 많은 영상 모델이 요구하는 것처럼 텍스트 프롬프트로 기술되지 않고 직접적인 기하학적 입력으로 전달된다. 모델은 1440p 해상도로 최대 1분 길이의 영상을 출력한다. World Labs의 표현을 빌리면, 사용자는 '슬롯머신 레버를 당기는' 대신 모든 샷을 직접 제어할 수 있으며, 이는 제어된 생성과 무작위 출력을 구분하는 기준이 된다.

공간 재구축 측면에서 아틀라스는 단 한 장에서 수십 장의 입력 이미지만으로 특수 촬영 장비 없이 실제 장면을 재구축한다. 입력 이미지가 많을수록 모델이 자체 지식으로 채워 넣어야 할 부분이 줄어든다. World Labs에 따르면 이미지 두세 장만으로도 아틀라스는 충실한 결과를 내며 전용 3D 모델을 능가한다. 100장이 넘는 입력도 처리할 수 있다. 한 시연에서 모델은 지상에서 촬영한 스탠퍼드 대학 메인 쿼드(Main Quad) 사진 2장에서 25장까지 점진적으로 장면을 조립하고, 캠퍼스 상공 훨씬 위에서 본 항공 뷰를 생성했다. 바로 이 지점에서 기존 모델들이 무너지는 경향이 있다. OpenWorldLib 프레임워크 내 비교에서 VGGT, InfiniteVGGT 같은 시스템은 카메라가 크게 이동하자마자 기하학적 비일관성과 흐릿한 텍스처를 보였다.

네이티브 3D 출력과 로봇 시뮬레이션

아틀라스는 RGB와 함께 깊이 정보를 처리하기 때문에 이미지나 영상이 아니라 실제 3D 데이터로 결과를 출력할 수 있다. 지원 포맷에는 포인트 클라우드(point cloud)와 3D 가우시안 스플랫(3D Gaussian splats)이 포함되며, 이는 수많은 작은 공간 데이터 포인트로 장면을 구성해 어떤 각도에서든 부드럽게 볼 수 있게 한다. 이는 회사의 기존 제품인 마블(Marble)이 사용하는 표현 방식과 일치한다.

시뮬레이터로서 아틀라스는 공간과 시간을 함께 모델링한다. 불과 몇 대의 카메라로 촬영한 영상만으로 장면을 멈춰 세우고 평소에는 불가능한 각도에서 볼 수 있는 '불릿 타임(bullet time)' 효과를 만들어낸다. 시연 영상은 전문 장비가 아닌 몇 대의 스마트폰과 액션캠으로 촬영되었다.

로봇 분야에서는 아틀라스가 real-to-sim 도구로 활용된다. 방을 재구축하고, 시뮬레이션된 로봇의 센서가 이동 경로를 따라 보게 될 이미지와 깊이 데이터를 생성한다. 몇 장의 사진만으로 객체, 위치, 조명, 배경을 바꿔가며 로봇의 파지(grasping)와 이동 작업을 시뮬레이션할 수 있다. 목표는 직접 촬영 없이 로봇을 위한 다양한 훈련 데이터를 생성하는 것이다.

원문 보기
원문 보기 (영어)
World Labs unveils Atlas, a single AI model that generates, reconstructs, and simulates 3D worlds from just a few photos Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 2, 2026 World Labs World Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas, a world model that generates, reconstructs, and simulates 3D scenes from just a few images. The company claims it beats specialized models at their own tasks, which could make many of them unnecessary. Since its founding, World Labs has pursued the goal of "spatial intelligence" , the idea that AI should understand 3D space the way humans do. Atlas is the company's first model built to do that at scale. Rather than producing flat images or video clips, it grasps how a scene looks from any angle and how it changes over time. World Labs describes Atlas as an omni-model trained from scratch on text, images, video, and 3D data. Every input gets anchored to a specific position in 3D space rather than processed as a flat sequence. The company calls this shared spatial understanding "spatial context," and it's what the model uses to generate each new frame or viewpoint. According to World Labs, this anchoring separates Atlas from pure language or video models. Fei-Fei Li laid out this exact problem in a November 2025 essay. Current multimodal language models and video diffusion models break data into one- or two-dimensional sequences, she argued, which makes even simple spatial tasks needlessly hard. What's needed are architectures that organize tokenization, context, and memory in a 3D- or 4D-aware way. One minute of video at 1440p For camera-controlled generation, Atlas takes one or more images and produces new views at freely chosen camera positions and angles. Camera movement is passed as a direct geometric input rather than described through text prompts, as many video models require. The model outputs up to one minute of video at 1440p. Users can control every shot themselves instead of "pulling the lever on a slot machine," as World Labs put it, drawing a line between controlled generation and random output. For spatial reconstruction, Atlas rebuilds real scenes from as few as one to several dozen input images without special capture equipment. The more images it receives, the less it has to fill in from its own knowledge. With just two or three images, Atlas delivers faithful results and outperforms specialized 3D models, according to World Labs. It can also handle over a hundred inputs. In one demo, the model progressively assembles Stanford's Main Quad from two to 25 ground-level photos and generates aerial views far above the campus. This is where existing models tend to fall apart. In a comparison within the OpenWorldLib framework , systems like VGGT and InfiniteVGGT showed geometric inconsistencies and blurry textures as soon as the camera moved significantly. Native 3D output and robotics simulation Atlas can output results as actual 3D data, not just images or video, because it processes depth information alongside RGB. Supported formats include point clouds and 3D Gaussian splats , which build a scene from many small spatial data points that can be viewed smoothly from any angle. This matches the representation used in Marble, the company's existing product . As a simulator, Atlas models space and time together. From footage captured by just a few cameras, it can produce a "bullet time" effect that freezes a scene and lets users view it from otherwise impossible angles. The demo footage was shot with a handful of smartphones and action cameras, not professional gear. For robotics, Atlas serves as a real-to-sim tool. It reconstructs a room and generates the image and depth data that a simulated robot's sensors would see along its path. From just a few photos, users can simulate and vary grasping and movement tasks by swapping out objects, positions, lighting, or backgrounds. The goal is to produce diverse training data for robots without capturing every situation in the real world. World Labs showed this approach in August 2026 with its real-to-sim-to-real engine as a standalone product. That engine creates thousands of variants from a single real-world task and trains control models entirely in simulation. On five robot platforms, the models ran for an hour each without human intervention, according to the company. The technology came from SceniX, a startup World Labs acquired in July. Text-to-image generation isn't the main focus, the company says, but Atlas can also follow complex prompts, render text, produce different visual styles, and create 360-degree panoramas. Speed from language models, quality from diffusion Atlas combines ideas from both language models and video models. It generates output piece by piece like a language model, so it can use the same speedup techniques, such as KV caching. But it also uses the diffusion principle from image and video models, gradually filtering output out of noise. That side gives it access to methods that shorten the denoising process or boost image quality. World Labs says no single benchmark captures what Atlas can do, but points to two sets of tests. In camera-controlled generation judged by external human evaluators, and in few-view 3D reconstruction, Atlas outperforms more specialized models. Human evaluators preferred Atlas in 75 percent of comparisons against MiniMax H3 , 81 percent against Gemini Omni Flash , 86 percent against Happy Horse 1.1, 93 percent against, and 94 percent against Seedance 2.5 . For reconstruction, Atlas leads with a median error of 25.3, ahead of Pi3X and VGGT-Ω 1B. The company says Atlas's performance improves with more training compute and expects that trend to hold as it scales. Atlas will power future versions of Marble and other products, and is currently available through an early-access program for select partners. From walkable photos to an omni-model World Labs was founded in 2024 by Fei-Fei Li, who created ImageNet and led Google Cloud's AI division from 2017 to 2018. The company at launch from Andreessen Horowitz, AMD, Intel, and Nvidia. A first system in late 2024 turned, though users could only move a few virtual meters before hitting invisible boundaries. Marble followed in November 2025. In February 2026 came a $1 billion funding round from Autodesk, Andreessen Horowitz, Nvidia, and AMD. Bloomberg had previously reported talks at a $5 billion valuation. What counts as a world model remains contested among researchers. An international team led by Peking University proposed a unified definition in April 2026 through OpenWorldLib, excluding pure text-to-video models because they lack feedback loops with the real world. 3D reconstruction and simulators like those in Atlas qualify as core building blocks in that framework because they provide environments where physical rules can be verified. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->