메뉴
BL
The Decoder 57일 전

엔비디아, 세계 모델 및 휴머노이드 로봇 공개

IMP
9/10
핵심 요약

엔비디아가 GTC 타이베이에서 새로운 세계 모델 'Cosmos 3', 자율주행 모델 'Alpamayo 2 Super', 그리고 휴머노이드 로봇 공개 플랫폼을 발표하며 물리적 AI 분야에 대대적인 투자를 단행했습니다. 이번 발표는 로봇 공학, 자율주행, 비디오 분석 시스템 개발자들이 합성 데이터를 생성하고 시뮬레이션을 고도화할 수 있는 강력한 인프라를 제공한다는 점에서 업계 실무자들에게 매우 중요한 의미를 갖습니다. 모든 모델과 프레임워크는 오픈소스 라이선스로 제공되어 관련 산업의 기술 발전 속도를 크게 앞당길 것으로 기대됩니다.

번역된 본문

엔비디아, 새로운 세계 모델, 자율주행 두뇌 및 오픈 소스 휴머노이드 로봇으로 GTC 타이베이에서 물리적 AI에 대대적인 투자 단행

Maximilian Schreiner | 2026년 6월 1일

엔비디아는 GTC 타이베이 행사를 활용해 로봇, 자율주행 차량, 비디오 시스템을 위한 일련의 새로운 모델을 공개했다. 이번 발표의 핵심은 새로운 세계 모델(World Model)인 'Cosmos 3', 대폭 확장된 자율주행 모델인 'Alpamayo 2 Super', 그리고 휴머노이드 로봇을 위한 오픈 레퍼런스 플랫폼이다.

Cosmos 3는 텍스트, 이미지, 비디오, 주변 소리, 그리고 동작 데이터를 단일 시스템에서 처리하는 엔비디아의 차세대 오픈 '옴니 모델(Omnimodel)'이다. 로봇, 자율주행 차량, 비디오 감시 시스템을 개발하는 개발자들은 이 모델을 활용해 실제 환경에서 위험하거나 재현하기 까다로운 상황을 일일이 테스트하지 않고도, 합성 훈련 데이터를 생성하고 장면을 해석하며 미래의 세계 상태를 예측할 수 있다.

엔비디아는 세 가지 주요 활용 사례를 제시했다. 첫째, 비전-언어 모델(Vision-Language Model)로서 Cosmos 3는 비디오를 분석한다. 예를 들어 파트너사인 Linker Vision이 이미 진행 중인 스마트 시티의 교통 이상 현상 감지가 여기에 해당한다. 둘째, 세계 모델로서 뼈 골절 위기 상황이나 창고 내의 특이한 물체 배치처럼 희귀한 상황의 사진 같은 실감 나는 비디오 시퀀스를 생성한다. 셋째, 이른바 세계-행동 모델(World-Action Model)의 기반으로서 산업용 파트너인 Agile Robots가 시연하는 것처럼 로봇이 물건을 집어 옮기는 등의 작업을 학습하는 데 사용되는 관절 각도나 그리퍼 위치 같은 수치적 동작 데이터를 생성한다.

이 아키텍처는 '혼합 트랜스포머(Mixture-of-Transformers)' 방식을 사용한다. 즉, 하나의 추론용 트랜스포머가 장면을 분석한 다음, 두 번째 생성용 트랜스포머가 해당 분석을 바탕으로 비디오, 설명 또는 동작 궤적을 생성한다. 훈련 데이터는 텍스트, 이미지, 비디오, 오디오 및 동작 데이터에 걸쳐 수십억 개의 예시로 구성되었다.

엔비디아는 세 가지 변형을 제공한다. Cosmos 3 Super는 현재 최고의 품질을 제공하며, Nano는 빠른 추론을 위해 구축되었고, 향후 출시될 Edge 모델은 임베디드 시스템에서의 실시간 작동을 목표로 한다. 이 모델들은 Hugging Face와 GitHub에서 OpenMDW-1.1 라이선스에 따라 공개된다. 이번 출시와 함께 Black Forest Labs, Runway, LTX, Generalist, Agile Robots, Skild AI 등이 포함된 파트너 그룹인 'Cosmos Coalition'도 발족했다. 실질적으로 이 연합은 엔비디아의 DGX Cloud 훈련 인프라를 활용하고, 그 대가로 모델과 데이터를 기여하는 형태의 동맹이다.

Alpamayo 2 Super, 로보택시를 위한 교사 모델로 설계되다

Alpamayo 제품군은 레벨 4(Level 4) 자율주행, 즉 정의된 지역 내에서 사람의 개입 없이 운행하는 로보택시를 위한 엔비디아의 오픈 모델 시리즈이다. 이 모델들은 카메라 이미지를 입력받아 주행 결정을 도출하고 구체적인 주행 궤적을 출력한다. 이전 버전인 Alpamayo 1 Nano 및 1.5 Nano는 각각 100억 개의 파라미터를 갖추고 있었다. Alpamayo 2 Super는 최상위 모델로서 이전 세대를 대체하며 320억 개의 파라미터를 탑재했다. 이러한 도약은 공간 이해력과 희귀 상황 처리 능력을 향상시키기 위함이다.

새로운 기능으로는 모델이 하위 궤적과 함께 하위 스트림 플래너(Planner)에 전달하는 '차선 변경', '정지', '양보'와 같은 이른바 메타 액션(Meta-action) 출력이 추가되었다. 또한 인식 범위가 전방 카메라에만 국한되지 않고 차량 전체로 확장되었다. 모든 결정에는 '인과 관계 체인(Chain of causation)'이라는 텍스트 형태의 추론 체인이 동반되며, 엔비디아는 이것이 안전 문서화 및 규제 검토를 위해 설계되었다고 밝혔다. 이는 AI 정렬(AI Alignment) 논쟁에서 익숙했던 질문을 주행 안전 논의로 가져온 것이다. 즉, 이러한 추론 과정이 네트워크 내부에서 실제로 일어나는 일을 얼마나 신뢰할 수 있게 반영하는가 하는 점이다.

엔비디아는 이 대형 모델이 교사 모델(Teacher Model)로 사용될 것이라고 밝혔다. 제조사들은 이 모델을 기반으로 더 작은 모델을 증류(Distill)하여 차량용 Drive AGX Thor 칩에 탑재해 실행할 수 있다. 또한 엔비디아는 시뮬레이션 환경에서의 폐루프 강화 학습(Closed-loop reinforcement learning)을 위한 오픈소스 프레임워크인 'AlpaGym'과 희귀 교통 시나리오를 위한 생성 모델인 'OmniDreams'도 함께 공개했다. 엔비디아는 Waymo나 Tesla의 시스템과 비교하는 신뢰할 수 있는 외부 벤치마크 수치는 제공하지 않았다. 코드와 모델 가중치는 곧 공개될 예정이다.

원문 보기
원문 보기 (영어)
Nvidia bets big on physical AI at GTC Taipei with a new world model, driving brain, and open humanoid robot Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jun 1, 2026 Nano Banana Pro prompted by THE DECODER Nvidia used GTC Taipei to launch a series of models for robots, autonomous vehicles, and video systems. The centerpieces are the new world model Cosmos 3, a significantly scaled-up driving model called Alpamayo 2 Super, and an open reference platform for humanoid robots. Cosmos 3 is Nvidia's next version of its open "omnimodel," which processes text, images, video, ambient audio, and action data in a single system. Developers building robots, autonomous vehicles, and video surveillance systems can use it to generate synthetic training data, interpret scenes, and predict future world states without having to painstakingly recreate those situations in the real world. Nvidia names three use cases. As a vision-language model, Cosmos 3 analyzes video, for example to detect traffic anomalies in smart cities, as partner Linker Vision is already doing. As a world model, it generates photorealistic video sequences of rare situations like near-misses or unusual object arrangements in a warehouse. And as the basis for so-called world-action models, it produces numerical motion data like joint angles or gripper positions that robots use to learn tasks such as picking and placing, as industrial partner Agile Robots demonstrates. The architecture uses a mixture-of-transformers approach: one reasoning transformer analyzes a scene, then a second generation transformer produces videos, descriptions, or motion trajectories from that analysis. Training data included billions of examples spanning text, images, video, audio, and action data. Nvidia offers three variants: Cosmos 3 Super delivers the best current quality, Nano is built for fast inference, and a forthcoming Edge model targets real-time operation on embedded systems. The models are available under the OpenMDW-1.1 license on Hugging Face and GitHub. The release comes alongside the "Cosmos Coalition," a partner group that includes Black Forest Labs, Runway, LTX, Generalist, Agile Robots, and Skild AI. In practice, it's an alliance that uses Nvidia's DGX Cloud training infrastructure and contributes models and data in return. Alpamayo 2 Super is meant to be a teacher model for robotaxis The Alpamayo family is Nvidia's open model series for Level 4 autonomous driving, meaning robotaxis that operate without a human driver within a defined area. The models take in camera images, derive a driving decision, and output a concrete trajectory. Previous versions included Alpamayo 1 Nano and 1.5 Nano, each with ten billion parameters. Alpamayo 2 Super replaces that generation at the top end with 32 billion parameters. The jump is supposed to improve spatial understanding and handling of rare situations. New is the output of so-called meta-actions like "lane change," "stop," or "yield," which the model delivers to a downstream planner alongside the trajectory. Perception now also covers the entire vehicle rather than just the front cameras. Every decision comes with a "chain of causation," a textual reasoning chain that Nvidia says is designed for safety documentation and regulatory review. This brings a familiar question from the AI alignment debate into the driving safety discussion: how reliably do these reasoning traces actually reflect what's happening inside the network ? Nvidia says the large model is intended as a teacher model. Manufacturers are supposed to use it to distill smaller models that then run on the vehicle-grade Drive AGX Thor chip. Nvidia is also releasing AlpaGym, an open-source framework for closed-loop reinforcement learning in simulation, and OmniDreams, a generative model for rare traffic scenarios. Nvidia doesn't provide any reliable external comparison numbers, for example against the stacks from Waymo or Tesla. Code and weights are expected to appear on GitHub and Hugging Face this summer. An open humanoid robot built on a Unitree chassis With the Isaac GR00T Reference Humanoid Robot , Nvidia is also releasing a reference platform for academic research in humanoid robotics. The roughly six-foot-tall robot is based on the Unitree H2 Plus chassis, paired with tactile five-finger hands from Sharpa, and powered by the Jetson AGX Thor T5000 with 2,070 FP4 teraflops. The system has 75 degrees of freedom in total. On the software side, it runs the Isaac GR00T stack , which covers teleoperation, simulation in Isaac Sim, foundation models, and ROS middleware. Nvidia isn't selling the robot itself. Instead, it points to Unitree, which plans to offer the hardware by late 2026. Research partners include Ai2, ETH Zurich, the Stanford Robotics Center, and the UC San Diego ARC Lab. In practice, Nvidia is trying to standardize a hardware-software bundle that deepens the robotics research community's reliance on Jetson chips and Isaac tooling. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->