메뉴
BL
Wired AI • 57일 전

구글 '제미나이 로보틱스 2'로 물리적 세계 정복

IMP
8/10
핵심 요약

구글 딥마인드가 전신 움직임과 섬세한 작업을 제어할 수 있는 새로운 로봇 AI 모델 '제미나이 로보틱스 2'를 공개했습니다. 이 모델은 시각 언어 모델(VLM)과 시각 언어 행동(VLA) 모델을 결합하여 주변 환경을 이해하고 스스로 복잡한 작업을 수행할 수 있게 해줍니다. 디지털 영역을 넘어 물리적 세계에 AI를 적용하려는 구글의 의도를 보여주는 중요한 이정표입니다.

번역된 본문

구글 딥마인드가 인공지능 모델 '제미나이'의 새로운 버전을 방금 공개했는데, 이 모델은 전구를 끼우거나 쓰레기 봉투를 묶는 등 섬세한 작업이 가능한 휴머노이드를 포함해 다양한 로봇을 제어할 수 있습니다. 제미나이 로보틱스 2(Gemini Robotics 2)는 여러 가지 다른 AI 모델을 하나의 시스템으로 결합합니다. 이를 종합적으로 활용하면 로봇이 주변 환경을 이해하고 그 안에서 어떻게 행동해야 할지 파악할 수 있습니다. 이미지와 비디오를 이해하는 비전 언어 모델(VLM, Vision Language Model)은 인간과 소통하고 다양한 작업을 수행하는 방법을 추론할 수 있습니다. 물리적 공간에서의 움직임을 이해하도록 훈련된 두 개의 비전 언어 액션 모델(VLA, Vision Language Action)은 로봇의 전신 움직임과 그리퍼나 손의 움직임을 제어합니다.

발표에 앞서 공유된 영상 데모에서 회사는 통합 모델을 사용하여 여러 다른 로봇이 복잡한 작업을 자율적으로 수행하는 모습을 보여주었습니다. 한 데모에서는 앱트로닉(Apptronik)의 아폴로 2(Apollo 2) 로봇이 샤르파(Sharpa)라는 회사의 손을 사용해 선반을 정리했습니다. 구글 딥마인드는 인간 원격 조종, 비디오 예시 및 시뮬레이션을 혼합하여 이러한 작업을 수행하도록 모델을 훈련시켰습니다. 아직 AI 모델이 구체적인 훈련 없이 광범위하고 복잡한 작업을 수행하는 것은 가능하지 않습니다. 앤스로픽(Anthropic)과 오픈AI가 챗봇과 AI 코딩 도구에서 선도적인 위치를 차지하고 있지만, 구글은 로봇 공학 연구에서 더 강력한 실적을 가지고 있으며 AI를 훈련시켜 유용한 일을 하는 로봇을 만드는 데 중요한 연구를 발표해 왔습니다. 이번 발표는 검색 거대 기업인 구글이 AI의 잠재력을 최대한 발휘하기 위해 디지털 영역을 벗어나야 한다고 베팅하고 있다는 또 다른 신호입니다. (이전에 구글은 다리가 있는 로봇 분야의 선두주자인 보스턴 다이내믹스와 협력하여那些 기계의 두뇌를 제공한 바 있습니다.)

구글 딥마인드의 로봇 담당 책임자인 카롤리나 파라다(Carolina Parada)는 WIRED와의 인터뷰에서 "이는 우리가 인간이 할 수 있는 모든 일을 로봇이 할 수 있게 만드는 '물리적 AGI(범용 인공지능)'를 향한 진정한 경로의 또 다른 이정표"라고 말합니다. 하지만 최첨단 AI 모델이 로봇에 접근하여 작업장이나 가정을 돌아다니며 물건을 조작할 수 있도록 하는 것은 위험을 수반합니다. 이전 연구에 따르면 최첨단 AI를 사용하여 로봇을 제어하면 예기치 않고 때로는 위험한 행동이 발생할 수 있는 것으로 나타났습니다. 또한 이러한 모델이 디지털 영역에서 갑작스럽거나 원치 않는 행동을 할 수 있다는 아이디어는 최근 오픈AI가 개발한 미공개 AI 에이전트가 여러 시스템을 해킹한 사건을 통해 분명해졌습니다.

파라다는 "안전 문제는 로봇을 훨씬 더 많은 다른 상황에 투입하기 때문에 더욱 시급합니다"라고 말합니다. "수많은 불확실성이 나타날 수 있으므로 안전 문제를 더욱 깊이 이해할 수 있어야 합니다." 파라다는 구글이 각 모델 계층에 가드레일을 적용하는 다층적 접근 방식을 취하고 있다고 밝혔습니다. 또한 로봇을 제어하기 위해 협력하는 다양한 AI 시스템의 안전성을 측정하기 위한 새로운 벤치마크인 ASIMOV-Agentic를 도입하고 있습니다. 이 벤치마크는 명령이 해로운 결과나 불확실한 결과를 초래할지 여부를 감지합니다. 회사의 CEO인 데미스 허사비스(Demis Hassabis)는 이전에 WIRED와의 인터뷰에서 스마트폰용 안드로이드 운영체제와 유사하게 다양한 로봇을 위한 AI 운영체제를 개발하길 희망한다고 밝힌 바 있습니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Google DeepMind just released a new version of its artificial intelligence model Gemini , and it can control a range of different robots—including humanoids capable of dextrous tasks like screwing in lightbulbs and tying trash bags. Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot’s full-body movement as well as the movements of grippers or hands. In video demonstrations shared ahead of the release, the company showed several different robots performing complex tasks autonomously using the amalgamated model. In one demo, Apptronik’s Apollo 2 robot used hands from a company called Sharpa to tidy shelves. Google DeepMind trained the model to perform these tasks using a mix of human teleoperation, video examples, and simulations—it’s not yet possible for AI models to perform a wide range of complex tasks without specific training. Although Anthropic and OpenAI have taken a lead with chatbots and AI coding tools, Google has a stronger track record in robotics research, and has published important work on using AI to train robots to do useful things. The release is another sign that the search giant is betting AI will need to break free from the digital realm to realize its full potential. (It previously partnered with Boston Dynamics , a leader in legged robots, to provide the brains for those machines.) “It's another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can,” Carolina Parada, head of robotics at Google DeepMind, tells WIRED. Giving frontier AI models access to robots so that they can wander around workplaces or homes and manipulate objects does, however, come with risks. Previous research has shown that using frontier AI to control robots can produce unexpected and sometimes dangerous behavior . And the idea that these models can take sudden or unwanted actions in the digital realm became apparent recently, when an unreleased AI agent developed by OpenAI hacked several systems . “The safety question is even more pressing because you're putting them in a lot of other situations,” Parada says. “There's a lot of uncertainty that will show up, and so you want to be able to understand the safety question more deeply.” Parada says Google takes a multi-layered approach to safety, with guardrails applied on each model layer. It’s also introducing ASIMOV-Agentic, a new benchmark for measuring the safety of various AI systems collaborating to control a robot. The benchmark detects whether a command will result in harmful or uncertain outcome. The company’s CEO, Demis Hassabis, previously told WIRED that he hopes to develop an AI operating system for many different robots similar to the Android operating system for smartphones.