메뉴
BL
Wired AI • 37일 전

현장에서 즉시 배우는 로봇에서 본 AI의 미래

IMP
7/10
핵심 요약

미국 매사추세츠의 스타트업 Generalist AI가 짧은 영상만 보고 특정 과제 훈련 없이 다양한 작업을 즉석에서 수행하고, 상황이 바뀌면 스스로 즉흥 대응하는 로봇 팔을 공개했다. 전통적인 수천 건 예시 학습 방식과 달리, 인간이 세상의 물리 법칙을 직관적으로 이해하듯 배우는 '범용 로봇 모델' 접근법이라는 점에서 주목할 만하다.

번역된 본문

지난주 나는 집에서 단 15분 거리를 이동해 로봇이 믿기 어려운 놀라운 일을 하는 모습을 보러 갔다. 매사추세츠주 케임브리지에 있는 Generalist AI라는 스타트업 사무실을 방문해, 로봇 팔이 컵 쌓기, 블록을 그릇에 넣기 같은 간단한 집안일을 수행하는 것을 지켜봤다. 로봇이 얼마나 빨리 상황을 파악하는지에 놀랐다—살아있는 인간을 연상시킬 정도였다. 로봇 팔은 짧은 교육 영상을 본 후 다양한 작업을 마스터했으며, 가장 인상적인 것은 특정 과제에 대한 별도 훈련이 전혀 없었다는 점이다.

가장 인상 깊었던 사례 중 하나는 쓰레받기와 솔을 사용해 블록을 그릇에 쓸어 넣으라는 지시를 받은 로봇이었다. 장면에서 솔을 치우자 로봇은 쓰레받기를 솔처럼 사용해 블록을 그릇 안으로 튕겨 넣는 즉흥적인 방법을 고안했다. 또 다른 사례에서 두 팔을 가진 로봇은 누군가 지갑을 열어 지폐를 꺼내는 영상 클립을 본 뒤, 다른 종류의 지갑을 열고 조심스럽게 지폐를 꺼냈다. 놀랍게도 지폐를 잡지 못하자 더 좋은 접근 각도를 얻기 위해 오른쪽 그리퍼에서 왼쪽 그리퍼로 전환했다. 옆에 서 있던 한 엔지니어가 말했다. "하, 저건 처음 보는군요."

"이것이 바로 사람들이 GPT-3에 대해 정말 흥분했던 종류의 것입니다." Generalist의 공동 창업자이자 CEO인 피트 플로렌스(Pete Florence)는 2020년 출시된 OpenAI의 획기적인 대형 언어 모델을 언급하며 나에게 말했다. "그 모델에 새로운 작업을 지시하면 실제로 수행할 가능성이 있었죠."

Generalist는 로봇에게 세상의 물리 법칙을 가르치는 데 집중하는 것으로 보이며, 이는 인간이 어릴 때부터 보여주는 직관적인 물리 감각에서 영감을 받은 듯하다. 이 덕분에 모델이 한 상황에서 배운 것을 다른 상황으로 전이하는 능력이 향상된 것으로 보인다. 실제로 이 회사의 일부 데모는 과제를 보여주면 즉흥적으로 실험하는 아이들의 모습을 떠올리게 했다. 연구자들은 로봇이 스스로 결정하는 행동에 종종 놀라곤 하는데, 예를 들어 한 로봇은 앞에 바나나가 놓이자 바나나로 물건을 쓸어 모으는 방법을 선택했다. 사소해 보일 수 있지만, 물리적 지능은 여전히 기계에 크게 부족한 요소이며, 아기들이 세상을 매우 효율적으로 학습하는 방식이 AI 연구자들에게 중요한 통찰을 제공할 수 있다.

나는 플로렌스와 공동 창업자이자 CTO인 앤드류 배리(Andrew Barry)를, 손에 특수 그리퍼를 끼고 로봇 훈련을 수행하는 팀들이 내려다보이는 회의실에서 만났다. 회사의 또 다른 공동 창업자이자 수석 과학자는 앤디 젱(Andy Zeng)이다. 이 세 사람은 인상적인 경력을 갖추고 있다. 이들은此前 구글 딥마인드(Google DeepMind)와 보스턴 다이내믹스(Boston Dynamics)에서 세계 최고 수준의 하드웨어 및 로봇 모델 연구에 종사했다.

전통적으로 AI 기반 로봇이 다양한 작업을 수행하도록 훈련하려면 수천 개의 예시를 모델에 공급해야 했다. 하지만 이는 유명할 정도로 불완전한 학습 방식이어서, 조명 같은 단순한 요소만 바꿔도 로봇은 작업에 어려움을 겪는다. Generalist를 비롯한 일부 로보틱스 스타트업은 인간이 훈련시키는 범용 로봇 모델에 많은 투자를 하고 있다. 이 회사는 카메라가 부착된 로봇 집게와 유사한 특수 장갑을 제작하며, 사람들이 이를 사용해 다양한 작업을 수행한다. 나는 멕시코 등지의 작업자들에게 보내질 수백 개의 그리퍼가 쌓인 상자를 보았다. 플로렌스와 팀은 로봇 훈련에 정확히 어떤 방법을 사용하는지에 대해서는 함구하지만, 이미 막대한 양의 고품질 훈련 데이터를 수집했다고 말한다. 더 똑똑한 로봇을 추구하는 다른 회사들과 달리, 이들은 오픈소스 언어 모델에 의존하지 않고 AI 모델을 완전히 처음부터 직접 구축했다.

Generalist의 연구를 잘 알고 있는 조지아공과대학교의 로보틱스 연구자인 다롄 쉬(Danfei Xu)는 이 스타트업이 범용 로봇 모델을 추구하는 기업들 가운데 돋보인다고 말한다. "그들은 이 방향성을 극한까지 밀어붙이고 있습니다."

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Last week, I ventured a whopping 15 minutes from my house to see robots do some mind-boggling, jaw-dropping stuff. I visited the Cambridge, Massachusetts, offices of a startup called Generalist AI , where I watched robot arms perform simple chores like stacking cups, putting blocks into bowls, and the like. I was astonished by how quickly they figured things out—it was reminiscent of a flesh-and-blood person. The arms mastered a range of tasks after ingesting a short, instructional video and, most impressively, no specific training for a given task. One of the most striking examples involved a robot that was instructed to sweep a block into a bowl using a dustpan and brush. When the brush was removed from the scene, the robot improvised by using the dustpan like a brush and flicking the block into the bowl. In another case, a two-armed robot watched a videoclip of someone unzipping a purse before removing some banknotes. I watched—somewhat slack-jawed—as the robot unzipped a different kind of purse and carefully removed the notes. Most amazingly, when it couldn’t grab the money, it switched from using its right gripper to its left to get a better angle of attack. “Ha,” said one engineer standing nearby. “It never did that before.” “This is exactly the kind of thing people were really excited about with GPT-3,” Generalist cofounder and CEO Pete Florence told me, in reference to OpenAI’s breakthrough large language model, released in 2020. “You could take that model and just prompt it to do a new task and it would have a real shot at doing it.” Generalist appears to be focused on teaching its robots about the physics of the world, which seems inspired by the intuitive sense of physics humans exhibit from an early age. That may well contribute to the model’s ability to transfer what it has learned in one scenario to another. In fact, some of the company’s demos made me think of how children improvise and experiment when shown a task. The researchers have often been surprised by what the robot decides to do—one chose to sweep up items with a banana when it was placed in front of it, for example. This might seem trivial, but physical intelligence is something still largely lacking in machines, and the way babies learn so efficiently about their world may offer important insights for AI researchers. I met Florence and Andrew Barry, cofounder and CTO, in a conference room overlooking teams of people doing robot training with special grippers on their hands. The company’s other cofounder and chief scientist is Andy Zeng. The trio have impressive backgrounds: They previously worked at Google DeepMind and Boston Dynamics on some of the most advanced hardware and robotic models around. Traditionally, training an AI-powered robot to do different tasks has meant feeding thousands of examples into the model. This is a notoriously imperfect kind of learning, though, and a robot will struggle with the task if you change something as simple as the lighting. Generalist and some other robotics startups are investing heavily in a general robotic model trained by humans. The company builds special gloves resembling robot pincers that have cameras attached to them, which people then use to perform different chores. I saw a crate piled high with several hundred of these grippers destined for workers in Mexico and elsewhere. Florence and team are cagey about exactly what recipe they’re using to train the robots, but they say the company has already gathered a huge amount of high-quality training data. In contrast to some other companies chasing smarter robots, they have also built their AI models entirely from scratch rather than relying on an open-source language model. Danfei Xu, a roboticist at Georgia Tech who is familiar with Generalist’s work, says that the startup stands out among companies chasing more general robot models. “They have pushed this to the extreme, and they’ve done a really good job executing,” Xu says. Besides gathering a huge amount of high-quality data, he says, “they are excellent roboticists, and they have done really good science.” Xu also says that the stuff Generalist has demo’d so far suggests that they have an eye on deploying robots in real commercial settings. “They are the closest to something that's deployable,” he says. “Generalist's data approach is collecting physical interaction data at large scale without tying it too closely to one particular robot,” says Karen Liu, a roboticist at Stanford University who also knows the company. “Their strongest results suggest that this bet may be working.” That said, Generalist says the learning skills of its models are not yet all that reliable. A robot is only able to complete a task it has been shown about 59 percent of the time, on average; ideally, its success rate would be somewhere upwards of 99 percent. It also seems unclear how well these skills will generalize to every imaginable task or setting. Even so, the potential for robots to quickly learn skills in, say, manufacturing seems huge. One of Generalist’s engineers seemed to discover this late one recent evening. A video that captured the episode shows the engineer stacking small cups on the table in front of a two-armed robot, just to see what the machine might do. The robot suddenly joined in, grabbing and stacking other cups with its two grippers. As the robot finished stacking the cups into one neat pile, the engineer began yelling to no one in particular, delighted by the maneuver. This is an edition of Will Knight’s AI Lab newsletter . Read previous newsletters here.