메뉴
BL
Wired AI 13일 전

AI보다 똑똑한 아기, 그 비밀을 찾아서

IMP
8/10
핵심 요약

최첨단 AI 모델조차 아기가 세상을 인식하고 학습하는 방식에는 미치지 못한다는 연구 결과가 나왔습니다. 스탠퍼드 대학교, 메타(Meta), 도쿄 대학교 등의 연구진은 아기의 시점에서 촬영된 영상으로 AI를 테스트하는 'EgoBabyVLM' 챌린지를 통해 이를 입증했습니다. 엄청난 데이터에 의존하는 현재 AI의 한계를 극복하고 물리적 환경을 자연스럽게 학습하는 로봇 등을 개발하기 위해 아기 뇌의 학습 방식을 연구하는 것이 매우 중요합니다.

번역된 본문

수천 개의 최첨단 컴퓨터 칩으로 구동되는 인공지능(AI) 모델이 똑똑하다고 생각된다면, 1세 아기의 개념을 소개해 드리겠습니다. 물론 아기들은 컴퓨터 프로그램을 작성하거나 고급 수학 문제를 풀거나 철학적 아이디어를 토론할 수는 없습니다. 하지만 바다와 같은 엄청난 양의 학습 데이터를 소비하고 소규모 국가 하나가 쓰는 수준의 에너지를 소모하는 오늘날의 AI 모델과 달리, 아기들은 놀라운 효율성으로 세상을 이해하는 법을 배웁니다. 아기들은 새로운 사물을 한두 번만 보고도 식별해 내며, 스치듯 보는 관찰과 물리적 상호작용을 통해 학습합니다.

AI를 개선한다는 측면에서 볼 때, 아기들, 그리고 그들의 뇌 구조는 핵심적인 실마리를 쥐고 있을지 모릅니다. 아기의 학습 방식과 더 비슷한 형태의 AI를 구축한다면 최첨단 모델의 비용과 에너지 소모를 줄일 수 있을 것입니다. 나아가 AI 기반 로봇이 자신의 주변 환경에 대해 더 자연스러운 방식으로 학습하게 만드는 데에도 이 방식이 가치를 발휘할 수 있습니다.

이 대담하고 새로운 지평을 탐구하기 위해 메타, 스탠퍼드 대학교, 도쿄 대학교, 그리고 프랑스 고등사범학교(École Normale Supérieure)의 연구진은 아기의 학습 능력을 조명하고 AI 연구자들이 이에 필적하는 알고리즘을 설계하도록 유도하는 새로운 테스트를 개발했습니다. 'EgoBabyVLM 챌린지'는 텍스트와 이미지 모두를 기반으로 학습하는 시각 언어 모델(VLM)이 아기의 시선으로 세상을 얼마나 잘 이해할 수 있는지 평가합니다. 이 테스트는 영유아의 머리에 카메라를 착여시켜 수집한 약 1,000시간 분량의 영상 데이터를 모델에 학습시킨 후, 그 모델이 세상을 어떻게 묘사하는지를 요구합니다. (네, 정말입니다.)

결과적으로 최첨단 모델들은 이 현실적이고 지저분한 영상을 제공받았을 때 끔찍하게 실패하는 것으로 나타났습니다. 이는 아기의 뇌 구조에는 매우 적은 정보만으로도 빠르게 학습할 수 있게 해주는 무언가 다른 설계가 포함되어 있음을 시사합니다. 아기들은 정제된 데이터셋을 통해 배우는 것이 아니라 만화경처럼 다채로운 시각에서 학습합니다. 예를 들어, 더 이상 보이지 않는 사물에 대해 이야기하는 부모, 시선이나 몸짓을 사용해 무언가를 가리키는 행동, 또는 지금 당장 일어나는 일이 아닌 과거의 사건이나 미래의 일에 대해 논의하는 상황 등이 여기에 해당합니다.

EgoBabyVLM의 개발에 참여했으며 언어 학습을 전공하는 스탠퍼드 대학교의 인지과학자 마이클 프랭크(Michael Frank)는 아기들이 단순히 언어뿐만 아니라 풍부한 멀티모달 및 촉각적 경험을 통해서도 학습한다고 말합니다. 이 테스트는 AI에 관한 한 "[단순한 언어 외에] 더 많은 것이 필요하다"는 사실을 명확히 보여준다고 프랭크는 덧붙였습니다.

언어 학습 EgoBabyVLM은 과학자들이 인간의 지능을 탐구하기 위해 AI를 활용하는 최신 사례 중 하나에 불과합니다. 2023년에 도입된 'BabyLM'이라는 챌린지는 AI 모델에게 AI 모델이 쓰는 수조 개의 단어가 아닌, 10살 아이가 접하는 정도의 데이터인 수천만 단어를 사용하여 언어의 통사론을 학습하는 과제를 부여했습니다. 놀랍게도 여러 문장에 걸쳐 단어 간의 관계에 주의를 기울여 언어를 처리하는 트랜스포머(Transformer) 기반 AI 모델이 이를 매우 잘 수행할 수 있다는 사실이 밝혀졌습니다. 이는 통사론이 인간의 뇌에 어떻게 하드와이어(Hardwired)되어 있는지에 관한 노엄 촘스키(Noam Chomsky)의 이론에 반하는 발견입니다.

BabyLM을 처음 개발한 취리히 연방 공과대학교(ETH Zurich)의 언어학자 라이언 코터렐(Ryan Cotterell)은 물리적 세계를 이해하는 데 있어서는 상황이 다르다고 말합니다. 그는 "인간의 상호작용에 대한 대규모 코퍼스는 존재하지 않을 것입니다. 인간의 상호작용이 담긴 인터넷 같은 공간은 없으니까요."라고 말합니다.

매사추세츠 공과대학교(MIT)의 인지과학자 조슈아 테넨바움(Joshua Tenenbaum)은 BabyLM 연구가 모델들이 물리적 세계, 사회적 역학, 또는 마음 이론에 대한 '상식'을 습득하지 못한다는 것을 보여주었다고 지적합니다. 테넨바움은 "트랜스포머는 데이터에서 패턴을 찾는 데 매우 뛰어납니다."라고 말합니다. "하지만 순수하게 패턴을 학습하는 시스템만으로는 아기나 어린아이가 받아들이는 종류의 데이터를 처리하여 그들이 배우는 모든 것을 학습할 수는 없는 것 같습니다."

결론적으로 남은 핵심적인 질문은 진화의 과정이 인간과 다른 동물의 특정 학습 기술을 최적화하는 방법을 찾아냈는지의 여부, 혹은 이것을 단순화할 수 있는지에 대한 것입니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story If you think an artificial intelligence model running on thousands of cutting-edge computer chips is smart, allow me to introduce you to the concept of a 1-year-old. OK, so babies might not be able to write computer programs, solve advanced math problems, or debate philosophical ideas. But unlike today’s AI models, which consume an ocean’s worth of training data and as much energy as a small country , babies learn to make sense of the world with amazing efficiency. They identify new objects after seeing them once or twice, and they learn through fleeting observation and physical interaction. When it comes to improving AI, babies—and the architecture of their brains—might hold crucial insights. Building a more baby-like version of AI could make frontier models less costly and less energy intensive, and it might also be valuable if AI-powered robots are to learn about their environments in a more natural way. To explore this bold new frontier, researchers at Meta, Stanford University, the University of Tokyo, and France’s École Normale Supérieure developed a new test that highlights the learning skills of babies and pushes AI researchers to design algorithms that match them. The EgoBabyVLM Challenge judges how well vision language models, or VLMs, which learn from both text and imagery, can make sense of the world as a baby sees it. It requires a model to describe the world after ingesting about a thousand hours of video collected from cameras strapped to the heads of infants and toddlers. (Yes, really.) It turns out that the cutting-edge models fail miserably when fed this realistic and messy footage, which suggests there may be something different about the design of the baby brain that enables it to learn so rapidly from so little information. Instead of curated datasets, babies learn from a kaleidoscopic view of things: parents talking about objects that are no longer visible, indicating things using their gaze or a gesture, or discussing events from the past or in the future rather than whatever’s happening right then. Babies learn not just from language but also from a rich multimodal and tactile experience, says Michael Frank, a cognitive scientist at Stanford University who specializes in language learning and was involved with EgoBabyVLM’s development. The test shows that when it comes to AI, “it’s clear that there’s more [than just language] that’s needed,” Frank says. Language Learning EgoBabyVLM is just the latest example of how scientists are using AI to explore human intelligence. A challenge called BabyLM , introduced in 2023, tasked AI models with learning the syntax of language using about the same amount of data a 10-year-old takes in—tens of millions of words, compared to trillions for AI models. Remarkably, it turns out that transformer-based AI models—which process language by paying attention to the relationship between words across different sentences—can do this quite well, a finding that challenges Noam Chomsky’s ideas concerning how syntax may be hardwired into the human brain. Ryan Cotterell, a linguist at ETH Zurich who first developed BabyLM, says the situation is different when it comes to understanding the physical world. “There isn't going to be a large corpus of human interactions—there's no internet of human interactions,” he says. Joshua Tenenbaum, a cognitive scientist at the Massachusetts Institute of Technology, notes that BabyLM showed models do not acquire “common sense” about the physical world, social dynamics, or theory of mind. “Transformers are very good at finding patterns in data,” says Tenenbaum. “But it does seem that just pure pattern learning systems are not able to take the kind of data that a baby or a child receives and learn all the things that they do.” An enduring question is whether evolution found a way to optimize certain learning skills in humans and other animals, or if simple learning algorithms can do everything we do. “There is a lot of debate in cognitive science and neuroscience about how much is built into the brain evolutionarily,” Tenenbaum says. “The brain is incredibly complex, and there's a lot of built-in structure and architecture.” In 2024, researchers showed that a basic VLM can learn simple things, like what a ball is, purely by consuming data recorded from the head of a single infant. But this is a ways away from reasoning about the world in sophisticated ways. “The mystery is how children get to the full capabilities that they have even at the age of 2,” says Brendan Lake, a cognitive scientist at Princeton University who was involved with the project. The authors of the EgoBabyVLM paper suggest that borrowing different ideas from cognitive science and neuroscience could enable progress toward more humanlike learning algorithms. This includes designing models that can pay attention over longer periods and can interpret social cues. Stanford’s Frank has already shown that novel approaches can get us closer to baby-like AI. Earlier this year, he and colleagues tested a new kind of model that’s adept at learning causality and visual and temporal relationships—or how objects affect one another over time—using the same baby-head video data. They found the new model was able to learn about the dynamics of different objects, a foundation for physical reasoning, much more effectively. It’s a tantalizing possibility: Perhaps models that are biased to learn more rapidly about things like physics and social relationships could be more efficient learners overall. “EgoBabyVLM is a wonderful challenge,” says Lake. “I'm excited to see what kinds of new architectures, approaches, and ingredients researchers come up with.” This is an edition of Will Knight’s AI Lab newsletter . Read previous newsletters here.