메뉴
BL
MIT Tech Review • 33일 전

아이는 AI보다 적은 데이터로 언어를 배운다—그 이유는 아직 미스터리

IMP
7/10
핵심 요약

LLM은 인간 아이보다 최대 10만 배 많은 언어 데이터를 소비해야 유창한 언어 구사가 가능한데, 이 '데이터 효율성 격차'는 인지과학과 AI 연구 모두에 중요한 난제다. 인터넷 데이터가 2030년대에 고갈될 수 있다는 전망 속에서 아이의 학습 방식을 역설계해 데이터 효율적인 AI 모델을 만들려는 연구가 주목받고 있다.

번역된 본문

우리가 판단할 수 있는 한, 인류는 최소 10만 년 전부터 서로 대화해왔습니다. 그리고 그 긴 세월 동안 인간의 언어를 완벽한 유창함으로 습득할 수 있는 존재는 세상에 단 하나뿐이었습니다. 바로 인간 아이였습니다. 이제 둘이 되었습니다. ChatGPT 출시 후 불과 4년 만에, 우리 중 많은 이들이 휴대폰이나 컴퓨터와 자연스럽게 대화하는 것을 당연하게 여기게 되었습니다. Claude, DeepSeek, OpenAI의 GPT 모델 같은 LLM(대규모 언어 모델)은 인간으로서 믿을 만할 정도로 위장할 수 있을 만큼 유창하고 유연합니다.

하지만 계산의 장막 뒤를 들여다보면 함정이 있습니다. 컴퓨터에게 인간 언어를 가르치는 것은 여전히 비인간적인 양의 데이터를 필요로 합니다. LLM은 모국어를 마스터하는 과정에서 한 사람이 경험하는 단어보다 10만 배 이상 쉽게 소화해낼 수 있으며, 아이들이 보통 언어를 잡기 시작하는 첫 번째 생일까지 듣게 되는 양보다 훨씬 더 많은 단어를 처리합니다.

스탠퍼드대 인지과학자 마이클 C. 프랭크는 LLM에 대해 "최근의 발전은 놀라웠다"며 말합니다. "하지만 우리는 거실에서 일 년 동안 일어나는 이정표를 재현하기 위해 숲 한 개를 태우고 인류 지식의 총체를 긁어모아야 합니다."

아이와 기계 사이의 이 벌어진 격차는 '데이터 효율성 격차(data efficiency gap)'라고 불립니다. 그리고 이것은 인지과학자들에게 매혹적인 질문을, AI 모델 설계자들에게 도전 과제를 던집니다. 어떻게 아이들이 지금까지 만들어진 가장 언어적으로 정교한 기계보다 여전히 더 나은 성능을 낼 수 있을까요? 답을 찾는 일은 AI 연구와 인지과학 모두에 중요한 의미를 갖습니다.

지난 10년간 언어 모델은 대부분 더 커짐으로써 더 좋아졌습니다. 2년 전 출시된 Meta의 개방형 가중치 LLM 라마 3.1(Llama 3.1)은 사전학습(pretraining)에서 15조 토큰(단어와 비슷한 언어 조각)을 소화했습니다. 사전학습은 챗봇처럼 특정 작업을 위해 미세조정되기 전에 이루어지는 모델 학습의 주요 단계입니다. 조지타운대의 인지과학자이자 언어학자인 에선 고틀리브 윌콕스는 최첨단 프론티어 모델들은 10배 더 많은 데이터로 사전학습될 수 있다고 말합니다. 하지만 학습에 사용할 수 있는 인터넷은 한계가 있으며, 어쩌면 2030년대에라도 쉽게 구할 수 있는 데이터의 우물은 말라버릴 수 있습니다.

아이들은 더 적은 것으로 더 많이 배울 수 있음을 보여줍니다. 훨씬 더 적은 것으로요. 언어적으로 풍부한 가정에서 자란 10대 초반 아이는 대략 1억 단어 정도를 들었을 수 있습니다. 문해력을 더하면 20세까지 어쩌면 3억 단어까지 늘어날 수 있습니다. 이 규모의 차이는 비유로만 짐작할 수 있습니다. 윌콕스는 "Claude는 한 도시 전체가 한 세대에 경험할 언어의 양을 접했다"고 말합니다. 현대 LLM을 학습시키는 데 사용된 모든 단어를 종이에 인쇄한다면 국제우주정거장을 넘어서는 높이의 종이탑을 만들 수 있습니다. 반면 인간 10대의 1억 단어는 겨우 20미터 높이에 불과합니다.

그리고 우리는 그보다 훨씬 적은 양으로도 해낼 수 있습니다. 과학자들은 아이들이 학습하는 방식을 역설계함으로써 더 데이터 효율적인 AI 모델을 만들 수 있기를 희망합니다. 이는 비디오로 AI를 효과적으로 학습시키는 것부터 소수 언어 공동체를 위한 챗봇을 만드는 것까지 모든 분야에 유용할 수 있습니다. 인간 학습에 대한 가설을 기계 모델에서 검증하는 것은 언어와 아이의 발달하는 마음에 관한 오랜 질문에 결론을 내릴 수도 있습니다. 우리는 언어 본능을 타고나는 걸까요, 아니면 원칙적으로 아이가 순수하게 경험만으로 언어를 배우는 것도 가능할까요? 우리가 언어를 처리하는 방식은 우리 생물학의 특이성일까요, 아니면 최소한 그 일부는 언어가 사용되고 학습될 수 있는 방식에 대한 보편적 제약을 반영하는 걸까요?

핵심 요소

우리 대부분은 어린 시절 이후에 새로운 언어를 배우려 할 때 비로소 언어가 어렵다는 것을 깨닫습니다. 과거완료시제, 혀를 굴리는 r 발음과 비모음, 속격, 구동사, 문법적으로 남성인 테이블과 여성인 숟가락—어른 언어 학습자를 괴롭히는 언어적 고문 도구는 무궁무진합니다. 하지만 아이에게 배우는 것은 대개 아무런 노력 없이 이루어집니다.

원문 보기
원문 보기 (영어)
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two. Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers. LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language. “The progress recently has been amazing,” Michael C. Frank, a cognitive scientist at Stanford University, says of LLMs. “But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.” This yawning divide between children and machines is called the data efficiency gap. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built? Finding answers has stakes for both AI research and cognitive science. For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry. Kids show that it could be possible to learn more with less. Far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20. The difference in scale is something that can only really be gestured at in analogy. “Claude has seen the amount of language that an entire city will experience in one generation,” says Wilcox. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station. The human preteen’s 100 million words, meanwhile, would stack up just 20 meters. And we can make do with far less than that. By reverse-engineering the way kids learn, scientists hope to be able to create more data-efficient AI models, which could be useful for everything from training AI effectively on video to creating chatbots that serve minority language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and children’s developing minds. Are we born with a language instinct, or would it be possible, even in principle, for a child to learn language purely from experience? Is the way we process language a quirk of our biology, or might at least some of it reflect universal constraints on how languages can be used and learned? The essential elements Most of us realize language is hard only when we try to learn a new one after childhood. The past perfect tense, rolled r s and nasal vowels, the genitive case, phrasal verbs, grammatically masculine tables and feminine spoons—many are the instruments of linguistic torment for the adult language learner. It’s typically effortless to learn our mother tongues, however. Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end. “It’s just totally miraculous,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” Exactly how babies pull this off is a mystery. Researchers know a lot about what kids learn and how they use language at different stages in development, but there’s still a lot we don’t know. Perhaps the most enduring question is why babies can learn language at all. The syntax of human language—the rules for combining words into sentences—includes recursive, nested structures that allow us to express virtually infinite ideas with a finite lexicon of words and pieces of words. This seems like something that should be a problem for babies. They only splash about in the shallows of a fathomless ocean of language. And yet, somehow, that’s enough. From a drop, they infer the depths. One solution, put forward in the 1950s by the MIT linguist Noam Chomsky, is that babies are born with hardwired knowledge of grammar. Chomsky was reacting to a rival view, championed by the psychologist B.F. Skinner, that language acquisition is entirely environmental. Skinner thought language was learned through conditioning and reinforcement, the way a dog figures out how to sit or shake for treats. Chomsky countered by citing the “poverty of the stimulus”—the idea that language, especially syntax, is too complex and children’s exposure to it too “impoverished” for them to learn entirely from experience. “His signature argument was, essentially, that language cannot be learned on the basis purely of statistics,” says Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Instead, Chomsky posited that language is based on a set of logical rules and argued that children needed innate knowledge of those rules to deduce the grammar of their language from scraps of speech. “It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” Michael C. Frank, cognitive scientist, Stanford University The Chomskyan view of language dominated linguistics in the US for decades under the moniker of generative grammar. And it was a major influence on computer science in the 1950s and ’60s, when AI was enjoying its first boom time and the lines between linguistics and natural-language processing dissolved in a flood of military funding; the Pentagon wanted computers that could understand English and translate Russian. Despite early successes of simple neural networks, which learn to recognize and reproduce statistical patterns, AI researchers in the United States largely adopted a rule-based framework influenced by Chomsky’s theories. They tried to teach language to computers by explicitly coding the rules into programs—think less immersion experience, more grammar class. This approach, part of a broader trend called symbolic AI, prevailed for decades. It also largely failed to produce models actually capable of handling human language at scale. Interest in natural-­language processing chilled in the “AI winter” that began in the 1970s. In the aftermath, neural networks started to make a comeback. But it wasn’t until the 2010s, when computer hardware was getting cheap and capable and the internet was getting big, that their performance began turning heads. By 2018 and 2019, the models BERT and GPT-2, which were built on a new architecture—the transformer—and trained on billions of tokens, made it clear to insiders that learning from a massive glut of data could work for language. In 2022, with the breakout success of OpenAI’s chatbot ChatGPT, it was clear to everyone. LLMs are not brains. What they are is powerful statistical learners—naïve pattern-learning machines without any of the evolved biological quirks folded into the human cortex. In other words, they are ex