메뉴
BL
MIT Tech Review • 31일 전

AI도 못 푸는 지능 테스트, 당신은 풀 수 있을까?

IMP
5/10
핵심 요약

퍼즐과 게임은 AI 발전 초기부터 핵심 평가 도구였으며, 2024년 말 18%에 그쳤던 NYT 커넥션즈 퍼즐을 최신 모델들은 거의 완벽하게 풀 정도로 능력이 빠르게 향상되고 있습니다. 그러나 공간 추론(심적 회전)과 시각 퍼즐은 여전히 치명적 약점이며, 훈련 데이터에 등장한 고전 문제의 변형(기사와 악당 퍼즐 등)에서는 암기한 답을 내놓아 실수를 반복합니다. 인간과 AI의 인지 차이를 보여주는 이런 퍼즐은 AI의 강점과 약점을 이해하는 유용한 창이 됩니다.

번역된 본문

퍼즐과 게임은 AI 개발의 처음부터 중심적인 역할을 해왔습니다. 우리 인간이 십자말풀이나 논리 퍼즐으로 자신의 두뇌를 시험하듯, 개발자들은 게임 시험대를 통해 모델이 얼마나 발전했는지 테스트할 수 있습니다. '머신러닝(machine learning)'이라는 용어는 1959년 체커 게임을 학습한 알고리즘에 관한 IBM 컴퓨터 과학자 아서 새뮤얼의 논문에서 대중화되었습니다. 체스와 중국 보드게임 바둑 역시 유명한 AI 시험대입니다.

순수하게 퍼즐 실력만 놓고 보면 AI는 매우 빠르게 발전하고 있습니다. 2024년 말 컬럼비아대학교 연구팀은 최고 수준의 모델조차 악명 높은 뉴욕타임스 '커넥션즈' 퍼즐의 18%만 풀 수 있음을 보여주었지만, 2025년 초에는 일부 모델이 거의 매번 완벽하게 풀어냈습니다.

하지만 퍼즐은 AI 역량의 멈출 수 없는 발전을 부각하는 것 이상의 의미가 있습니다. 모델이 성공하고 실패하는 지점, 그리고 우리 인간이 여전히 AI를 이기는 지점을 관찰하면 이 기술의 강점과 약점을 이해하는 유용한 창이 됩니다. 발전에도 불구하고 오늘날의 모델은 여전히 헤맵니다. 고전 수수께끼의 미묘한 변형에 자주 걸려 넘어지고, 시각 퍼즐은 특히 취약점입니다.

여기서 한때 모델들을 곤란에 빠뜨린 퍼즐으로 여러분의 두뇌를 시험해볼 기회가 있습니다. 어떤 문제는 AI에게 어려웠듯 여러분에게도 까다로울 수 있고, 다른 문제는 너무 간단해서 AI가 정말 지능이 있는지 의심하게 만들 것입니다. 각 퍼즐은 기계와 인간의 인지가 다른 방식을 최소 하나씩 보여줍니다. 테스트에서 만점을 받는다면, 적어지 않은 지금은 AI를 퍼즐로 이길 수 있음을 증명하는 셈입니다.

공간 추론 인간이 큰 우위를 가진 영역부터 시작해봅시다. 공간 추론입니다. IQ 테스트를 받아본 적이 있다면 심적 회전 문제를 풀어봤을 것입니다. 이 퍼즐은 서로 다른 이미지들이 같은 물체를 다른 각도에서 표현한 것인지 판단하라고 요구합니다. 오늘날의 언어 모델은 대체로 시각 입력을 분석할 수 있지만, 이런 퍼즐에서는 여전히 처참하게 실패합니다. 세계 모델(world model)이 AI가 물리적 환경을 이해하는 데 도움을 준다는 논의가 많지만, LLM은 여전히 건축가나 기계 엔지니어 같은 공간적 사고가처럼 3D 물체를 조작하지 못하는 듯합니다.

심적 회전 문제: 제시된 물체를 다른 각도에서 보여주는 답을 고르세요. 각 문제에는 정답이 하나뿐입니다! (원본 이미지 A B C / 원본 이미지 A B C D)

기억력과 적응력 최첨단 LLM은 놀라운 기억력을 갖고 있습니다. 훈련 중 방대한 양의 사실을 접했고 그중 많은 내용을 충실하게 암송할 수 있습니다. 이는 상식 퀴즈에서 인간을 이기는 자산이지만, 부채가 될 수도 있습니다. 퍼즐이 훈련 중 본 문제와 매우 비슷하면 모델이 핵심 차이를 놓치고 암기한 내용으로 답할 수 있기 때문입니다.

이는 2024년 구글과 일리노이대학교 어배너-섐페인 캠퍼스 연구진이 '기사와 악당(Knights and Knaves)'이라는 고전 퍼즐 유형의 살짝 변형된 버전으로 모델을 훈련·테스트한 연구에서 사실로 확인됐습니다. 이 문제에서는 일부 등장인물은 항상 진실만 말하고 다른 이들은 항상 거짓만 말하며, 누가 누구인지 추리해야 합니다.

같은 원리가 'SimpleBench'라는 테스트에서도 작동할 수 있습니다. 이 문제들은 모델이 훈련 중 접했을 법한 더 복잡한 문제들과 닮아 있습니다. 인간은 함정을 알아채지만 최고 수준 모델조차 넘어집니다.

기사와 악당 문제: 풀기 위해 알아야 할 것은 단 하나, 기사(knight)는 항상 진실을 말하고 악당(knave)은 항상 거짓을 말한다는 것입니다. 각 인물의 말을 바탕으로 누가 기사이고 누가 악당인지 판단하세요.

  1. 섬 주민 두 명을 만났습니다. 이름은 에드워드와 월러스입니다. 월러스: 에드워드는 진실을 말한다. 에드워드: 월러스와 나는 같은 유형이다.

  2. 섬 주민 세 명을 만났습니다. 이름은 조지프, 프랜신, 앨리스입니다. 프랜신: 조지프는 악당이다. 프랜신: 앨리스는 진실을 말한다. 앨리스: 조지프는 나와 같은 유형이 아니다.

  3. 섬 주민 세 명을 만났습니다. 이름은 로버트, 빈센트, 미셸입니다. 미셸: (본문 생략)

원문 보기
원문 보기 (영어)
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too. Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time. But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot. Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now. Spatial Reasoning Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can. Mental Rotation Instructions : Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer! Original A B C Original A B C D Memory & Adaptability Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized. This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip. Knights and Knaves Instructions : The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says. 1 You have met a group of two islanders. Their names are Edward and Wallace. Wallace says: Edward tells the truth. Edward says: Wallace and I are the same type. 2 You have met a group of three islanders. Their names are Joseph, Francine, and Alice. Francine says: Joseph is a knave. Francine says: Alice tells the truth. Alice says: Joseph is not my type. 3 You have met a group of three islanders. Their names are Robert, Vincent, and Michelle. Michelle says: Robert always lies. Vincent says: Michelle is truthful. Robert says: Vincent is untruthful. Robert says: Vincent is not my type. SimpleBench Instructions : Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time. 1 Beth places four whole ice cubes in a frying pan at the start of the first minute, then five at the start of the second minute and some more at the start of the third minute, but none in the fourth minute. If the average number of ice cubes per minute placed in the pan while it was frying a crispy egg was five, how many whole ice cubes can be found in the pan at the end of the third minute? A 30 B 0 C 20 D 10 E 11 F 5 2 A juggler throws a solid blue ball a meter in the air and then a solid purple ball (of the same size) two meters in the air. She then climbs to the top of a tall ladder carefully, balancing a yellow balloon on her head. Where is the purple ball most likely now, in relation to the blue ball? A At the same height as the blue ball B At the same height as the yellow balloon C Inside the blue ball D Above the yellow balloon E Below the blue ball F Above the blue ball Abstract & Visual Reasoning AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous ­puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell. Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-­generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them. ARC-AGI Instruction s: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.) Pair 1 Pair 2 Pair 3 Now you try it Puzzle Your answer Intuition It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI does not share. Psychologists have designed problem suites that invert the SimpleBench phenomenon: For these questions, humans often give knee-jerk answers, whereas models will respond deliberatively. Some of the problems exploit errors in the ways that we intuitively do math; others are phrased so as to suggest obvious answers that fall apart if the question is read carefully. Lightning Round Instructions : Answer the questions below as quickly as you can. 1 In a cave, there is a colony of bats whose population doubles each day. Given that it takes 60 days for the entire cave to be filled with bats, how many days would it take for the cave to be half-filled with bats? 2 In what famous novel does Alice state "I'm late, I'm late, for a very important date"? Increasing Complexity In some cases, whether an LLM can complete a puzzle is a matter of scale. One study from researchers at Apple found that LLMs can ace simple versions of the Tower of Hanoi problem, which involves moving a stack of disks one at a time without ever putting a larger disk atop a smaller one, and river-crossing puzzles, in which a group of people must traverse a river according to certain rules. But only up to a point: As the number of disks or people hits six and higher, the models began to falter. In another study, researchers at the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle similarly with logic grid puzzles, which require deducing the attributes of a set of individuals from a list of clues. The Apple paper went vir