메뉴
BL
The Decoder • 57일 전

언어 모델은 불가능, 세계 모델이 과학 혁명을 이끌 것이다

IMP
8/10
핵심 요약

구글 딥마인드의 톰 자하비(Tom Zahavy)는 언어 모델이 귀납과 연역은 잘 수행하지만, 완전히 새로운 가설을 만들어내는 '가추(Abduction)' 능력이 부족하다고 지적합니다. 특히 그는 수학적 계산이 아닌 신체적 감각을 기반으로 한 사고의 도약이 필요하며, 이는 언어 모델의 한계를 넘어 '세계 모델(World models)'만이 달성할 수 있다고 주장합니다.

번역된 본문

언어 모델은 과학 혁명을 일으킬 수 없지만, 세계 모델(World models)은 가능할 수 있습니다.

언어 모델이 과학 혁명을 일으킬 수 있을까요? 구글 딥마인드의 톰 자하비(Tom Zahavy)는 'LLMs can't jump(LLM은 도약할 수 없다)'라는 제목의 논문에서 그렇지 않다고 주장합니다. 언어 모델에는 진정으로 새로운 것을 창조하는 데 필요한 인지적 메커니즘이 빠져 있기 때문입니다.

자하비는 알베르트 아인슈타인이 친구 모리스 솔로빈(Maurice Solovine)에게 보낸 편지에 스케치한 프레임워크를 바탕으로 자신의 주장을 전개합니다. 아인슈타인은 발견이 하나의 순환이라고 썼습니다. 감각적 경험은 공리(Axioms)를 향한 직관적인 '도약(leap)'으로 이어지고, 그곳에서 논리적 연역은 테스트 가능한 결론을 도출합니다. 여기서 공리는 이론의 증명되지 않은 기초 가정을 의미합니다.

세 가지 유형의 추론 중 두 가지를 처리하는 AI

이 격차가 어디에 있는지 정확히 짚어내기 위해, 자하비는 철학자 찰스 샌더스 퍼스(Charles Sanders Peirce)의 고전적인 구분을 차용합니다. 퍼스는 모든 추론을 규칙, 사례, 결과를 연결하는 방식에 따라 분류했습니다.

연역(Deduction)은 고정된 규칙에서 보장된 결론을 도출합니다. 증명 가능한 정확한 결과를 내는 프로그램을 실행하는 것과 같습니다. 귀납(Induction)은 데이터에서 패턴을 발견합니다. 천 마리의 흰 백조를 관찰하고 모든 백조가 희다는 일반화를 이끌어내는 식입니다. 가추(Abduction)는 창조적인 도약입니다. 놀라운 현상을 설명하기 위해 원인을 창안해 내는 것입니다.

자하비는 세 번째 형태인 가추가 핵심 병목 현상이라고 보며, 이를 두 가지 수준으로 나눕니다. 보통의 가추는 알려진 후보들의 집합 중에서 가장 그럴듯한 설명을 선택하는 것입니다. 의사가 증상을 질병에 대입하는 것과 같은 방식입니다. 그는 언어 모델이 이 정도는 할 수 있다고 인정합니다.

하지만 더 어려운 버전은 그가 말하는 '조작적 가추(manipulative abduction)'입니다. 아직 언어적 틀이 존재하지 않는 원인을 발명해내는 것입니다. 그는 이것이 과학적 발명의 진짜 병목이며, 기계는 이것을 할 수 없다고 주장합니다.

논문에 따르면, 귀납과 연역은 AI가 충분히 다다를 수 있는 영역입니다. 언어 모델은 이미 통계적 패턴 인식을 훌륭하게 수행하고 있으며, 형식적 파생(formal derivation) 또한 빠르게 정복하고 있습니다. AlphaProof, Gemini, GPT-5와 같은 시스템은 이제 국제 수학 올림피아드 문제에서 금메달 수준의 점수를 기록합니다. 자하비는 심지어 언어 모델에게 시작점으로 아인슈타인의 가정이 주어진다면 일반 상대성 이론을 도출해 낼 수 있을 것이라고 인정합니다. 하지만 무엇보다 그 가정을 처음 공식화하고 그것에 도달하기 위해 조작적 도약을 만드는 것이 여전히 병목입니다.

기계가 이 도약에 어려움을 겪는 이유를 자하비는 바로 그 이론(상대성 이론)을 이용해 설명합니다. AI 모델은 일반적으로 자신의 예측을 현실과 비교하고 예측과 결과 사이의 오차를 기반으로 조정하며 학습합니다. 감지할 수 있는 오차가 없다면 시스템이 작동할 방법이 없습니다. 자하비는 이것이 아인슈타인이 직면했던 상황이라고 주장합니다.

아인슈타인이 연구할 당시, 데이터의 위기는 없었습니다. 뉴턴의 물리학은 극도로 정확하게 확인되었습니다. 알려진 유일한 이상 현상인 수성 궤도의 미세한 변화는 '불칸(Vulcan)'이라는 가상의 숨겨진 행성 때문인 것으로 돌려졌습니다. 자하비는 최적화 주도 AI에는 물리학을 뒤엎을 이유가 없었을 것이라고 주장합니다. 논증의 논리를 따르자면, 그 시대의 천문학자들처럼 시공간을 재고하기보다 작은 불일치를 설명하기 위해 행성을 하나 더 만들어냈을 것입니다. 에딩턴의 빛 휘어짐 측정과 같이 아인슈타인의 이론을 확증하는 데이터는 이론이 공식화된 지 수년이 지나서야 나왔습니다.

도약에는 몸이 필요하다 n 그렇다면 아인슈타인을 공리로 이끌었던 조작적 가추는 어디서 왔을까요? 자하비는 아인슈타인의 '가장 행복했던 생각'을 가리킵니다. 더 이상 중력을 느끼지 못하는 자유 낙하하는 관찰자입니다. 이 통찰은 체화된 시뮬레이션(embodied simulation), 즉 방정식을 푸는 것이 아니라 물리적 감각을 마음속으로 재현해 본 것에서 비롯되었습니다. 그는 우주에서 가속하는 엘리베이터 안의 물리학자를 상상했고, 가속과 중력이 내부에서는 구별할 수 없다는 결론을 내렸습니다.

자하비는 계산을 통해서가 아니라 물에 뛰어든 아르키메데스의 신체적 경험을 통해 부력의 원리를 깨달은 아르키메데스와의 평행선을 그립니다.

원문 보기
원문 보기 (영어)
Language models can't spark scientific revolutions, but world models might Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jul 30, 2026 Nano Banana Pro prompted by THE DECODER Can language models spark a scientific revolution? In a position paper titled "LLMs can't jump," Google Deepmind's Tom Zahavy argues they can't. They're missing the cognitive mechanism needed to create something truly new. Zahavy builds his case on a framework Albert Einstein sketched in a letter to his friend Maurice Solovine. Discovery, Einstein wrote, is a cycle: sensory experience leads to an intuitive "leap" toward axioms, and from there, logical deduction produces testable conclusions. Axioms are the unproven foundational assumptions of a theory. AI handles two of three types of reasoning To pinpoint where the gap lies, Zahavy draws on a classic distinction from philosopher Charles Sanders Peirce, who categorized all reasoning by how it connects rules, cases, and results. Deduction derives guaranteed conclusions from fixed rules, like running a program that produces a provably correct output. Induction spots patterns in data: observe a thousand white swans, and you generalize that all swans are white. Abduction is the creative leap. It invents a cause to explain a surprising phenomenon. This third form is where Zahavy sees the critical bottleneck, and he draws a line between two levels of it. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease. Language models can do this, he concedes. The harder version is what he calls "manipulative abduction": inventing a cause for which no linguistic template exists yet. That, he argues, is the real bottleneck of scientific invention, and machines can't do it. Induction and deduction, the paper argues, are well within reach. Language models already excel at statistical pattern recognition, and they're rapidly conquering formal derivation too. Systems like AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could probably derive general relativity if given Einstein's assumptions as a starting point. But formulating those assumptions in the first place, making the manipulative leap to reach them, remains the bottleneck. Why machines struggle with this leap, Zahavy illustrates using that very theory: AI models typically learn by comparing their predictions to reality and adjusting based on the error, the gap between prediction and outcome. Without a detectable error, there's nothing for the system to work with. And that's the situation Einstein faced, Zahavy argues. When Einstein was working, there was no data crisis. Newton's physics had been confirmed with extreme precision. The only known anomaly, a tiny shift in Mercury's orbit, had been attributed to a hypothetical hidden planet called "Vulcan." An optimization-driven AI would have had no reason to overthrow physics, Zahavy argues. Following the logic of the argument, it would have done what the astronomers of the era did: invented an extra planet to account for the small discrepancy, rather than rethinking space and time. The data confirming Einstein's theory, such as Eddington's measurement of light deflection, didn't arrive until years after the theory was formulated. A jump requires a body So where did the manipulative abduction come from that led Einstein to his axioms? Zahavy points to Einstein's "happiest thought": the freely falling observer who no longer feels gravity. This insight came from embodied simulation, Einstein mentally playing through a physical sensation rather than grinding through equations. He imagined a physicist inside an accelerating elevator in space and concluded that acceleration and gravity are indistinguishable from the inside. Zahavy draws a parallel to Archimedes, who didn't discover his buoyancy principle through calculation but, as the story goes, through the physical feeling of water rising as he stepped into a bathtub. In both cases, a foundational principle emerged that didn't yet exist in the language of the time. Language models lack exactly this sensory grounding. Zahavy compares them to philosopher John Searle's "Chinese Room," a famous thought experiment where a person shuffles Chinese characters according to a rulebook without understanding a single word. Language models shuffle the symbols of physics in much the same way, without access to the physical experience that gives those symbols meaning. Sakana's AI Scientist and Deepmind's AlphaEvolve automate scientific workflows impressively. But the AI Scientist only recombines existing concepts, while AlphaEvolve optimizes brilliantly yet needs a clear error signal it can shrink step by step. Einstein never had that signal. Neither system, Zahavy argues, can make the leap into an entirely new framework of thought. World models as a path to abduction As a possible way forward, Zahavy points to physically consistent world models. He draws a line here: video generators like Veo simply predict the most likely next frame. A falling apple falls not because the model understands gravity, but because falling is the most common continuation in the training data. That's still just pattern matching. Action-controllable world models like Genie , on the other hand, let an agent actively intervene in a simulation and run counterfactual experiments, like mentally cutting an elevator cable. A "synthetic lab" like this could provide the feedback loop needed to invent new axioms where no linguistic template exists yet. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->