메뉴
BL
The Decoder 57일 전

튜링상 수상자 리처드 서튼: 순수 생성 AI는 진짜 과학을 할 수 없다

IMP
8/10
핵심 요약

튜링상 수상자 리처드 서튼은 일반적인 생성 AI가 자체적인 '평가' 능력이 부재하여 진정한 과학적 발견을 이뤄낼 수 없다고 주장합니다. 모방이나 무작위 생성에 그치는 순수 생성 모델과 달리, 알파고나 알파폴드, 코딩 에이전트 등은 명확한 피드백과 평가 루프를 갖추고 있어 진정한 창의성과 발견이 가능하다고 설명했습니다.

번역된 본문

튜링상 수상자 리처드 서튼(Richard Sutton)은 일반적인 생성 AI가 과학적 발견에 필수적인 핵심 능력, 즉 자신의 결과를 평가하고 발전시키는 능력이 부족하다고 주장합니다.

대규모 언어 모델, 이미지 생성기, 비디오 모델은 방대한 예시 데이터를 학습하고 이와 유사한 결과물을 만들어냅니다. 서튼에 따르면, 이러한 결과물이 우수할 때는 대개 모델이 학습한 원천 자료인 텍스트, 이미지 또는 데이터의 덕분입니다. 결과물이 진정으로 새로울 때는 그 원천 자료를 넘어서는 것입니다. 사실에 관한 질의의 경우, 이를 소위 '환각(Hallucination)'이라고 부릅니다.

서튼은 오래된 연구자들의 농담을 인용하며 자신의 비판을 설명합니다. "이 연구는 새롭기도 하고 훌륭하기도 합니다. 안타깝게도 훌륭한 부분은 새롭지 않고, 새로운 부분은 훌륭하지 않습니다." 서튼은 이 진단이 오늘날 생성 AI의 상당 부분에 들어맞는다고 말합니다. 생성 AI는 유용한 것을 모방하거나 무작위로 새로운 것을 만들어낼 수는 있지만, 스스로 어떤 새로운 아이디어가 진정으로 좋은 것인지 판단할 수는 없습니다.

서튼은 생성 AI가 요약, 연구, 비서 또는 엔터테인먼트 분야에서 유용할 수 있다는 사실을 부정하지 않습니다. 새로움(novelty)은 애초에 목표가 아닌 경우가 많습니다. 요약본은 새로운 사실을 지어내서는 안 되며, 연구 결과에 추가적인 주장을 슬쩍 끼워 넣어서도 안 됩니다. 서튼은 "생성 AI는 단지 모방하더라도, 모방 대상보다 더 빠르거나, 더 저렴하거나, 더 작거나, 더 맞춤화 가능하거나, 더 쉽게 복사될 수 있다면 매우 유용할 수 있다"고 말합니다.

과학에서 모방의 한계 서튼의 관점에서 이러한 한계는 과학 전반에 있어 가장 중요하게 작용합니다. 과학의 핵심 목적은 이미 알려진 것을 재현하는 것이 아니라 새로운 것을 발견하고 이를 테스트하여 지속 가능한 지식으로 전환하는 데 있기 때문입니다.

서튼은 진정한 발견을 '변이(Variation)', '평가(Evaluation)', '선택적 유지(Selective retention)'의 세 단계 과정으로 설명합니다. 시스템은 다양한 옵션을 생성하고 이를 테스트한 뒤, 효과가 입증된 접근 방식을 계속 사용해야 합니다. 서튼은 이 원리가 진화, 과학적 방법, 계획, 탐색 및 강화학습(Reinforcement Learning)에 모두 존재한다고 말합니다.

순수 생성 AI가 가장 결여된 부분은 바로 '평가' 능력입니다. 언어 및 이미지 모델은 다양한 변형을 생성해 냅니다. 하지만 테스트가 없다면 최적의 결과를 선택할 수 없고, 발견도 이뤄내지 못합니다. 서튼은 "새로움은 순간적으로 생겨났다가도 그 가치를 인정받지 못하면 곧 사라져 버린다"고 말합니다.

평가는 인간으로부터 올 수 있습니다. 예를 들어 사용자가 AI가 생성한 여러 이미지 중 최고의 결과를 직접 고르는 경우입니다. 하지만 평가는 명확한 목표에서도 비롯될 수 있습니다. 체스나 바둑의 체크메이트, 형식적으로 유효한 수학적 증명, 프로그램의 성공적인 실행, 시뮬레이션 환경에서의 높은 보상 등이 그 예입니다. 오직 이러한 종류의 피드백만이 단순한 생성을 탐색 및 발견의 과정으로 바꿔놓을 수 있습니다.

알파고, 알파폴드, 클로드 코드(Claude Code)가 보여주는 차이점 서튼은 순수 생성 AI의 한계를 넘어선 일부 AI 시스템이 이미 "진정한 창의성과 진정한 발견을 할 수 있다"고 말합니다. 그는 famously 기발한 수로 알려진 알파고(AlphaGo)의 37수, 알파제로(AlphaZero)의 독창적인 체스 스타일, 단백질 구조 예측의 알파폴드(AlphaFold), 수학 분야의 알파프루프(AlphaProof), 프로그래밍 분야의 클로드 코드(Claude Code), 시뮬레이션 레이싱의 GT-Sophy 등을 그 예로 듭니다.

이러한 시스템들의 공통점은 순수한 텍스트나 이미지 생성을 넘어선 '평가 루프(Evaluation loop)'를 갖추고 있다는 것입니다. 바둑의 수는 승리할 확률을 높이거나 그렇지 않거나 둘 중 하나입니다. 수학적 단계는 형식적으로 검증되거나 그렇지 않습니다. 코드는 테스트를 통과하고 올바르게 실행되거나 실패합니다. 이를 통해 더 나은 솔루션을 선택하고 추구하는 것이 가능해집니다.

서튼은 "이 모든 시스템들은 진정한 창의성과 발견을 가능하게 하는 몇 가지 추가적인 기능을 갖추고 있다"고 밝혔습니다.

서튼의 비판은 명시적으로 런타임 시 자체 출력을 평가하지 않는 '일반적인(ordinary)' 생성 AI를 겨냥합니다. 검색, 검증자(Verifiers), 도구, 강화학습 또는 형식적 검증자(Formal validators)가 확장된 언어 모델은 진정한 발견 시스템의 일부가 될 수 있습니다. 하지만 이러한 구조가 프로그래밍, 게임 및 명확하게 테스트 가능한 작업을 넘어 어디까지 확장될 수 있는지는 여전히 미해결 과제로 남아있습니다.

원문 보기
원문 보기 (영어)
Turing Award winner Richard Sutton says pure generative AI can't do real science Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 1, 2026 Screenshot via YouTube Turing Award winner Richard Sutton argues that ordinary generative AI lacks a key ability for scientific discovery: it can't evaluate and develop its own results. Large language models, image generators, and video models learn from massive amounts of examples and produce outputs that resemble them. According to Sutton, when these outputs are good, it's usually thanks to the source material: the texts, images, or data the model learned from. When the outputs are truly novel, they go beyond that material. For factual queries, that's called hallucination. Sutton illustrates his critique with an old researcher's joke: "This work is both novel and good. Unfortunately, the parts that are good are not novel, and the parts that are novel are not good." That diagnosis fits large parts of today's generative AI, Sutton says. It can mimic useful things or randomly produce new things, but it can't tell on its own which new ideas are actually good. Sutton doesn't deny that generative AI can be useful for summaries, research, assistants, or entertainment. Novelty often isn't even the goal: a summary shouldn't invent new facts, and research shouldn't sneak in extra claims. "Generative AI can be extremely useful, even when it just mimics, if it is faster, or cheaper, or smaller, or more customizable, or more copy-able, than the thing being mimicked," Sutton says. Imitation falls short for science In Sutton's view, this boundary matters most for science in general, where the point isn't to reproduce what's already known but to discover new things, test them, and turn them into lasting knowledge. Sutton describes genuine discovery as a three-step process: variation, evaluation, and selective retention. A system has to generate different options, test them, and keep using the approaches that work. Sutton says this principle exists in evolution, in the scientific method, in planning, in search, and in reinforcement learning. What pure generative AI lacks most is evaluation. Language and image models do generate different variants. But without testing, there's no selection of the best and no discovery. "The novelty flickers into existence, but if its value is unrecognized, it flickers away and is lost," Sutton says. Evaluation can come from humans, for example, when users pick the best image from several AI-generated options. But it can also come from a clear goal: a checkmate, a formally valid proof, a successful program run, or a high reward in a simulated environment. Only that kind of feedback turns mere generation into a search and discovery process. AlphaGo, AlphaFold, and Claude Code show the difference Sutton says some AI systems that go beyond pure generative AI are already "capable of true creativity and true discovery." He points to examples like AlphaGo with its famous move 37 , AlphaZero with its unique chess style , AlphaFold in protein structure prediction, AlphaProof in math, Claude Code in programming , and GT-Sophy in simulated racing . What these systems share is an evaluation loop that goes beyond pure text or image generation. A Go move either raises the chance of winning or it doesn't. A math step can be formally checked, or it can't. Code passes tests, runs correctly, or fails. This makes it possible to select and pursue better solutions. "All these systems have some additional features that make them capable of true creativity and true discovery," Sutton says. Sutton's critique explicitly targets "ordinary" generative AI: models that don't evaluate their own output at runtime. Language models extended with search, verifiers, tools, reinforcement learning, or formal validators can become part of genuine discovery systems. But how far that structure can stretch beyond programming, games, and clearly testable tasks remains an open question. Sutton sees another issue in how neural networks are trained. Standard networks start with random settings and then learn from data. That initial randomness is a source of variation, but it mostly happens at the beginning. Over time, models can lose their ability to learn as their internal structures get rigid. A truly learning system shouldn't just be trained once, Sutton argues. It would need to renew its structure on an ongoing basis: try new possibilities, keep what works, and discard what doesn't. His goal is an AI that manages variation, evaluation, and selective retention on its own over long stretches of time. "Let's fully automate Creativity and Discovery!" he says. Sutton has been critical of the AI industry's direction for a while Sutton recently criticized the AI industry more broadly, saying it has "lost its way." The researcher is mainly pushing back against the heavy focus on ever-larger language models that absorb vast knowledge during training but don't learn from their own experience over time. Instead, Sutton calls for AI agents that interact with their environment continuously, learn from it, build internal models of the world, and plan new strategies. Meta-learning also factors into his vision: systems should learn how to learn better instead of just mimicking individual tasks. In his Oak architecture, Sutton lays out a possible path to powerful AI systems. The core idea is that agents start with no built-in specialist knowledge, act in an environment, get feedback, and form increasingly abstract concepts over time. Useful concepts become the foundation for the next stage of learning. The big open prerequisite for this, Sutton says, is reliable continual learning . Today's neural networks often struggle to absorb new knowledge without overwriting old knowledge or losing the ability to adapt. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->