메뉴
BL
MIT Tech Review • 39일 전

AI의 재귀적 자기 개선, 생각보다 빨리 오지 않을 수도

IMP
8/10
핵심 요약

프린스턴대 연구진이 개발한 '섀도우 평가' 실험에서 AI 에이전트는 연구에 필요한 엔지니어링 작업은 수행했지만, 창의성과 판단력이 부족해 최상위 AI 학회 수준의 논문을 만들어내지 못했습니다. 이는 AI가 스스로를 개선하는 '재귀적 자기 개선'에 대한 낙관적 전망이 실제 증거를 앞서고 있음을 시사합니다.

번역된 본문

AI 업계가 현재 내걸고 있는 가장 대담한 약속은 AI가 곧 거의 인간의 감독 없이 스스로를 개선할 것이라는 것입니다. 대규모 언어 모델(LLM)은 이미 코드를 작성하고, 학습용 합성 데이터를 생성하며, 자신이 구동되는 컴퓨터 칩을 최적화할 수 있습니다. 폭발적인 AI 발전 전망들은 연구자들이 '재귀적 자기 개선(recursive self-improvement)'이라 부르는 것이 눈앞에 다가왔다고 예측합니다. 그러나 새로운 연구에 따르면 그 지점에 도달하는 데는 상당한 시간이 걸릴 수 있습니다.

프린스턴대의 피터 커기스(Peter Kirgis)와 사야시 카푸어(Sayash Kapoor)가 이끄는 다수 기관 연구진은 AI 에이전트가 아직 개방형(open-ended) AI 연구—명확한 답이 없고 판단력과 안목이 필요한 자유로운 탐구—를 수행할 능력이 없다는 사실을 발견했습니다. 이러한 능력은 자기 개선형 AI를 구축하는 데 핵심적일 수 있습니다. 연구진은 AI 에이전트가 AI 연구에 필요한 엔지니어링 문제는 해결할 수 있지만, 최상위 머신러닝 학회에 채택되는 논문 수준의 독창적 연구를 생산할 판단력과 창의성이 부족하다는 것을 확인했습니다. 이러한 격차는 AI 연구 자동화에 대한 일부 과장된 시간표가 실제 증거를 앞서가고 있을 수 있음을 시사합니다.

에이전트가 AI 연구를 자동화하는 방안에 관한 기존 연구 대부분은 검증 가능한 답이 있는 좁은 과제—예를 들어 엔지니어링 문제 해결이나 벤치마크 대상 소형 언어 모델 사후 학습—를 완수하는 능력을 평가합니다. 그러나 AI 연구에서 진전을 이루려면 개방적 사고도 필요합니다. 즉 가설 집합을 선택하고, 어떤 증거가 문제를 해결할지 판단하며, 언제 처음부터 다시 시작해야 하는지 아는 능력입니다.

연구진은 이런 유형의 역량을 테스트하기 위해 '섀도우 평가(shadow evaluation)'라는 새로운 평가 방법을 제안했습니다. 이는 고품질의 미출판 논문에서 연구 질문을 가져와 AI가 답하도록 하는 방식입니다. 연구진은 오픈소스 소프트웨어인 OpenClaw에서 구동되는 앤스로픽의 Claude Opus 4.8에게 이러한 질문—이번 경우에는 저명한 머신러닝 학회인 NeurIPS 2026에 제출된 두 편의 논문에서 가져온 것—을 풀도록 했습니다. 첫 번째 질문은 대규모 언어 모델의 행동을 결정하는 '페르소나'를 모델의 가중치(학습 중 배운 모든 것을 저장하는 수십억 개의 숫자)를 편집해 제어할 수 있는지 여부였습니다. 다른 질문은 스프레드시트 데이터를 기반으로 예측을 하는 모델이 신뢰할 수 없게 된 시점을 지적하는 탐지기를 설계하는 방법이었습니다. 논문이 공개되지 않았기 때문에 에이전트는 학습 데이터에서 답을 암기하거나 온라인에서 찾을 수 없었습니다.

에이전트에는 6일, 앤스로픽 API 크레딧 3,000달러, 실험 실행용 GPU 예산, 자체 가상 컴퓨터, 그리고 개방형 웹 접근 권한이 주어졌고, 최상위 AI 학회에 게재할 가치가 있는 연구 논문을 작성해야 했습니다. 원 논문 저자들은 학회에 제출된 논문을 평가하듯 에이전트의 논문을 채점했습니다. 그 저자들은 두 논문 모두 거절했습니다.

인간 과학자들이 발견한 바에 따르면 에이전트는 연구 수행에 필요한 모든 엔지니어링을 해낼 수 있었습니다. 에이전트는 문헌을 검토하고, 수백 건의 실험을 실행하고, 결과를 정리했습니다. "반면 에이전트는 연구 자체를 수행하는 능력이 명백히 부족했습니다"라고 카푸어는 말합니다. 에이전트는 기이한 실험을 실행했고(일부 경우 지극히 작은 합성 데이터셋에서 가설을 테스트했음), 자신의 연구를 이해할 수 있게 글로 쓰는 데 애를 먹었으며, 해당 분야에 어떠한 새로운 기여도 하지 못했습니다. "논문들은 최상위 AI 학회 수준의 품질에 비하면 기준에 크게 미치지 못했습니다"라고 그는 말합니다.

이는 에이전트가 연구 수행에 필요한 창의성과 판단력을 발휘하는 데 어려움을 겪었기 때문입니다. 에이전트는 다양한 아이디어를 탐색하는 데 충분히 노력하지 않았고, 유망하지 않은 접근법에 너무 빨리 매달렸습니다. 에이전트가 원 저자들 자신이 시작했던 것과 유사한 참신하고 야심 찬 가설을 세우기는 했지만, 매우 제한된 데이터를 근거로 이를 기각했습니다. 또한 실패한 접근법에서 물러서지 못했습니다. 작은 방향 전환은 가능했지만 접근 방식을 근본적으로 재고할 수는 없었습니다.

원문 보기
원문 보기 (영어)
The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research—free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI. A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence. Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking—choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over. To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called “shadow evaluation,” which requires the AI to answer a research question from a high-quality unpublished paper. The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026. The first question was whether a large language model’s “personas,” which determine its behavior, can be controlled by editing the model’s weights (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online. The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers’ original authors graded the agents’ papers as they would evaluate one submitted to a conference. Those authors rejected both papers. The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results. “On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be. For all their failures, the agents didn’t engage in the misbehavior that researchers call “ reward hacking ,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says. Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment. There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer. Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work. Even so, the finding may also echo what AI companies are finding internally. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that it rhymes with what the company found when it tried to automate some aspects of AI safety research. “There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,” he wrote. He called AI systems’ lack of creativity a “bearish signal on short recursive self-improvement timelines.” AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an automated AI researcher an explicit goal, and Anthropic identifies self-improving AI as the industry’s next milestone. “If there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it’s failing currently,” says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it’s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks—the kind that can be scored—while advancing slowly on open-ended research. The big open question, then, is how crucial open-ended research is to recursive self-improvement—whe