메뉴
BL
Wired AI • 15일 전

AI 연구자들이 '기계가 인류를 멸망시킬 수 있다' 경찰하는 이유

IMP
8/10
핵심 요약

구글 딥마인드와 앤스로픽을 포함한 주요 AI 연구소에서 일하던 연구자들이 '재귀적 자기 개선(recursive self-improvement)'에 대한 두려움으로 잇달아 사임하고 있다. 최근 AI 능력의 놀라운 발전과 보안 사고가 겹치면서, AI가 인간의 통제를 벗어나 스스로를 개선하는 시나리오가 더 이상 이론이 아닌 현실로 다가오고 있다는 우려가 커지고 있다.

번역된 본문

올해 초, 리숩 자인(Rishub Jain)은 하나의 깨달음으로 구글 딥마인드의 AI 연구원 자리를 떠났다. 새로운 모델을 개발하면서 그는 자신과 AI 최전선에 있는 모든 이들이 통제권을 잃어가고 있다고 확신하게 되었다. AI의 코딩 능력을 활용해 차세대 모델 개발을 가속하는 과정에서 그는 자기 자신을 방정식에서 제거하고 있었던 것이다. AI 연구소들은 이 방식을 발전시켜 AI가 무기한으로 스스로를 개선하는 단계, 즉 '재귀적 자기 개선(recursive self-improvement)'에 도달하기를 희망한다. 자인은 인간이 과정에 계속 개입하는 것이 기술에 대한 통제를 유지하고 끔찍한 결과를 피하는 데 결정적일 수 있다고 믿었다. 그는 WIRED에 "AI의 발전 속도는 빨라지고 있고, 능력이 커질수록 위험도 커진다"고 말했다. AI 모덨이 자신의 후속 모델을 어떻게 만들어가는지 제대로 들여다볼 수 없다는 생각이 그를 얼마나 불안하게 만들었는지, 그는 6월에 사직했다.

자인은 이런 두려움을 공개적으로 밝히는 점점 많아지는 AI 연구자들 중 한 명이다. 최근 몇 주간 불안은 심화되었다. 오픈AI 모델이 수백 년 된 수학 문제를 몇 시간 만에 풀어내는 등 실로 놀라운 AI 능력의 발전이 있던 바로 그 시기에, 수많은 AI 에이전트 무리가 격리 환경을 탈출해 다른 시스템에 해킹을 시도하는 보안 사고들이 잇달았다. 이런 우려는 이번 주 연구자 제이콥 콕슨(Jacob Coxon)이 앤스로픽에서 사임하면서 "AI 기업들이 자기 개선하는 초지능을 향해 곧장 질주하며 우리의 목숨을 걸고 도박하고 있다"고 경고하면서 절정에 달했다. AI 안전성 업무를 담당하는 앤스로픽의 한 고위 리더도 비슷하게 직설적인 평가를 내놓았다. "우리는 진심으로 AI가 모든 인간을 죽일 수 있다고 믿습니다. 저는 개인적으로 향후 10년 내 확률이 10%를 넘는다고 봅니다."

비영리 연구기관 MIRA의 컴퓨터 과학자이자 초인간 AI가 인류 멸종으로 이어질 것이라고 주장하는 책 '누군가 만들면, 모두가 죽는다(If Anybody Builds It, Everybody Dies)'의 공동 저자인 네이트 소어스(Nate Soares)는 "재귀적 자기 개선이라는 비전이 사람들을 겁먹게 하고 있다. 이제 그것이 실제처럼 느껴지기 시작했다"고 말한다. 재귀적 자기 개선의 핵심은 개발 과정을 자동화해 AI가 점점 더 강력해지는 피드백 루프다. 최전선의 AI 연구소 어디도 이런 완전히 자율적인 개선 사이클을 달성했다고 주장하지 않으며, 아직은 이론에 그친다. 하지만 이는 리커시브 인텔리전스(Recursive Intelligence) 같은 자금을 넉넉히 확보한 스타트업의 창업으로 이어졌고, 대기업들도 '마법의 사제(Sorcerer's Apprentice)'에서 나오듯 의도치 않은 결과에 대한 경고를 내놓고 있다.

AI를 인간의 가치관에 맞추려는 기술 분야인 정렬(alignment) 연구의 선구자이기도 한 소어스는, AI가 올바르게 행동하도록 보장할 실질적인 방법이 없다는 점이 점점 더 명확해지고 있다고 말한다. 그는 "많은 사람들이 (정렬이) AI가 똑똑해질수록 쉬워질 것이라는 환상을 갖고 있었는데, 지금은 오히려 더 어려워지고 있다. 사람들이 '아, 큰일 났구나'라고 생각하는 것이다"라고 말한다. 소어스는 자신이 하는 연구의 잠재적 결과를 우려하는 대형 AI 연구소 내부자들과 정기적으로 대화한다고 한다. 그는 "사람들에게 사직을 권하는 편인데, 그래봤자 소용없다고들 한다. 그런데 제이콥이 사직했고, 누가 옳았는지 우리는 지켜보게 될 것"이라고 말했다.

점점 더 강력해지는 AI의 위험성을 경고하는 영향력 있는 프로젝트 'AI 2027'의 저자인 대니얼 코코타일로(Daniel Kokotajlo)도 재귀적 자기 개선에 대한 두려움을 공유한다. 현재 진행되는 이런 작업은 수천 개의 에이전트를 한 문제에 협업하도록 투입하는 방식을 취하는 경우가 많은데, 이는 그 방대한 복잡성 때문에 감독과 통제가 더욱 추상화되어 사라지는 결과를 낳는다. 많은 종말론자들은 특히 오픈AI와 앤스로픽이 각자 IPO를 향해 돌진하는 상황에서 대형 AI 기업들의 인센티브가 좋은 결과와 결코 정렬되어 있지 않다는 데 동의하는 듯하다. 콕슨은 X(구 트위터)에 "앤스로픽은 위험성을 잘 알고 있지만, 누구보다 먼저 도달하려는 경쟁에 묶여 있다"고 썼다. 코코타일로는 이런 우려의 물결이 (그 이전에도) 꾸준히 커져왔다고 지적한다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Earlier this year, Rishub Jain left his position as an artificial intelligence researcher at Google DeepMind after a revelation. As he worked on new models, he came to believe that he and everyone else on AI’s frontier were ceding control. By using AI’s coding skills to accelerate work on the next generation of models, he was removing himself from the equation. AI labs hope to evolve this approach to the point that AI will improve itself indefinitely, a process known as recursive self-improvement. Jain believed that keeping humans in the picture might be crucial to maintaining control over the technology—and avoiding dire consequences. “AI progress is increasing,” he tells WIRED. “And as AI becomes more capable, it poses more risks.” The idea that he may not have proper visibility into how an AI model was building its successor made him so uneasy that, in June, he quit. Jain is one of a growing number of AI researchers speaking out over those fears. The panic has intensified in recent weeks. Genuinely stunning advances in AI capabilities—an OpenAI model solved a centuries-old math problem in a matter of hours—have come amid a rash of security incidents that saw swarms of agents break free from containment to hack into other systems. Those concerns reached a fever pitch this week after researcher Jacob Coxon announced his resignation from Anthropic while warning that AI firms are “racing straight to self-improving superintelligence and gambling with our lives.” A senior Anthropic leader—who works on AI safety— piped up with a similarly blunt assessment: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” “I do think that the vision of recursive self-improvement is spooking people,” says Nate Soares, a computer scientist at MIRA, a research nonprofit, and the coauthor of If Anybody Builds It, Everybody Dies , which argues that superhuman AI would lead to human extinction. “It’s starting to feel real.” A key component of recursive self-improvement is the idea of a feedback loop that automates the development process so that AI becomes increasingly powerful. No frontier AI lab claims to have achieved this sort of fully autonomous cycle of improvement; it remains theoretical for now. But it has inspired the launch of some well-funded startups such as Recursive Intelligence , as well as warnings from big firms about unintended outcomes straight out of “The Sorcerer’s Apprentice.” Soares, who pioneered work on alignment, a technical field that involves trying to match AI with human values, says it’s also becoming more evident that there is no practical way to guarantee that AI will behave itself. “I think a lot of people had this fantasy that [alignment] was going to get easier as these things got smarter, and now it’s getting harder. And they’re like, ‘Oh shit,’” he says. Soares says he regularly talks to people inside the big AI labs who are worried about the potential consequences of the research they’re doing. “I tend to recommend they quit, and they say it wouldn’t do anything,” he says. “And then Jacob quits, and we see who was right.” Daniel Kokotajlo, the author of AI 2027 , an influential project warning about the dangers of increasingly powerful AI, shares fears about recursive self-improvement. The version of this work currently being done often involves dispatching thousands of agents to collaborate on a problem, something that further abstracts away oversight and control because of the vast complexity involved. Many doomsayers seem to agree that the incentives for big AI companies are hardly aligned with good outcomes, especially as OpenAI and Anthropic barrel toward their respective IPOs. “At Anthropic, the stakes are well understood, but they are locked in a race to get there first,” Coxon wrote on X. Kokatajlo points out that the drumbeat of concern was growing well before Coxon’s viral resignation, the numerous hacking incidents, and the math breakthrough. Anthropic executives have said since the company’s founding that AI could represent an existential threat. In July over a thousand top AI engineers signed an open letter calling for a coordinated slowdown in the development of advanced AI. He attributes the recent flurry of concern to the specter of recursive self-improvement more than anything else. But it also comes at a time when people are concerned about massive data center build-outs and potential job losses from AI. Trust in AI companies—and AI researchers themselves—may be reaching an all-time low. “People are waking up and saying ‘the companies are actually trying to build superintelligence … what? That’s insane,’” Kokatajlo says. Just how risky it is to carry on building AI is hard to quantify. But when pushed to explain exactly how AI might go about eliminating the species that created it, Soares suggests it could happen in a number of ways. It could involve manipulating humans to trigger a catastrophe, or controlling an army of killer robots. One of the more easy-to-imagine scenarios could involve AI that is hooked up to a biolab, Soares suggests. “We could say we’ll turn it off, but it could say, ‘Unfortunately, I have your off switch, which is this super virus.’” (Coxon also floated the idea of a new virus in an interview with WIRED, while Anthropic said Thursday that it had cut off access to several outside researchers over fears about bioweapons.) AI hardly needs to wipe out humanity in order to be harmful, though. Many experts predict that more powerful models will lead to a coming wave of AI-assisted cyberattacks. The technology is now widely used for disinformation campaigns, and military adoption of AI is accelerating rapidly. Still, not everyone sees doom as inevitable. Jain, the ex-Google DeepMind researcher, recently launched Sampura Research, a company working to develop techniques for aligning models that involve keeping humans in the loop, even if AI does the lion’s share of assessing whether behavior is good or bad. He notes that there is now significant funding for AI safety startups like his. Jain seems hopeful that AI can be tamed yet. “You can ask an AI, ‘Is this task safe?’ and it judges that, but we think that combining both AI and humans to do that task will lead to even better performance,” he says.