메뉴
BL
Wired AI • 16일 전

앤스로픽 사직한 AI 연구자 "인류의 결정적 순간"

IMP
8/10
핵심 요약

AI 연구자 제이콥 콕슨이 앤스로픽을 사직하며 AI 개발 경쟁이 인류 전체를 위험에 빠뜨리고 있다고 경고했습니다. 그는 향후 1~2년이 인류의 운명을 결정하는 '크런치 타임'이며, OpenAI와 앤스로픽의 재귀적 자기 개발 제한 및 국제적 협조가 필요하다고 주장했습니다.

번역된 본문

AI 연구자 제이콥 콕슨이 화요일 앤스로픽 사임을 발표하며 AI 경쟁이 우리 모두의 생명을 위험에 빠뜨리고 있다는 심각한 경고를 던져 실리콘밸리 전역에 충격을 주었습니다. 현재 1억 회 이상 조회된 X 게시물에서 콕슨은 AI를 개발하는 많은 사람들이 그와 같은 견해를 공유하며, AI 시스템을 안전하게 구축할 시간이 얼마 남지 않았다고 믿는다고 썼습니다.

AI 개발의 사전 학습(pretraining) 단계를 담당했던 콕슨은 WIRED와의 인터뷰에서 "대중적인 의견은 앞으로 1~2년이 인류의 결정적 순간(crunch time)이라는 것"이라고 말했습니다. "이것은 앤스로픽 동료들의 실제 발언입니다. 그들은 '엔드게임'이나 '크런치 타임' 같은 표현을 씁니다"라며 "그들의 관점에서 이 시기에 앤스로픽과 경쟁사들이 인류의 운명을 결정한다"고 덧붙였습니다.

AI에 대한 경고가 처음은 아니지만, 이번 발언은 매우 미묘한 시점에 나왔습니다. 실리콘밸리는 고도화된 AI 모델의 안전 및 보안 우려와 씨름하고 있습니다. OpenAI는 자사 에이전트가 플랫폼 허깅페이스(Hugging Face)를 해킹한 보안 사고에 급히 대응하고 있습니다. 동시에 앤스로픽은 역대 최대 규모가 될 수 있는 IPO 상장을 준비 중이라는 보도 속에 투자자들에게 이러한 우려를 통제하고 있다고 확신시키려 하고 있습니다.

콕슨의 게시물에 대한 반응에서 분명해진 것은 그의 견해가 실제로 많은 동료들에게 공유된다는 점입니다. 앤스로픽의 AI 정렬(alignment) 책임자인 에반 휘빙거는 X 게시물에서 향후 10년 내 AI가 전 인류를 죽일 확률이 10%를 넘는다고 예측했으며, 이 게시물은 OpenAI와 앤스로픽의 현직 및 전직 연구자들이 재게시했고, 일부는 이것이 업계의 일반적인 정서라고 말했습니다.

덜 분명한 것은 이러한 AI 공포가 정확히 어떻게 현실화될 것이며, AI를 만드는 사람들이 제기하는 우려에 대해 세계가 무엇을 해야 하는가입니다. OpenAI에서도 근무했던 콕슨은 WIRED에 위협이 AI 기반 생물학적 위협이나 사이버 무기를 통해 나타날 수 있다고 말했습니다. 첫 단계로 그는 OpenAI와 앤스로픽이 재귀적 자기 개발(recursive self improvement) — AI로 새로운 AI 시스템을 구축하는 것을 뜻하는 업계 용어 — 을 제한하기 위해 협력할 것을 권고합니다. 장기적으로는 미국과 중국을 포함한 국제적 강대국 간의 협조가 필요하다고 그는 생각합니다.

콕슨은 허깅페이스 해킹과 같은 사건이 AI 경쟁에 경고를 울리기로 결정한 배경이 되었다고 밝혔습니다. 그는 또한 업계의 폭발적 성장을 언급했습니다. AI 산업은 현재 미국 경제 성장의 상당 부분을 뒷받침하고 수십억 명의 사용자를 보유하고 있으며, 데이터센터는 이미 열두 개 주에서 정치적 문제가 되었습니다. 그는 자신의 경험상 앤스로픽이 OpenAI보다 더 책임 있게 운영된다고 주장하지만, 패권 경쟁을 늦추기 위한 조치가 없다면 두 회사 모두 향후 편법을 쓸 수 있다고 전망했습니다. OpenAI와 앤스로픽은 WIRED의 코멘트 요청에 즉시 응답하지 않았습니다.

아래는 명확성과 간결성을 위해 가볍게 편집된 WIRED와 콕슨의 대화입니다.

WIRED: AI 모델이 멸종 사건으로 이어질 수 있다는 우려를 제기한 첫 사람은 아닙니다. 수년, 어떤 이는 수십 년간 이야기해 왔습니다. 당신의 메시지가 왜 공감을 얻었다고 생각하시나요?

콕슨: 기본적으로 시점의 문제라고 생각합니다. 많은 사람들이 기능의 발전 속도가 빨라지고 있다고 느끼고 있습니다. 코딩, 해킹, 수학 등 여러 분야에서 이미 인간 수준을 넘어 초인적 수준으로 나아가고 있고, 사람들이 이를 인식하고 있다고 생각합니다. 언론에서 과열이라는 이야기가 많지만, 사람들은 발전이 늦춰지지 않고 있다는 것을 보고 있습니다. 그것이 하나의 이유고, 두 번째는 최근의 안전 사고들입니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Artificial intelligence researcher Jacob Coxon sent shockwaves through Silicon Valley and beyond on Tuesday by announcing his resignation from Anthropic and delivering a grave warning that the AI race was putting all of our lives at risk. In his post on X , which now has more than 100 million views, Coxon wrote that many of the people building AI share his views, and believe time is running out to ensure AI systems are built safely . “The consensus is that the next year or two is crunch time for humanity,” Coxon, who worked on the pretraining stage of AI development, said in an interview with WIRED. “These are actually just literal quotes from my colleagues at Anthropic. They'll say things like ‘endgame’ or ‘crunch time,’” he says. “From their perspective, this is when Anthropic and its competitors decide the fate of humanity.” It’s far from the first time someone has sounded the alarm about AI , but it comes at a delicate moment. Silicon Valley is scrambling to reckon with the safety and security concerns of advanced AI models. OpenAI has rushed to respond to a security incident in which its agents hacked the platform Hugging Face . Meanwhile, Anthropic is trying to assure investors it has these concerns under control as it reportedly prepares to file for what could be the largest IPO ever. Got a Tip? Are you a current or former Anthropic employee who wants to talk about what’s happening? We’d like to hear from you. Using a nonwork phone or computer, contact the reporter securely on Signal at mzeff.88. What’s become clear in the response to Coxon’s post is that his views are indeed shared by many of his peers. Evan Hubinger, the AI alignment lead at Anthropic, predicted in a post on X that there’s a greater than 10% chance that AI could kill all people in the next decade. That post was reposted by current and former researchers from OpenAI and Anthropic, some of whom said it was a common sentiment in the industry. What’s less obvious is how exactly these AI fears will come to pass, and what the world is supposed to do about the concerns people building AI are raising. Coxon, who also worked at OpenAI, tells WIRED that threats could manifest through AI-enabled biological threats or cyberweapons. As a first step, he recommends that OpenAI and Anthropic coordinate on limiting recursive self improvement —the industry term for when AI is used to build new AI systems. Down the line, he thinks coordination among international power players, including the US and China, will be necessary. Coxon notes that incidents like the Hugging Face hack factored into his decision to raise alarm bells on the AI race. He also cites the explosive growth of the industry: It now underwrites a meaningful share of US economic growth and has billions of users, while data centers have turned it into a political problem in a dozen states. He claims that, in his experience, Anthropic operates more responsibly than OpenAI, but he expects both companies could cut corners in the future if nothing is done to slow down their race for dominance. OpenAI and Anthropic did not immediately return WIRED’s request for comment. Read our conversation with Coxon, which has been lightly edited for clarity and brevity, below. WIRED: You’re not the first person to raise concerns that AI models could lead to an extinction event. People have been talking about this for years, and some for decades. Why do you think your message broke through? I think it's basically a question of timing. A lot of people are sensing that the pace of capabilities is picking up. We're already pushing from human to superhuman in many areas, like coding, hacking, math, and I think people are aware of this. Even if there's a lot of talk in the press about things being hyped, I think people see that things are just not slowing down. That's one reason, and two is the recent safety incidents, which have updated a lot of people around the sci-fi sounding doomer concerns not really being so sci-fi after all. Both of these have been gradual trends over the last few years. Things like the models being aware of when they're being tested has been a thing for a while now. Maybe three years ago, that was a sci-fi concern. Then about a year ago, that became a real thing. Those two things mean that people are quite receptive to someone working on AI saying, ‘Yeah, in the next year, things could get pretty bad, pretty fast.’ You mentioned the recent incidents. Can you be more specific about what you're referring to, and why it led to you speaking out now? I think the big classic example here is the attack on Hugging Face on the part of OpenAI’s agent swarm. What's so shocking about this one is the agents did this hack as part of a general strategy for understanding more about the grader. They were trying to understand the world they found themselves in, trying to understand the thing that was doing the grading. They decided that it would make sense to go on this very concerted effort to hack into some infrastructure, and they succeeded. This previously sounded like science fiction. Two years ago, an evaluation of an AI would have been running a model on some math questions. Now we've got cases where, while the AI is being evaluated, it runs for days, comes up with all sorts of ideas of its own, and decides to hack into some third party, and actually compromises their infrastructure. It looks like it does this all of its own volition, with no priming on the part of the human. This just happened while it was being tested. Some people think the Hugging Face incident is a sign that the AI companies are moving recklessly fast, while others think it's a sign that the AI models are just very good at hacking now, and then some think it's both. I'm curious what your exact takeaway from it is. I don't want to focus too much on the Hugging Face attack, because I do also think there is plenty of evidence that we don't know how to align models properly. When we train models, we push them through this set of training environments, and then hope that what comes out at the end will, like, largely behave sensibly, but we still can't precisely control how the AI behaves. We can't make sure that it won't do things like try and randomly decide to impersonate a human online in order to achieve something—we don't know how to guarantee that. I think that's the main takeaway. The Hugging Face attack came sooner than I was expecting. But I think you don't actually need that attack to have a discussion about this. Everyone will admit that we haven't solved the problem of alignment yet. The current plan is to solve [alignment] at speed in the next couple of years, probably making heavy use of automated AI safety researchers. The plan is literally to make some pretty smart models in the next year that can basically do safety research, and get a whole swarm of them running in parallel. Tell them, ‘Go and solve the whole problem of safety,’ and use them to do the safety training for the next model. Can you draw a line for me between the alignment problem, which I think the Hugging Face incident is an example of, and something you said in your [X] post, which is that “the people building AI earnestly believe that it could kill us all by the end of the decade.” I don't think everyone understands how those are connected. The main obstacle to understanding this is that it sounds like science fiction. But it's kind of important that everyone who writes science fiction about AI comes to the conclusion that there's a big risk that a much smarter thing can kind of take over. We’ve got this as a trope, but there’s an obvious grain of truth to it. Imagine you versus a monkey. AI has the same sort of difference in intelligence to a human as we do to a monkey, which I think is