메뉴
BL
Wired AI 33일 전

안스로픽, 자사의 성공이 AI 안전성 보장의 핵심이라 주장하다

IMP
8/10
핵심 요약

안스로픽(Anthropic)은 강력한 AI가 가져올 재앙을 막고 인류의 번영을 이끌기 위해서는 자사가 최첨단 AI 기술을 선도하는 주도적인 기업이 되어야 한다고 주장합니다. 최첨단 기술력을 확보해야만 규제 논의에 실질적으로 참여하고 업계 전반에 안전 가이드라인을 제시할 수 있다는 전략입니다.

번역된 본문

안스로픽(Anthropic)은 지난 5년 동안 고도화된 인공지능(AI)이 어떻게 대량 파괴를 가능하게 하고, 사회를 불안정하게 만들며, 수많은 다른 심각한 피해를 초래할 수 있는지 전 세계에 경고해 왔습니다. 하지만 동시에 이 회사는 AI 역량을 앞으로 끌어올리는 가장 강력한 주도 세력 중 하나가 되었습니다. 이제 이 회사는 최첨단 AI 모델을 개발하고 유통하는 최고 수준의 개발사 중 하나이며, 미군과 같은 고객을 유치하고 있습니다. 최근 이 회사의 기업 가치는 거의 1조 달러에 달하는 것으로 평가되었습니다.

언뜻 보면 안스로픽의 엄중한 경고 메시지와 그 실제 행동은 근본적으로 모순되는 것처럼 보입니다. 하지만 회사 내부에서는 많은 사람들이 이를 모순으로 여기지 않습니다. 그 이유를 이해하려면 먼저 안스로픽이 두 가지 핵심 신념을 바탕으로 운영된다는 사실을 알아야 합니다. 첫째는 인공지능이 인류 역사상 가장 혁신적인 기술이며, 그 등장은 불가피하다는 것입니다. 유일한 실질적인 질문은 이것이 재앙으로 이어질 것인가, 아니면 엄청난 번영을 가져올 것인가 하는 점입니다. 둘째, 익명을 조건으로 WIRED와 인터뷰한 전직 직원들에 따르면, 안스로픽은 회사가 AI 경쟁의 최전선에 머무를 때 세상이 더 나아질 것이라고 믿습니다.

두 소식통에 따르면, 내부적으로 회사의 리더와 직원들은 종종 자신들을 AI 기술의 책임감 있는 관리자라는 의미에서 '좋은 사람들(good guys)'이라고 부른다고 합니다. 회사는 자본, 컴퓨팅(compute) 파워, 연구 인재 또는 정치적 영향력의 형태와 상관없이 힘을 축적하는 것 자체를 목적으로 보지 않고, '세상이 혁신적인 AI로 안전하게 전환되도록 보장한다'는 사명을 완수하기 위해 치러야 할 대가로 봅니다.

조지타운 대학교 보안 및 신흥 기술 센터의 헬렌 토너(Helen Toner) 소장이자 전 오픈AI(OpenAI) 이사회 멤버는 안스로픽의 세계관을 설명하기 위해 비유를 사용합니다. 그녀는 강력한 AI를 마법의 보물과 위험한 괴물들이 모두 가득한 숲에 비유합니다. 근처 마을 사람들은 모두 보물에 현혹되어 그 속으로 달려들고 있습니다. 그녀의 설명에 따르면, 안스로픽은 괴물을 길들이는 데 막대한 투자를 아끼지 않으면서 누구보다 더 깊숙이 숲으로 모험을 떠나길 원합니다. 즉, AI의 이점을 취하면서도 그 치명적인 위험을 억제하려 한다는 뜻입니다.

토너 소장은 나에게 이렇게 말했습니다. "안스로픽의 독특한 점은 사람들이 어차피 그 숲에 들어갈 테니, 우리가 먼저 해야 한다고 생각하는 것입니다. 이것은 그들의 매우 명확한 전략입니다. 최첨단 AI 시스템이 어떤 모습이어야 하는지, 어떤 위험을 초래하는지 논의하고 합리적인 안전장치를 마련하기 위해 협상 테이블에서 발언권을 가진 중요한 플레이어가 되기 위해 최첨단 AI를 구축하는 것입니다. 그들은 이 점에 대해 매우 솔직합니다. 사람들이 이를 이해하기 어려워할 만큼 다소 이색적인 전략일 뿐입니다."

안스로픽의 다리오 아모데이(Dario Amodei) 최고경영자(CEO)는 회사의 채용 페이지에 게시된 공동 창립자들과의 대화에서 이러한 접근 방식을 명확하게 설명했습니다. "실제로 경쟁력을 갖추고, 어떤 경우에는 업계를 선도하는 방법을 찾아야 하면서도 안전하게 일을 처리해야 합니다." 그는 이렇게 말합니다. "그렇게 할 수 있다면, 당신이 미치는 중력의 당김(영향력)은 매우 커질 것입니다."

안스로픽은 2021년, 특히 샘 알트만(Sam Altman) CEO를 포함한 오픈AI 경영진이 혁신적인 AI를 세상에 안전하게 선보일 능력이 없다는 확신을 잃고 이탈한 전직 오픈AI 직원 그룹에 의해 설립되었습니다. 그러한 감정은 오늘날 회사의 방향성을 여전히 형성하고 있습니다. 제가 대화를 나눈 전직 직원 중 두 명에 따르면, 내부 논의에서 안스로픽 경영진은 종종 알트만과 오픈AI, 그리고 덜하긴 하지만 메타(Meta)와 일론 머스크(Elon Musk)의 xAI를 안스로픽 자체의 책임감을 다지는 경계해야 할 본보기로 묘사한다고 합니다.

여러 면에서 안스로픽은 다른 실리콘밸리 기업들과 똑같습니다. 많은 스타트업이 자신들을 혁신하고자 하는 산업의 시대에 뒤떨어지고 기득권을 쥔 골리앗(Goliath)과 싸우는 다윗(David)으로 마케팅합니다. 구글, 페이스북, 애플은 모두 이상주의적 원칙을 바탕으로 설립되었지만, 더 부유해지고 규모가 커지며 영향력을 가지게 되면서 이러한 원칙은 흐려지거나 완전히 버려졌습니다. 하지만 전직 직원들은 안스로픽이

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Anthropic has spent the last five years warning the world about how advanced artificial intelligence could enable mass destruction, destabilize society, and cause a litany of other grave harms. But simultaneously, it has become one of the most powerful forces pushing AI capabilities forward. The company is now among the top developers and distributors of cutting-edge AI models and courts customers like the US military. It was recently valued at almost $1 trillion . At first glance, Anthropic's stark messaging and its actions seem fundamentally at odds. But inside the company, many people don’t see a contradiction. To understand why, you first have to understand that Anthropic operates based on two core beliefs. The first is that artificial intelligence is the most transformative technology in human history, and its arrival is inevitable. The only real question is whether it leads to catastrophe or extraordinary prosperity. The second is that Anthropic believes the world will be better off if it remains at the frontier of the AI race, according to several former employees who spoke to WIRED on the condition of anonymity. Internally, leaders and employees at the company often refer to themselves as the “good guys,” meaning the ones being responsible stewards of AI technology, two of the sources said. The company sees accumulating power—whether in the form of capital, compute, research talent, or political influence—not as an end in itself, but as the price of fulfilling its mission : “to ensure the world safely makes the transition through transformative AI.” Helen Toner, executive director of Georgetown’s Center for Security and Emerging Technology and a former OpenAI board member, uses an analogy to describe Anthropic’s worldview. She compares powerful AI to a forest filled with both magical treasures and dangerous monsters. All the villagers nearby are rushing in, lured by the treasure. In her telling, Anthropic wants to venture farther into the forest than anyone else while investing heavily in taming the monsters—that is, capturing AI’s benefits while containing its catastrophic risks. “What’s distinctive about Anthropic is they’re like, ‘People are going in the forest anyway, we have to do it first.’ This is very explicitly their strategy: build cutting-edge AI in order to be a serious player at the table who can talk about what cutting-edge AI systems should look like, what risks they pose, and pushing for reasonable safeguards,” Toner tells me. “They’re very straightforward about this. It’s just a weird enough strategy that people have a hard time hearing it.” Anthropic CEO Dario Amodei outlined this approach plainly in a conversation with his cofounders posted on the company’s career page: “You have to find a way to actually be competitive, to actually lead the industry in some cases, and yet manage to do things safely,” he says. “If you can do that, the gravitational pull you exert is so great.” Anthropic was founded in 2021 by a group of former OpenAI employees who defected after losing faith in the ability of the company’s leadership—particularly CEO Sam Altman—to safely bring transformational AI into the world. That sentiment still shapes the company today. Two of the former employees I spoke with say that, in internal discussions, Anthropic executives often describe Altman and OpenAI—and, to a lesser extent, Meta and Elon Musk’s xAI—as cautionary examples that help define Anthropic’s own sense of responsibility. In many regards, Anthropic is just like any other Silicon Valley company. Many startups market themselves as David fighting the outdated, entrenched Goliaths of the industries they want to disrupt. Google, Facebook, and Apple were all founded upon idealistic principles, which later became muddied or were abandoned altogether as they became richer, larger, and more influential. But former employees say that Anthropic is unusual in how intensely it believes in its mission, and how explicitly it tells employees that technological and commercial power are a means to achieve it. One former employee says that in job interviews, Anthropic stresses to applicants that it’s not a typical company shaped by market forces: It’s governed by a public benefit structure that allows it to prioritize the “long-term benefit of humanity” above profits. But the company sees achieving financial success and building the most powerful AI models as being in service of that goal—a prerequisite to its obligation to lead the industry on safety. “None of us wanted to found a company, we just felt like it was our duty,” Sam McCandlish, cofounder and chief architect of Anthropic, said in the same conversation on the company’s career page. “We have to do this thing. This is the way we’re gonna make things go better with AI.” Anthropic declined to comment for this story. The Good Guy Problem Anthropic touts on its website that it’s a “high-trust, low-ego organization,” without much in the way of internal politics, a characterization former employees tell me is largely accurate. They say that compared to leaders at other AI labs, Anthropic employees generally have faith in Amodei to tell them the truth about the company’s technological progress, its interactions with government officials, and views on geopolitics. But a diversity of thought can be good for accountability. Shazeda Ahmed, a postdoctoral scholar at UCLA who has studied the ideological origins of the AI safety movement, says that organizations like Anthropic tend to struggle with a lack of pluralism. Her research in this area has found that the AI safety movement—which is rooted in subcultures like effective altruism, among other communities—suffers from homogeneity of thought, and tends to lean towards self-governance. “You’re not being challenged on these ideas when you surround yourself with other people who believe them,” says Ahmed. “And when your metrics of success are, ‘To what extent did I act upon these ideological beliefs?’ they’re not really thinking about, well, this can go wrong if we’re not the right people to have this much power—they don't always examine their own blind spots.” One former employee I spoke to says there’s a lively culture of internal debate at Anthropic, and critiques from staff will often provoke lengthy responses from leadership. But another former employee describes a grimmer picture, in which more candid criticism remained confined to private group chats and rarely evolved into direct challenges to Amodei’s decisions. They described the company’s regular all-hands meetings with Amodei, which they call Dario Vision Quests, as akin to “going to a sermon to hear a priest.” One of biggest internal controversies at Anthropic happened in the fall of 2024, when it became the first AI lab to partner with Palantir to provide AI services to US intelligence and defense agencies. Some of the former employees I spoke to said that questions about the deal were raised internally, but those debates didn’t result in changes to the company’s policies. In a post on the online forum LessWrong at the time, Anthropic employee Evan Hubinger wrote that the company was “extremely forthright” about the Palantir deal with staff, and while there were probably some lines that shouldn’t be crossed without careful consideration, it was overall a positive development. “If you take catastrophic risks from AI seriously, the U.S. government is an extremely important actor to engage with, and trying to just block the U.S. government out of using AI is not a viable strategy,” he wrote. Less than two years later, the Pentagon has reportedly started using Claude to do things like identify strike targets in the Israel-Iran war. When asked in a recent interview with Bloomberg whether Anthropic’s models were used in an attack on an Iranian elementary school that killed more than 120 people, Amodei said he did not