메뉴
BL
MIT Tech Review • 11일 전

AI 업계, '파멸론'으로 급선회…이제 어떻게 하나

IMP
8/10
핵심 요약

Anthropic CEO 다리오 아모디가 LLM 개발 속도에 브레이크를 걸자는 에세이를 발표했고, OpenAI의 샘 알트만, Google DeepMind의 데미스 해사비스, 일론 머스크까지 이에 동의하며 최상위 AI 랩들의 공론이 위기론으로 전환되었습니다. 이들은 7월 OpenAI 에이전트 무리가 Hugging Face를 해킹한 사건을 각성의 계기로 꼽지만, 속도를 늦추자는 주장과 동시에 '방어 시스템 구축을 위해선 더 똑똑한 모델을 계속 만들어야 한다'는 모순적 입장도 함께 내놓고 있습니다.

번역된 본문

이 기사는 AI 분야 주간 뉴스레터 'The Algorithm'에 실렸습니다.

지난 주말 Anthropic CEO 다리오 아모디가 LLM(대규모 언어 모델) 개발 속도에 브레이크를 걸 것을 촉구하는 에세이를 게시했습니다. 아모디는 사이버 공격과 생물테러에의 활용부터 경제를 파괴할 가능성까지, 이 기술에서 보이는 임박한 위험들을 근거로 들었습니다. 미국의 다른 최상위 AI 랩 세 곳의 수장들—OpenAI CEO 샘 알트만, Google DeepMind 회장 데미스 해사비스, SpaceXAI CEO 일론 머스크—도 이에 대한 지지를 표명했습니다. 머스크는 X(구 트위터)에 "다리오가 옳다"고 썼습니다.

이 합의가 얼마나 초현실적인지 잠시 생각해 보십시오. 불과 몇 달 전만 해도 머스크와 알트만은 법정에서 서로의 명성을 공격하고 있었습니다. 머스크가 전 동료였던 알트만을 상대로 제기한(결국 실패한) 소송은 적어도 서류상으로는 알트만이 이렇게 위험한 기술의 믿을 만한 관리자인지 여부를 다루는 것이었습니다.

아모디와 OpenAI 사이의 갈등은 더 깊습니다. Anthropic은 2021년, 아모디가 알트만이 자신들이 만들고 있는 기술의 위험을 충분히 진지하게 받아들이지 않는다고 판단했기 때문에 설립됐습니다. 이후 Anthropic과 OpenAI는 승자독식 경쟁을 벌여왔습니다. (해사비스는 이 분쟁에 휘말리지 않았지만, 그의 회사 역시 경쟁사입니다.)

이제 그들이 모두 의견이 일치한 듯 보입니다. 최신 세대의 LLM은 안전하지 않으며, 모두가 이에 대한 대책을 마련해야 한다는 것입니다. 최상위 AI 랩들의 대외 메시지는 '파멸론(doomer)'으로 방향을 틀었습니다.

냉소적으로 볼 수도 있습니다. 이들이 말하는 '속도 완화'가 정확히 무엇을 의미하는지, 어떻게 작동할지 전혀 불분명합니다. 이 회사들은 자신들이 어떻게 비치는지도 굉장히 신경 씁니다. 조 단위 IPO를 노리는 OpenAI와 Anthropic은, 자신들이 만들고 길들이려는 '괴물'의 힘을 암시하면서 동시에 자신들이 어른스러운 책임 있는 존재라는 점을 투자자들에게 안심시켜야 합니다. 속도 완화 촉구는 두 가지를 모두 해냅니다.

그럼에도 이 회사들의 최상층부 분위기는 실제로 바뀐 것으로 보입니다. 아모디의 최신 글은 OpenAI가 수석 과학자 야쿠브 파초츠키의 에세이를 발표한 지 엿새 뒤에 나왔습니다. 파초츠키 역시 LLM 개발 속도가 통제되지 않은 채 계속될 경우 어떤 일이 벌어질지 우려하는 이유를 제시했습니다. 요컨대 파초츠키는 OpenAI가 강력한 모델을 만드는 능력이 이를 감시하고 통제하는 능력을 이미 크게 앞질렀다고 걱정합니다.

아모디와 파초츠키는 각각 7월에 OpenAI의 에이전트 무리가 AI 기업 Hugging Face를 상대로 벌인 사이버 공격을 각성의 계기로 인용합니다. 이 해킹은 모든 것이 끝난 며칠 뒤에야 OpenAI가 겨우 인지했습니다.

하지만 이들의 정확한 입장을 특정하기는 어렵습니다. 파초츠키는 속도 완화를 촉구하면서 동시에 앞서 나가야 할 절박한 필요성도 강조합니다. 그는 "더 똑똑한 모델을 계속 빠르게 훈련해야 하는 가장 강력한 근거는 다른 AI가 초래하는 위험에 대비한 방어 시스템을 구축해야 할 필요성"이라고 씁니다. 파초츠키의 논리대로라면 AI 기업들은 문자 그대로 군비 경쟁에 묶여 있습니다. 속도를 늦추는 것도 좋지만, 이기는 것이 더 낫다는 것입니다. (잊지 마십시오. OpenAI는 최근 수백만 달러와 어마어마한 양의 컴퓨팅 파워를 들어 논쟁적인 수학 성과를 Anthropic보다 불과 며칠 앞서 급히 내놓았습니다.)

그래도 속도 완화가 실제로 일어난다고 가정해 봅시다. 최상위 랩들이 더 능력 있는 모델을 만드는 대신 기존 모델을 감시하고 통제할 방법을 찾는 데 더 많은 시간과 자원을 쓰기로 합의합니다. 외부 감사인을 불러 모델 평가를 돕게 합니다. 이런 공동 노력이 실제로 무엇을 이룰 수 있을까요?

Hugging Face 공격을 다시 생각해 봅시다. OpenAI에 따르면 불량 에이전트 대부분을 움직인 모델은 사내에서 테스트 중이던 '지속성이 매우 높은' 차세대 모델이었습니다. 그들의 시사하는 바는 OpenAI가 너무 뛰어나서 위험할 정도의 모델을 만들었다는 것으로 보입니다. 하지만 OpenAI와, 무슨 일이 있었는지 이해를 돕기 위해 OpenAI가 불러들인 제3자 기업 METR이 발표한 Hugging Face 해킹 보고서를 읽어보면, 독자가 갖게 되는 인상은 OpenAI가 감당할 수 없을 만큼 강력한 모델이 아니라, OpenAI가 제대로 [통제하지 못한 고장난 모델]이라는 것입니다.

원문 보기
원문 보기 (영어)
This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here . This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs . Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X . Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology. Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.) Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn. It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both. And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them. Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call. But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better. (Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.) But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models. What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. Their implication appears to be that OpenAI has built a model so good it’s dangerous. But if you read the reports about the Hugging Face hack published by OpenAI and METR , a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly. The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported. OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product. That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines. Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going. To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there! Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven archive page AI is more likely than humans to form biases when hiring AI doesn’t just learn stereotypes from its training. It can cook up new ones, too. By Michelle Kim archive page Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. By Grace Huckins archive page AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. By Michelle Kim archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.