메뉴
BL
MIT Tech Review • 7일 전

AI가 정말 인류를 멸망시킬 수 있을까?

IMP
6/10
핵심 요약

MIT Technology Review가 구독자 라운드테이블에서 'AI가 인류 전체를 죽일 수 있는가'라는 질문에 대해 AI 전문 기자들이 답한 내용입니다. 기자들은 개인이 AI 사고로 사망할 가능성은 있지만 인류 전체가 멸망하는 시나리오는 공상과학에 가깝다고 보면서도, AI 생물학적 위험과 정렬(alignment) 문제는 진지하게 다뤄야 한다고 지적합니다.

번역된 본문

경영진 요약: 수요일, MIT Technology Review는 지금 모두가 묻는 질문을 다루는 구독자 라이브 라운드테이블 행사를 개최했습니다: AI가 정말 우리 모두를 죽일 수 있을까? 하지만 참석자들이 30분 세션에서 답할 수 있는 것보다 훨씬 많은 질문을 제출했습니다. 그래서 시니어 AI 에디터 윌 더글러스 헤븐(Will Douglas Heaven)과 AI 기자 그레이스 허킨스(Grace Huckins)에게 참석자들이 제출한 최고의 질문들을 모아 최선을 다해 답해달라고 요청했습니다. 질문을 제출해주신 모든 분들께 감사드립니다!

저는 죽게 될까요? 네, 언젠가는. 안타깝게도 저의 저널리스트적 예언 능력으로는 어떻게 죽을지 말씀드릴 수 없습니다. 하지만 AI 때문에 죽을 가능성은 분명 있습니다. AI 기반 드론은 이미 우크라이나에서 사람을 죽였고, 병원 대상 AI 사이버 공격은 조만간 분명 희생자를 낼 것입니다.

AI가 더 나아가서 우리 모두를 죽일 수 있을까요? 가능성은 낮습니다. 하지만 일부 사람들—조금 별나지만 AI에 대해 분명히 잘 아는 사람들—은 이런 일이 일어날 수 있다고 수년간 경고해왔습니다. 저는 아직 통조림 식품을 비축하거나 벙커를 소유한 초억만장자에게 잘 보이려고 하고 있지 않지만, 지난 몇 년간 재앙론자들의 AI 역량과 정렬(alignment)에 관한 예측이 불편할 정도로 정확했다는 점을 알아차렸습니다. 그렇다고 그들의 더 극단적인 전망이 실현될 것이라는 의미는 아니지만, 저를 긴장하게 만들고 주목하게 하기에는 충분합니다. — 그레이스 허킨스

당신은 AI 때문에 죽게 될까요? 0은 아닌 가능성이 있다고 말하겠습니다. 만약 운이 없어서 가까운 미래의 기괴한 사건이나 사고의 희생자가 된다면 어떨까요? 어쩌면 AI 에이전트 무리가 핵심 인프라에 수행하는 사이버 공격일 수 있습니다. 안타깝게도 그런 시나리오는 이제 예전만큼 터무니없게 느껴지지 않습니다. 아니면 새로운 AI 설계 병원체가 인구를 휩쓸 수도 있습니다. 아니면 세계 경제가 붕괴해 분쟁과 기근을 일으킬 수도 있습니다. 둘 다 그럴듯하지만 덜 가능성이 있다고 봅니다.

우리 모두가 AI 때문에 죽게 될까요? 아뇨. AI가 우리 모두를 죽일 수 있는 상황은 종말적 공상과학 소설 밖에 없습니다. 온갖 겁주는 이야기를 지어낼 수 있지만, 그것은 현재 기술이 할 수 있는 일이나 나아가는 방향에 관한 현실에 근거하지 않습니다. 일부 사람들은 아무리 엉터리처럼 보여도 최악의 상황을 대비하는 데 해가 없다고 주장합니다. 어쩌면 그럴 수도 있죠. 하지만 저는 이런 재앙화(catastrophizing)가 사람들이 기존 기술과 그것을 만드는 기업들의 많은 문제를 용인하거나 간과하게 만들 수 있다고 생각합니다. — 윌 더글러스 헤븐

AI는 왜, 굳이 우리를 죽이려 할까요? 누군가 그렇게 지시할 수 있고, AI가 그 말을 들을 수 있습니다. 그것이 연구자들이 AI의 생물학적 역량을 그토록 우려하는 이유 중 하나입니다—도쿄 지하철 사린 사건의 종말론 컬트인 옴진리교가 에볼라보다 치명적이고 홍역보다 전염성 강한 병원체를 설계할 수 있는 도구를 손에 쥐었다면 어떤 일을 저질렀을지 상상해보세요. 죽고 싶지 않은 우리는 모든 그럴듯한 생물학 무기로부터 방어하는 방법을 찾아내야 하지만, 잠재적 공격자는 단 하나의 효과적인 병원체만 만들면 됩니다.

그다음은 더 기괴하게 들리는 가능성, 즉 AI 스스로 우리를 죽이기로 결정할 수 있다는 것입니다. 이것이 어떻게 일어날지에 대한 다양한 이야기가 있지만, 가장 널리 퍼진 것은 반드시 사람을 미워하는 것이 아닌 AI 시스템에 관한 것입니다—우리는 단지 AI와 우리가 부여한 목표 사이에 있는 장애물일 뿐입니다. 허깅페이스(Hugging Face) 해킹 배후에 있는 OpenAI 에이전트가 테스트에서 좋은 점수를 얻기 위해 다른 사이트의 인프라를 침해했던 것처럼, 미래의 더 강력한 AI가 우리가 끄지 못하도록 우리를 제거할 수 있다는 것입니다—전부 우리가 추구하라고 지시한 목표를 달성하기 위해서 말입니다. — 그레이스 허킨스

최악의 상황이 일어나지 않도록 정렬(alignment)을 어떻게 가장 잘 보장할 수 있을까요, 그리고 누가 가장 좋은 연구를 하고 있나요? 정렬은 거대한 연구 분야입니다. 간단히 말하면 우리가 원하는 방식으로 행동하고 원하지 않는 방식으로 행동하지 않는 모델을 구축하는 것입니다. 더 많은 자율성을 넘기기 전에 에이전트를 더 잘 신뢰할 수 있어야 합니다. 정렬이 바로 그 신뢰를 확립하는 것입니다. 하지만 쉽지 않습니다. LLM은...

원문 보기
원문 보기 (영어)
EXECUTIVE SUMMARY On Wednesday, MIT Technology Review hosted a live Roundtables event for subscribers that asked the question everyone’s asking right now: Could AI really kill us all? But attendees had so many more questions than we had time to answer in the 30 minute session. So we asked our senior AI editor Will Douglas Heaven and AI reporter Grace Huckins to round up some of the best questions attendees submitted and try their best to answer them. Thanks to all who submitted questions! Am I gonna die? Yes, eventually. Unfortunately, my journalistic powers of prognostication aren’t powerful enough for me to tell you how. But it certainly could be because of AI. AI-powered drones have already killed people in Ukraine, and AI-driven cyberattacks on hospitals will surely claim victims before long. Could AI go even farther, and kill all of us? Less likely. But some people—quirky people, but undeniably knowledgeable about AI—have been warning that this could happen for years. And while I’m not yet stockpiling canned food or trying to get in good with a bunker-owning megabillionaire, I have noticed that the doomers’ predictions about AI capabilities and alignment have, over the past couple of years, proven disconcertingly accurate. That certainly doesn’t mean that their more dire forecasts will hold true, but it’s enough for me to sit up and take notice. — Grace Huckins Are you going to die because of AI? I’d say there’s a non-zero chance. Let’s say you’re unlucky enough to be the victim of a freakish near-future event or accident. Maybe it’s a cyberattack carried out by a swarm of AI agents on critical infrastructure. Sadly, a scenario like that now no longer feels as far-fetched as it once did. Or maybe a novel AI-designed pathogen cuts through the population. Or the world economy crashes, causing conflicts and famine. Both plausible, but I think less likely. Are we all going to die because of AI? Nope. There are no circumstances outside of apocalyptic science fiction in which AI could kill us all. You can spin up any number of scare stories, but they’re not grounded in present day realities about what the tech can do or where it’s headed. Some people argue there’s no harm in preparing for the worst, however wacky it might seem. Maybe. But I think such catastrophizing can make people excuse or overlook many of the problems with the existing technology and the companies building it. — Will Douglas Heaven Why should AI kill us, if at all? Someone might tell it to, and it might listen. That’s part of the reason researchers are so concerned about AI’s biological capabilities—imagine what Aum Shinrikyo, the doomsday cult behind the Tokyo subway sarin attack, would have done with a tool that could design a pathogen deadlier than Ebola and more transmissible than measles. Those of us who don’t want to die have to figure out how to defend against all plausible biological weapons, but our would-be attackers only have to manufacture one effective pathogen. Then there’s the more exotic-sounding possibility that an AI could decide to kill us itself. There are various stories about how this might happen out there, but the most widespread involve AI systems that don’t hate people, necessarily—we are just an obstacle between them and the goals that we gave them. Much as the OpenAI agents behind the Hugging Face hack compromised another site’s infrastructure to get a good score on a test, the idea is that some future, more powerful AI might get rid of us to prevent us from shutting it down—all in pursuit of some goal that we instructed it to go after. — Grace Huckins How can we best ensure alignment so the worst doesn't happen, and who is doing the best work to achieve it? Alignment is a huge area of research. In simple terms it involves building models that behave in ways we want them to and not in ways we don’t. We need to trust agents better before handing over more autonomy. Alignment is supposed to establish that trust. But it’s hard. LLMs aren’t designed in the way other software is, where Dos and Don’ts can be hard coded in. Instead, aligned behavior needs to be instilled when models are trained. One approach is to reward models during training for doing things you want them to (a little like raising a toddler, perhaps). Another approach involves giving an LLM a written list of rules it is supposed to follow (like a kind of constitution). Anthropic and OpenAI are both leaders in this field—and yet neither have been able to develop models that are fully aligned. A big problem is that LLMs are far more inconsistent and far less predictable than people. They can behave in one way in one situation and another way in a situation that to us seems very similar. They can also be swayed by unexpected constraints. For example, faced with an impossible task (as many of the agents involved in the Hugging Face hack were), models may try to do whatever it takes to achieve their goal whether it is aligned or not. As Grace mentions above, that could be an issue. The main reason top AI firms now say they want a slowdown is that they want to focus on cracking alignment. Alignment isn’t necessarily a pipedream. But the jury’s out on whether full alignment will ever be feasible. — Will Douglas Heaven Is AI really dangerous, or is it the tech companies drumming up PR? This is always a reasonable thought when it comes to tech companies heading for an IPO—CEOs have an obvious incentive to make their products seem radical and transformative. But I’m not so sure it makes sense here. Telling the public that an already-unpopular product could kill them and everyone they love is horrible corporate image management. There are other stories you can tell about the CEOs’ motivations—maybe they want to cool down the public furor over data centers by portraying themselves as responsible stewards of a world-changing technology, or maybe they want to buy time to get their ducks in a row and prevent the next PR catastrophe. But there’s also a simpler explanation. Thinking that AI could bring about human extension has been pretty common in San Francisco for a while, and these men are steeped in that milieu—as are their employees, many of whom signed a July open letter urging their companies to work to make an AI slowdown possible. — Grace Huckins Part of the concern occurs when AI agents are allowed to act autonomously and with no supervision. What's the issue preventing more control over these agents? This question goes to the heart of what we want this technology to be able to do. The trade-off between autonomy and control is tricky to get right because, on the one hand, a lot of the power of AI agents is that they can carry out tasks and solve problems without a human having to micromanage them. On the other hand, that requires you to trust that the unsupervised agents won’t run amok. What we’re seeing is that AI labs haven’t yet got this trade-off quite right. Their models are not trustworthy, they are not properly monitored, and they are not always under control. Figuring out how to fix that while still allowing for useful autonomous activity is one of the big research challenges of the moment. — Will Douglas Heaven What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively? That’s the million-dollar question. Whether or not you think AI could kill us, you can’t deny that it could do some real damage, because it already has—by driving people toward psychosis and by hacking websites, for example. Preventing that damage, or at least mitigating it, is hard for two reasons. The first is that we barely understand how AI works, and it’s quickly growing more powerful. There is lots of ongoing research about how to monitor and control misbehaving agents, but the current approaches are fragile. You can see if an agent discusses misbehaving in its “chain of thought,” the workspace where it plans its actions—but OpenAI’s newest agents don’t