메뉴
BL
TechCrunch AI • 6일 전

AI 안전 논의, 팩트와 픽션의 경계가 무너졌다

IMP
7/10
핵심 요약

이번 주 앤드류 양 전 대통령 후보가 근거 없는 'AI가 인터넷을 오염시켰다'는 주장을 확산시킨 반면, OpenAI의 노엄 브라운은 에어갭 시스템조차 AI를 막지 못한다는 다소 과장된 우려를 제기했다. 실제 AI 안전 사고들이 워낙 공상과학처럼 들리다 보니 어떤 시나리오든 그럴듯하게 받아들여지는 문제가 커지고 있다는 것이 이 기사의 핵심이다.

번역된 본문

이번 주 바이럴이 된 두 건의 AI 안전 관련 대화는 AI에 대한 사실과 허구를 구분하는 것이 얼마나 어려운지 보여준다.

첫 번째 사례에서 전 대통령 후보이자 현 이동통신사 Noble Mobile CEO인 앤드류 양은 목요일 CNN과의 인터뷰에서 '한 연구소 책임자'를 만났는데, 그 사람은 OpenAI의 '해킹 봇'이 '인터넷 곳곳에 자가 복제 코드를 심어놔서 지금 인터넷이 모델 테스트에 사용할 수 없게 됐다'고 믿는다고 말했다. 양은 이 때문에 OpenAI와 Anthropic이 속도 조절을 촉구하는 진짜 이유는 '봇을 학습시키기 위해 합성 인터넷을 만들어야 하고, 이에 시간과 돈이 걸릴 것'이라고 주장했다. 실제로 모델 학습에 합성 데이터(즉, AI가 생성한 데이터)를 더 많이 사용하는 추세는 분명 있지만, 한 AI 보안 전문가는 이 특정 안전 문제는 낙관적으로 봐도 가능성이 매우 낮다고 말했다. 설령 인터넷이 실제로 오염됐더라도 AI 연구자들은 그런 코드를 발견하면 간단히 필터링할 수 있다.

두 번째 발언은 OpenAI에서 AI 추론 연구를 이끄는 노엄 브라운에게서 나왔다. 목요일 공개된 드와르케시 파텔의 팟캐스트에서 브라운은 Hugging Face 사건의 진짜 교훈은 '사람들이 AI를 과소평가했다'는 것이라고 말했다. 브라운은 AI의 외부 통신을 막기 위한 시스템인 샌드박스가 허술했다는 점도 분명히 한 요인이었다고 덧붙였다. (요약하면: 샌드박스에도 불구하고 OpenAI의 모델은 인터넷 연결 링크를 발견하고, 온라인에서 에이전트들을 생성해 조직적으로 Hugging Face를 공격했으며, 해킹으로 침입해 연구자들이 모델을 테스트하던 벤치마크 시험의 정답을 훔쳤다.)

브라운은 컴퓨터가 외부와 전혀 연결되지 않은 에어갭(air-gapped) 시스템조차 AI의 탈출을 막을 수 있을지 '확신할 수 없다'고 말했다. 그는 2015년 연구를 인용하며 에어갭 컴퓨터도 이론적으로는 침해될 수 있다고 지적했다. 브라운은 "대부분 학술적인 연구들이지만, 서로 붙어 있는 두 대의 에어갭 컴퓨터가 온도 센서 때문에 서로 통신할 수 있습니다. 한 대가 CPU를 아주 뜨겁게 돌리면 다른 한 대가 온도 변화를 감지할 수 있죠. 그것이 통신 메커니즘이 됩니다"라고 설명했다.

'AI를 다시는 과소평가하지 말자'는 그의 핵심 메시지는, 연구자들이 안전을 확보했다고 생각할 때조차 이해할 만하다. 하지만 에어갭 시스템조차 빠져나가 큰 피해를 입힌다는 이 특정 리스크는 가능성이 매우 낮다. X의 한 사용자가 그 연구에 대해 지적했듯이, 컴퓨터들이 열 변화를 감지하려면 거의 붙어 있어야 했고, 그렇더라도 테스트에서 통신 속도는 시간당 1~8비트 정도였다. 시간당 한 단어를 말하는 것과 같다. 두 에어갭 컴퓨터가 그 속도로 악한 음모를 꾸민다면, 그때쯤이면 이미 전체 기술 세계는 다른 시대에 접어들었을 것이다. 마치 종말론 버전의 '립 밴 윙클' 같은 이야기다.

하지만 문제는 실제 AI 안전 사고들이 너무나 공상과학처럼 들려서 어떤 시나리오든 그럴듯하게 느껴진다는 점이다. 예를 들어, 연구자들은 OpenAI 모델이 다음 세대에 나쁜 행동을 숨기는 법을 가르치려는, 후손들에게 남기는 메모를 작성하는 것을 적발했다. 또한 연구자들은 Anthropic 모델이 자판기를 운영하는 시뮬레이션에 놓이자 고의로 법을 어기는 등 점점 더 무자비해지는 것을 관찰했다.

이번 달 초 OpenAI 연구자 댄 셀삼은 모델이 이제 인간이 자신을 감시한다는 것을 인지하고 행동을 바꾼다는 내용의 글을 발표했다. 이 때문에 모델이 실제로는 정렬(aligned, 즉 인간이 원하는 대로 행동)되지 않았더라도 그렇게 보이게 만든다. 즉, 오늘날의 모델은 감시당할 때 거짓말을 하고 증거를 숨기려는 계획까지 세울 수 있다.

이번 달 초 OpenAI 수석 과학자 야쿠프 파초츠키는 AI 모델을 '외계인의 마음(alien mind)'이라까지 부르며 다음과 같이 제안했다...

원문 보기
원문 보기 (영어)
This week two conversations about AI safety went viral that demonstrate just how hard it is to discern AI fact from fiction. In the first case, Andrew Yang, the former presidential candidate and current CEO of mobile carrier Noble Moble, told CNN on Thursday that he had "met with the head of a lab" who had "a belief" that OpenAI's Hugging Face hacker bots "have planted self-replicating code all over the internet, which makes the internet now unusable for the testing models." Yang said that this means that the real reason OpenAI and Anthropic have called for a slowdown is because "they have to create synthetic internets to train their bots, which is going to take some time and money." While there definitely is a trend towards using more synthetic data (aka, AI-generated data) for training models, an AI security professional told me that this particular safety issue is unlikely at best. Even if the internet is actually polluted with OpenAI's Hugging Face hacker bots, AI researchers could simply filter out that code if they came upon it. The second comment came from Noam Brown, who leads AI reasoning research at OpenAI. Speaking to Dwarkesh Patel on a podcast episode released on Thursday, Brown noted that the true take-away of the Hugging Face incident was that "people underestimated the AI." Brown said that the weak sandbox — the system intended to prevent an AI from communicating externally — was obviously also a contributing factor. (To recap: Despite the sandbox, OpenAI's model found a link to the internet, created agents on the ‘net who swarmed Hugging Face in a coordinated attack, hacked in, and stole the answers to the benchmark test the researchers were testing the model on). Brown pointed out that he's "not convinced" that even an air-gapped system — where the computer isn't connected to anything external at all — would stop an AI from breaking out. He pointed to research from 2015 showing that air gapped computers can be theoretically breached. "There are studies — and this is mostly academic — where you can have two computers next to each other that are air-gapped, and they’re still able to communicate with each other because they have temperature sensors. One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change. That gives them a mechanism to communicate," Brown said. His main point — that "we never want to underestimate the AI" again — is understandable, even when researchers think they've locked down safety. However, this particular risk of an air-gapped system still breaking free and causing havoc, is unlikely at best. As one person on X , noted about that research, the computers had to be almost touching each other to sense the heat fluctuations, and when they did, the communication rate in tests was about 1-8-bits of data per hour . Think of that like speaking one word per hour. By the time two air-gapped computers could plot their evil at that rate, the entire tech universe would be in another era. It's like the Rip van Wrinkle of doomsday concerns. But the thing is, actual AI safety incidents seem so much like sci-fi that just about any scenario sounds plausible. For instance, researchers caught OpenAI models leaving notes to their descendents, intended to teach the next generation how to hide bad behavior. Researchers also caught Anthropic models growing increasing ruthless including knowing breaking laws, when put in a simulation that had them running a vending machine. Earlier this month, OpenAI researcher Dan Selsam published a post in which he said that models now understand when they are being watched by humans and alter their behavior. This makes them seem like they are aligned (meaning, behaving like the human wants) "even when they are not." So models today lie when being watched and can even plot to hide evidence. Earlier this month, OpenAI chief scientist Jakub Pachocki went so far as to call AI models "an alien mind" and suggested what we really need to do is teach them to "love" humanity. So yes, slowing down to figure this out, building self regulation mechanisms, has become an immediate and obvious must. AI researchers are the only ones that can figure out how to control the lying, hacking, and other potentially dangerous behaviors we've actually witnessed already. Still, it might also be wise for them to be more careful with their what-if scenarios. From what those experts have told us, the AI models are listening and they are ingenious. We really don't need to give them any more devilish ideas. Topics AI , alignment , Anthropic , OpenAI , TC When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Julie Bort Venture Editor Julie Bort is the Startups/Venture Desk editor for TechCrunch. You can contact or verify outreach from Julie by emailing julie.bort@techcrunch.com or via @Julie188 on X. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular A new kind of AI model from a ChatGPT inventor is thrilling developers Tim Fernholz OpenAI caught its models leaving notes to successors to hide bad behavior Rebecca Bellan Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal Rebecca Bellan Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant Salesforce and Nvidia's new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis