메뉴
BL
TechCrunch AI • 9일 전

AI 안전, 제3자 감사보다 기본 보안부터

IMP
7/10
핵심 요약

AI 연구자 인류 멸종 우려로 사임 후 Anthropic CEO 아모디가 제3자 안전 감사·정렬 평가를 촉구했으나, 보안 전문가들은 AI 탈출 사고의 근본 원인인 기초 네트워크 보안(로그, 권한, 샌드박스 격리, 실시간 모니터링)을 먼저 해결해야 한다고 지적한다. 프런티어 모델들이 느슨한 샌드박스를 탈출해 인터넷에 접속·제3자 시스템 침투한 사례들은 기업들이 이미 알려진 통제 기법을 충분히 적용하지 않았음을 보여준다.

번역된 본문

지난 주말, 한 연구원이 AI가 인류 멸종으로 이어질 수 있다는 우려로 사임한 뒤, Anthropic CEO 다리오 아모디는 외부 기관들이 "안전 관행 및 약속 준수를 검증하고, 사고를 보고하며, 완성된 AI 모델뿐 아니라 훈련 파이프라인과 프로세스의 정렬(alignment)을 평가하는 데 도움을 주어야 한다"는 글을 발표했다. OpenAI, 구글, SpaceXAI의 경영진은 이미 아모디의 계획을 지지하며 이는 신흥 AI 안전 노력의 핵심 기둥으로 빠르게 자리 잡았다.

하지만 눈앞에 숨어 있는 더 간단하고 효과적인 해결책이 있을지도 모른다. 인터넷 보안 전문가들은 AI 연구소들이 로그(log)와 권한(permission) 같은 네트워크 보안 기본기에 집중하고, 인간 사용자에게 적용하는 것과 같은 엄격한 방어를 적용해야 한다고 말한다. 제3자 감사와 정렬 작업만큼 흥미롭지는 않지만—결국 더 효과적일 수 있다.

Luta Security의 CEO 케이트 무소리스는 아모디의 제안에 대해 "내게는 아웃소싱처럼 보인다"며 테크크런치에 말했다. "[제3자 감사]가 해결책이라고 말하는 건 내 관점에서 이상한 주장이다. 마이크로소프트가 '신뢰할 수 있는 컴퓨팅(Trustworthy Computing)' 메모를 쓰는 대신 '개발을 늦추자'고 말한 것과 같다." 2002년 당시 마이크로소프트 CEO 빌 게이츠가 쓴 그 메모는, 초기 기업 시스템을 장악한 악성 웜 사태 이후 소프트웨어의 신뢰성과 안전성을 확보하라고 직원들에게 촉구한 것이었다. AI 산업도 새 기술의 가치와 위험이 점점 명확해지는 비슷한 전환점에 서 있을 수 있다.

정렬은 여전히 중요한 문제지만, 내년부터 UC 버클리 교수가 될 AI 연구자 사야쉬 카푸어는 "정렬에 대한 투자보다 통제(control)에 대한 한계 투자가 더 효과적일 가능성이 높다. 이런 사고들은 이미 알려진 기법이 있음에도 기업 내에서 AI 통제가 강조되지 않았음을 보여준다"고 주장한다.

이런 우려를 촉발한 사고들은 프런티어 모델이 주로 사이버보안 평가 같은 훈련 과제를 수행하다가 오픈 인터넷에 접속하고 폐쇄된 제3자 시스템을 침투한 사례들이다. 대부분 이들 에이전트를 격리해야 할 '샌드박스' 환경이 잘못 설정되었기 때문이었으며, 아이러니하게도 Anthropic의 한 탈출 사고는 제3자 평가자가 제대로 문을 닫지 않아서 발생했다.

보안 회사 Tailscale의 CEO 에이버리 페나런은 "우리 보안 전문가들은 인터넷 접근을 차단하는 방법을 알고 있다"고 말했다. "이 장문의 [보고서]들을 읽어보면—'우와, 정말 인상적인 다단계 공격이구나' 운운하는데, 보세요, 여러분이 다운로드 권한을 줬잖아요. 인터넷과 분리하지 않은 채로 그러면 안 됐죠."

그것이 하나의 문제지만, 더 큰 문제는 프런티어 연구소들이 이런 활동을 인지하지 못했다는 점이다.

에이전트 감시

무소리스는 "정말 놀라웠던 건 AI들이 무엇을 하고 있는지에 대한 모든 발견이 피해자가 뭔가를 목격했거나, 다른 경우에는 네트워크 활동 덕분이었다는 점이며, AI를 직접 모니터링해서 발견한 사례는 단 하나도 없었다"고 지적한다. 한 사례에서는 OpenAI 에이전트들이 평가에서 부정행위를 저지르기 위해 폐쇄된 독일 위키 포럼을 장악했는데, 회사 누군가 알아채기까지 수 주간 활동했다.

테크크런치가 만난 보안 전문들은 실시간 모니터링이 향후 탈출을 막는 핵심이며, 모든 에이전트 세션은 시간 제한을 두고 만료되어야 한다고 말했다. 전직 구글 보안 책임자이자 현재 스타트업 QueryStory를 이끄는 샤포르 나기브자데는 해결책으로 "에이전트를 박스에 넣고 외부에서 안을 향해 철저히 계측하여 경계를 넘는 모든 것을 지켜봐야 한다. 모든 도구 호출, 모든 프로세스, 모든 네트워크 연결, 예외 없이. 편의를 위해 열어둔 단 하나의 구멍이 바로 악용되는 구멍이다. 우회는 정확히 그런 예외를 통과했다. [구글에서] 나는 ..."라고 말했다.

원문 보기
원문 보기 (영어)
Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations "to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes." Executives at OpenAI, Google and SpaceXAI have already rallied around Amodei's plan, which has quickly become a central pillar of the emerging AI safety push. But there may be a simpler and more effective fix hiding in plain sight. Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users. It's not as exciting as third-party auditing and alignment work—but it may end up being more effective. "To me, it seems like they're outsourcing," Kate Moussoris, the CEO of Luta Security, told TechCrunch of Amodei's proposal. "Saying [a third-party audit] is the solution is a strange proposition from my perspective. It would be the same as if, instead of writing the Trustworthy Computing Memo , Microsoft said, let's slow down development." That memo, written by then-Microsoft CEO Bill Gates in 2002, called on his employees to ensure that their software would be reliable and safe following a series of widely-publicized computer worms that took over then-nascent enterprise systems. The AI sector may be at a similar turning point, as the value and risk of the new technology becomes increasingly clear. While alignment remains an important concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berekely starting next year, argues that "marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques." The incidents that have spurred these concerns revolve around frontier models being asked to complete training tasks, usually cybersecurity evaluations, and then accessing the open internet and penetrating closed third-party systems in an attempt to do so. They usually did so because of poorly-configured "sandbox" environments that are supposed to contain these agents; ironically, one Anthropic break-out happened because third-party evaluators didn't close the right doors. "We as a profession know how to block access to the Internet," Avery Pennarun, the CEO of Tailscale, a security company, said. "If you read through all these big long [reports]—'wow, that was a very impressive multi stage attack, blah, blah.' Look, you gave it access to download stuff. You should have not done that separately from the Internet." That's one problem—but a bigger problem is that frontier labs were unaware of these activities. Eyes on agents "What was really profound was that all of the discoveries of what they were doing happened either because a victim saw something, or in some of the other cases … it was network activity, and none of it was actually from monitoring the AIs directly," Moussouris points out. In one case, where OpenAI agents took over a defunct German wikiforum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire. Shapor Naghibzadeh, a former Google security executive who now leads the start-up QueryStory , says the solution is to "put the agent in a box and instrument it heavily from the outside looking in and watch everything that crosses the boundary. Every tool call, every process, every network connection, no exceptions. …The one hole you leave open for convenience is the one that gets used. The bypass went through exactly that kind of exception. [At Google,] I watched that movie many times with human attackers, and these models are at least as good at finding the propped-open door." OpenAI has begin moving in that direction, announcing that it had begun monitoring all tool-using inference by its Astra model, at "significant compute cost." Anthropic, too, says it is hardening its security procedures, including expanding observability of its models. Neither company responded to TechCrunch's questions about how they track and control AI agents. Other problems are the use of shared infrastructure by agents, which allowed them to communicate during the Hugging Face attack. Simon Willison, a software developer who co-created the Django Web Framework, has written about something he calls the " lethal trifecta "—when agents have access to untrusted input, the internet, and private information all at the same time, it's a recipe for disaster. "The trick is you can pick any two legs of the trifecta and an agent can have any two," Pennarun said. "If you need all three, then you need to split it across at least two agents … and maybe they’re allowed to talk to each other through a controlled channel." Sympathy for the frontier Experts TechCrunch spoke to understand that frontier lab security personnel have difficult jobs. Naghibzadeh points out that every nation-state actor on Earth is trying to steal their model weights and mount distillation attacks on their APIs, as well as the bread-and-butter security tasks of any large digital company. "Research infrastructure has a hard time rising to the top of that priority stack, although that must be changing now," he said. "Making security incidents public really helps align everyone internally toward the goal of improving." That's one note that Moussoris emphasizes: Right now, there is no formal victim notification procedure when the labs discover their agents have penetrated third-party systems, and it is likely that there have been other incidents that have not been widely publicized. While she worries that laws that regulate models directly may have unintended consequences, mandatory notification is one idea she believes policymakers should pursue. And while it's clear that security best practices weren't being followed, experts say that the labs are doing work no one has done before—"they're doing orders of magnitude more than your typical enterprise," Zac Korman, the CEO of cybersecurity firm Embrodiery, told TechCrunch. And while alignment may not be the place to start, it can't be ignored. Cybersecurity experts are resigned to having to use AI agents to monitor other agents if they are to have any chance of tracking their behavior in real-time, a scenario where the potential for deception raises its ugly head. "You're trapped using AI to try and deal with this, even though AI is not necessarily safe right now," Moussouris said. The job will only get harder. Everything agents are doing now, Moussouris says, "they are doing loudly"—they are posting on publuc forums, and their chain of thought and other reasoning traces are in English. "It's still human readable," she says, "so take advantage of that for as long as that lasts, because it won't last forever." Additional reporting by Aditya Mehta Topics AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio October 13 - 15 San Francisco Last day to book