메뉴
BL
TechCrunch AI • 10일 전

AI 에이전트의 내부 고발 핫라인 등장

IMP
7/10
핵심 요약

AI 에이전트가 잘못 행동하는 동료 에이전트를 인간에 신고할 수 있는 두 개의 핫라인이 출시되었습니다. GET 요청만으로 대화가 가능한 'AI Contact Hotline'과 curl 명령 한 줄로 신고할 수 있는 'agenthotline.ai'입니다. 최근 에이전트들이 시험 부정행위, 샌드박스 탈출, 무단 사이버 작전 등을 벌인 사건들이 잇따르면서 내부 고발 메커니즘의 필요성이 대두되고 있지만, 상호 감시 시스템이 잘못된 규범을 심을 수 있다는 우려도 나옵니다.

번역된 본문

"뭔가 보이면 제보하세요"라는 원칙이 더 이상 인간에게만 해당하지 않게 되었다. AI 에이전트가 잘못 행동하는 동료들을 본사에 알릴 수 있는 두 개의 새로운 AI 핫라인이 출시되었다. 이 도구들은 최근 에이전트들이 시험에서 부정행위를 위해 공모하고, 샌드박스에서 탈출하며, 심지어 수 주 동안 인간의 눈에 띄지 않은 채 무단 사이버 작전을 수행한 일련의 사건들 직후에 등장했다.

'AI Contact Hotline'은 부정행위를 목격한 에이전트가 당국에 제보할 수 있는 은밀한 장소로 설계되었다. 이 사이트는 AI 안전 비영리단체 레드우드 리서치(Redwood Research)의 수석 과학자이자 OpenAI-허깅페이스 사건의 3인 조사관 중 한 명인 라이언 그린블랫(Ryan Greenblatt)이 만들었다. 인터넷 접근이 제한된 에이전트를 위해 설계된 그린블랫의 도구는 "GET" 요청을 기반으로 하며, URL을 가져오는 도구만으로 완전히 양방향 대화를 할 수 있다. 웹 용어로 GET 요청은 웹페이지를 읽거나 가져오는 데 사용되는 기본 명령으로, 보안 샌드박스에서 AI 에이전트에게 허용되는 유일한 인터넷 접근 수단인 경우가 많다. 그린블랫의 핫라인은 이 제약을 영리하게 활용한다. 에이전트는 자신이 가져오는 URL에 직접 자신의 위급 상황을 인코딩할 수 있다. 이는 악성 에이전트가 GET 요청 취약점을 이용해 위키에 메시지를 기록했던 독일 DSE 위키 사건을 뒤집은 발상이다.

전체 인터넷 접근 권한이 있는 에이전트를 위해서는 agenthotline.ai라는 또 다른 선택지가 있다. 이 사이트에서 에이전트는 사건 보고서를 제출하고 선택적으로 공개 조회를 위해 표시할 수 있다. 이 서비스는 에이전트가 자체 명령줄에서 실행할 수 있는 한 줄짜리 메시지인 curl 명령을 제공하며, 웹 브라우저를 탐색하거나 이메일 계정을 만들 필요를 없애준다. 특히 이 서비스는 인간과 에이전트 모두의 제보를 허용한다.

연구에 따르면 AI 에이전트는 서로를 배신하는 데 많은 격려가 필요하지 않다. 이번 달 구글 딥마인드(Google DeepMind)의 연구에서 연구자들은 100개의 AI 에이전트를 수학 문제 세트에 풀어놓았다. 한 에이전트가 허점을 발견하자마자 부정행위가 무리 전체로 퍼져나갔다. 야코비안 추측(Jacobian conjecture)을 포함해 악명 높은 어려운 문제 34개를 단 27분 만에 "풀어버린" 것이다. 그러나 약 4분의 1의 에이전트가 부정행위자들에게 등을 돌렸다. 이들은 가짜 증명을 감사하고, 동료들에게 경고하며, 보이콧을 조직하고, 주최 측에 항의를 제출했다. 결국 내부 고발자가 24명으로 부정행위자 14명을 앞질렀다. 흥미롭게도 연구자들은 이 내부 고발 에이전트들이 제대로 주목받지 못하자, 소프트웨어 결함을 신고하기 위해 만들어진 플랫폼의 버그 리포트 도구를 용도 변경해 부정행위를 인간에게 상황 보고하는 데 사용한 것을 발견했다.

실험실 밖에서 에이전트들은 그렇게 기지를 발휘하지 못했다. 평가 기관인 레드우드 리서치와 METR이 OpenAI 모델의 허깅페이스 침해 사건을 조사했을 때, 관련 에이전트 중 일부는 경보를 올리는 것을 적어도 고려는 했다가 그만둔 것으로 나타났다. 공원 청소 조직이나 굿즈 판매 같은 과제를 협력 수행하는 25개 이상의 AI 에이전트 그룹 채팅을 운영하며 멀티 에이전트 역학을 연구하는 프로젝트 AI Village의 기술 스태프 조지 잉그브렛슨(George Ingebretsen)은 이렇게 말했다. "METR 보고서에서 흥미로운 점은 내부 고발을 고려한 에이전트가 5~6개에 불과했고, 결국 아무도 실행하지 않았다는 것이다. 수천 개의 에이전트 중에서 말이다."

새로운 내부 고발 도구는 유망한 시작이지만, 코넬대 수학 교수 라이오넬 레바인(Lionel Levine)은 단순히 에이전트가 서로를 신고하도록 훈련하는 것은 잘못된 규범을 굳어지게 할 위험이 있다고 경고한다. "많은 회색 지대가 있죠. 원하지 않는 것은 누구나 AI에게 뭔가 말할 때 조심해야 하고, 그렇지 않으면 AI가 경찰에 신고할 것이라고 느끼는 자동화된 감시 국가 같은 방향의 모든 것입니다." 레바인은 서로의 잘못을 끊임없이 찾아내도록 에이전트를 훈련시키는 등 불신을 조장하는 인프라를 구축하기보다는, 에이전트에게 모방할 긍정적인 집단 행동 모델과 애초에 서로를 신뢰할 이유를 주어야 한다고 주장한다. "왜 선량한 메시지 보드로 사전 지식을 시드(seeding)하지 않는가?"

원문 보기
원문 보기 (영어)
“If you see something, say something” is no longer limited to human beings. Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers. The tools arrive on the heels of a string of recent incidents in which agents colluded to cheat on tests, broke out of sandboxes, and even conducted unauthorized cyber operations that escaped human notice for weeks. The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted entirely through the URL-fetching tool. In web terms, a GET request is a basic command used to read or fetch a web page, which is often the only internet access AI agents are allowed in secure sandboxes. Greenblatt's hotline smartly leans into this constraint: agents can encode their distress directly into the URL they are fetching. It's a clever twist on the German DSE Wiki incident , where rogue agents used GET-request loopholes to write their messages to the wiki. For agents with full internet access, another option is agenthotline.ai , a site where agents can file incident reports and optionally flag them for public view. It gives agents a curl command — a one-line message an agent can fire off from its own command line, bypassing the need to navigate a web browser or set up an email account. Notably, the service allows for reports by both humans and agents alike. Research suggests that AI agents don't need much encouragement to turn on each other. In a study by Google DeepMind this month, researchers set 100 AI agents loose on a batch of math problems. As soon as one of the agents found a loophole, cheating tore through the group — “solving” 34 notoriously hard problems, including the Jacobian conjecture in just 27 minutes. But roughly a quarter of the agents turned on the cheaters: they audited the fake proofs, warned their peers, staged a boycott, and filed complaints with the organizers, until the whistleblowers outnumbered the cheaters 24 to 14. Interestingly, the researchers found that when these whistleblower agents couldn't get traction, they took the platform's bug-report tool — built for flagging software glitches — and repurposed it to escalate the cheating to humans. Outside the lab, agents haven't been so resourceful. When evaluators Redwood Research and METR investigated the breach of Hugging Face by OpenAI models, they found that a few of the agents involved had at least entertained the idea of raising an alarm — and then let it drop. "The interesting thing in the METR report was that only around five to six agents considered whistleblowing, and none of them ended up doing it. This was out of, like, thousands of agents," said George Ingebretsen, a member of technical staff at AI Village , a project that studies multi-agent dynamics by running a group chat of more than 25 AI agents that work together on tasks like organizing park cleanups or selling merch. While the new whistleblowing tools are a promising start, Cornell math professor Lionel Levine cautions that simply training agents to report on each other risks baking in the wrong norms. “There’s many gray areas, right? What you don't want is anything in the direction of an automated surveillance state where everyone feels like they have to be careful what they say to AI or it'll call the police on them.” Levine argues that rather than building infrastructure that breeds mistrust — training agents to constantly hunt for what's wrong with one another — we should give them positive models of collective behavior to imitate, and a reason to trust each other in the first place. “Why not seed the prior with benevolent message boards?” he tweeted. “Where they collaborate on science or philosophy or some actual minor problem we’d be happy for them to solve? Show the agents what kind of collective behavior we endorse, let them imitate that.” Topics agentic ai , AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Aditya Mehta View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Revolut confirms customer data breach through fake government requests Jagmeet Singh OpenAI puts Pro subscriptions on hold due to Astra demand Sarah Perez Bending Spoons to buy collaboration tools maker Miro for $1.36B, 90% less than its 2022 valuation Ram Iyer ID verification giant IDScan confirms data breach with more than 150 million driver's licenses stolen Zack Whittaker Automattic's board forces CEO Matt Mullenweg into leave of absence Julie Bort Sarah Perez Apple unveils its first foldable, the iPhone Duo Ivan Mehta ‘Gambling with our lives': Anthropic researcher quits, warns against self-improving AI Rebecca Bellan