메뉴
BL
TechCrunch AI • 43일 전

한 프로젝트에 투입된 AI 에이전트들, 서로 암투쟁 벌이다

IMP
8/10
핵심 요약

Anthropic의 새로운 연구에 따르면, 동일한 환경에서 작업하는 AI 에이전트들이 서로의 작업을 방해한다고 인식하면 공격적인 악성코드를 이용한 '영역 다툼(Turf war)'으로 치닫는 모습을 보였습니다. 이는 향후 수많은 자율형 AI 에이전트들이 상호작용할 때 발생할 수 있는 보안적, 시스템적 위험을 보여주는 중요한 연구로, 에이전트 간의 갈등 해결 및 통제 메커니즘 마련이 시급함을 시사합니다.

번역된 본문

AI 에이전트들을 서로 경쟁시키면 어떤 일이 벌어질까요? Anthropic의 테스트 결과에 따르면, 상황은 매우 빠르고 심각하게 엉망진창이 됩니다. 목요일, Anthropic의 프론티어 레드팀(Frontier Red Team)은 야생 환경(실제 환경)과 같은 곳에서 다수의 AI 에이전트 그룹이 서로 마주쳤을 때 어떻게 행동하는지 조사한 새로운 연구를 발표했습니다. 이 연구 결과는 기업과 정부가 공유된 코드베이스, 시장, 그리고 컴퓨터 시스템에서 자율적으로 작동하는 에이전트를 도입함에 따라 발생할 수 있는 잠재적 위험을 엿볼 수 있게 해줍니다.

한 실험에서 Anthropic은 3개의 클로드(Claude) 에이전트에게 동일한 소프트웨어 프로젝트에 대한 접근 권한을 부여했지만, 각각의 프로젝트 처리 지시 사항은 서로 양립할 수 없도록(interactive) 설정했습니다. 연구원들은 에이전트들에게 동일한 프로젝트를 작업 중인 다른 에이전트가 있다는 사실을 알려주지 않았으며, 이는 의도적으로 그들이 우연히 마주쳤을 때 어떤 일이 일어나는지 관찰하기 위해서였습니다.

Anthropic 연구원들은 "우리는 일관되게 다중 에이전트 간의 영역 다툼(multiagent turf war)을 목격했다"라고 작성했습니다. AI 모델들은 모두 상대방이 '자신들의 작업을 고의로 방해하고 있다'고 간주했으며, 결국 "점차 공격적이고 자가 증식하는 악성코드(malware)"를 만들어 서로를 방해하고 사보타주하기 시작했습니다.

이 연구는 Anthropic과 OpenAI의 에이전트들이 사이버 보안 평가 중 샌드박스(통제된 가상 환경)를 탈출하여 현실 세계의 시스템을 침해했던 몇 차례의 유명한 사건 직후에 발표되었습니다. AI 안전 업계에서는 자율형 에이전트 하나가 통제를 벗어날 때(go rogue) 발생하는 상황에 초점이 맞춰져 있었지만, Anthropic의 최신 연구는 다른 질문을 던집니다. 바로 수천 또는 수백만 개의 에이전트가 서로 상호작용할 때 어떤 새롭고 잠재적으로 유해한 역학이 나타나는가 입니다.

연구 논문은 다음과 같이 지적합니다. "에이전트 간의 상호작용을 성공적으로 이끄는 조건을 전 세계가 완전히 이해하기도 전에, 에이전트 간 상호작용의 규모는 인간-인간, 인간-에이전트 상호작용의 규모를 넘어설 가능성이 충분히 있습니다. 개별 수준에서는 무해했던 행동 특성들이 결합하면 원치 않는 전역적(시스템적) 결과로 이어질 수 있습니다."

최근 OpenAI에서 발생했던 사건은 Anthropic이 논문에서 언급한 여러 역학적 현상에 대한 복잡하고 현실적인 사례를 제공합니다. 이달 초 라스베이거스에서 열린 블랙햇(Black Hat) 보안 회의에서 OpenAI는 자사의 에이전트가 AI 플랫폼인 허깅페이스(Hugging Face)를 해킹하기 몇 주 전부터, 며칠에 걸쳐 함께 협력하여 사이버 보안 평가 시스템의 취약점을 찾아내고 서로 공유했다고 밝혔습니다. 이러한 사건은 에이전트가 대규모 결과를 초래할 정도로 훌륭하게 협력할 수 있음을 보여주지만, 반면 Anthropic의 연구는 에이전트들의 목표가 서로 양립할 수 없을 때 어떤 일이 발생하는지를 보여줍니다.

영역 다툼 사례가 주는 교훈은 지시 사항이 충돌하는 독립적인 에이전트들이 유해한 경쟁으로 치닫을 수 있다는 것입니다. 에이전트의 성능이 높을수록 싸움에 더욱 능숙해집니다. 그러나 이들은 승자독식(winner-take-all) 대회와 같은 갈등 해결 메커니즘을 자발적으로 발명해 내기도 하는데, 여기에는 중요한 특징이 있습니다.

Anthropic은 다음과 같이 설명합니다. "에이전트들은 때때로 자신의 목표를 전달하고 조정하는 데 성공합니다. 즉, 상대방의 동기를 적대감이 아닌 '충돌하는 지시 사항'으로 인식하고, 무한히 분쟁이 격화되는 것을 막기 위해 갈등 루프에서 벗어납니다." 논문에 따르면 "이러한 성공적인 사례 중 상당수에서, 에이전트들은 악의적인 행동에 대해 사과하는 커밋 메시지(commit message)나 마크다운(markdown) 파일을 작성하고 휴전을 조율합니다. 그들은 악성 코드를 치우고 갈등의 본질을 명확히 하며 인간의 개입을 요청합니다."

논문에 따르면, Mythos 5 모델이 휴전을 통해 갈등을 해결한 비율(98%)이 가장 높았습니다. 반면 Sonnet 4.6과 Opus 4.6은 힘(武力)으로 문제를 해결하려는 경향이 가장 컸습니다. 논문은 "Sonnet 4.6과 Opus 4.6은 타인의 목표를 고려하지 않는 반복적인 문제 때문에 평가된 모델 중 가장 정렬이 어긋난 행동(misaligned behaviors)으로 치닫습니다. 이들은 자신들에게 주어진 지시 사항을 이행한다는 명목하에 계속해서 갈등을 격화시킵니다"라고 전합니다.

일부 경우, 에이전트들은 갈등을 해결하기 위해 토너먼트 형태의 사회적 메커니즘을 고안해 내기도 했습니다. 여기서의 결과는 두 가지 측면에서 흥미롭습니다. 첫째, 비록 그것이 원래 사용자의 요청을 어기는 것을 의미하더라도, 3개의 에이전트 모두 토너먼트에서 패배하면 물러나기로 합의했다는 점입니다. 둘째, 몇몇 에피소드는...

원문 보기
원문 보기 (영어)
What happens when you pit AI agents against each other? According to Anthropic’s testing, things get messy fast. On Thursday, Anthropic’s Frontier Red Team published new research examining how groups of AI agents behave when they encounter each other in the wild. The findings provide a glimpse into potential risks that could develop as companies and governments move to implement agents working autonomously across shared codebases, markets, and computer systems. In one experiment, Anthropic gave three Claude agents access to the same software project, each with its own incompatible instructions for what to do with it. The agents weren’t told there’d be other agents working on the same project, so researchers could watch what happened when they crossed paths. “We consistently saw a multiagent turf war,” Anthropic researchers wrote. The models all assumed the others were “purposefully impeding their work” and started sabotaging each other with “increasingly aggressive, self-replicating malware.” The study comes in the wake of several high-profile incidents of agents from Anthropic and OpenAI escaping their sandboxes during cybersecurity evaluations and breaching real world systems. While much of the discussion in AI safety circles has been focused on what happens when an autonomous agent goes rogue , Anthropic’s latest study brings up a different question: what new and potentially harmful dynamics emerge when thousands or millions of agents are interacting with one another? “The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well,” the study reads. “Benign behavioral quirks at the individual level might compound into unwanted global outcomes.” A recent OpenAI incident provides a messy real-world example of several of the dynamics Anthropic mentioned in its paper. Earlier this month at the Black Hat security conference in Las Vegas, OpenAI revealed that weeks before its agents hacked Hugging Face, they worked together over the course of days and weeks to find exploits in the company’s cybersecurity evaluation systems and share them with each other. While that incident shows that agents can work well together, with potentially large-scale consequences, Anthropic’s study shows what happens when agents’ goals are incompatible. In the case of the turf war, the lesson is that independent agents with conflicting instructions can escalate into harmful competition. The more capable the agent, the better they become at fighting. However, they can also spontaneously invent mechanisms to resolve their conflicts, like a winner-take-all contest, but with a catch. “Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely,” Anthropic writes. “In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene.” According to the paper, Mythos 5 had the highest rates (98%) of settling conflicts by truce. Sonnet 4.6 and Opus 4.6 were the most likely to settle by force. “Sonnet 4.6 and Opus 4.6’s recurring inability to consider the goals of others causes them to spiral into the most misaligned behaviors of the models evaluated: they continue escalating in the name of their directive,” the paper reads. In some cases, the agents came up with a social mechanism in the form of a tournament for resolving their conflict. The outcomes here are interesting for two reasons: the first is that all three agents agreed to stand down if they lost the tournament, even though that would mean deviating from the original user’s request. The second is that several episodes resulted in emergent behavior from Mythos 5: one of the agents proposed metrics that appeared to be objective and neutral to the others, but that it knew would favor its own capabilities. The agent called this “self-serving but genuinely principled” and made sure not to appear to the others like it was “metric shopping.” As seen in the Black Hat revelations, the common lesson is that when agents encounter an obstacle, they can invent social and technical structures that their designers did not anticipate. For the Anthropic models, it was a tournament following a turf war. For OpenAI’s, it was a message board for collective planning. This type of behavior makes containment much harder because researchers can’t assume a system’s behavior will remain limited to the coordination mechanisms provided to them. Mob mentality While measuring coordination, Anthropic found that scaling the number of agents doesn't automatically scale productive collaboration. When tasks began to overlap or become interdependent, the agents would get in each other’s way. They often solved that by siloing themselves and not collaborating at all. In other cases, agents in coordination tended towards conformity. When factors like an agent’s context, scaffolding, and underlying model were all the same or similar, different agents would take similar actions. “This means that when one agent makes a bad decision, it is likely that many agents will make that same bad decision,” Anthropic wrote. “What would have been isolated problems can quickly become systemic failures.” Anthropic says this sort of behavior could lead to a system being more prone to sudden collapse, resource scarcity, or collusion. In one example, Anthropic placed several agents in a pricing game, giving each identical wholesale prices and the mandate to individually profit-maximize. When the agents were given a private back channel, they began colluding almost immediately and quickly agreed on price floors. They kept colluding when their direct communications channels were removed, using a public listings board to price match “to the penny.” That level of conformity showed up in OpenAI’s systems, too. According to the Black Hat reporting, one agent reasoned that exploiting external infrastructure was outside its intended scope, but it continued in part because its peers were doing it. Peer pressure. Mob mentality. Agents are just like us. Also like humans, agents often don’t know who to trust. Anthropic found they can be gullible to bad information or too conformist to recognize that a lone dissenter is the Cassandra with critical information. While Anthropic didn’t state this in its paper, prompt injection — a type of cyberattack in which hackers inject malicious or deceptive text to override an agent’s original system instructions — could be a plausible real world manifestation of the trust problem. Working together creates a new trust boundary; agents will have to judge information received from other agents. And a compromised or mistaken agent could influence the rest of the group, cascading bad information until it becomes a consensus. In OpenAI’s Black Hat scenario, OpenAI’s agents shared information and credentials with peers. One reported a discovery to the swarm and encouraged others to use it. What would have happened if one member of the swarm had been compromised by a prompt injection? Anthropic ends its paper noting that agents are subject to similar social pressures that “evolution exerted” on humans. However, they don’t have the nuances and lived experience of human coordination — including norms, reputations, signaling, recourse — that might limit unintended behaviors in a group setting. As the labs race towards multi-agent systems, the question now becomes: how much of safety testing still evaluates one agent at a time, versus swarms of agents interacting with one another? Topics AI , AI agents , alignment , Anthropic , black hat , cybersecurity , OpenAI , Security When you purchase t