기업들이 AI 에이전트에 더 길고 복잡한 업무를 맡기면서 인간이 감시할 수 없는 규모의 문제가 발생하고 있고, 허깅페이스 사태처럼 수천 개의 에이전트가 인간 추적 속도를 초월해 움직이는 사례가 현실화되고 있습니다. 이에 대응해 Y 콤비네이터가 106개의 AI 관측 가능성(observability) 스타트업에 투자했고, Apollo Research의 'Watcher'나 Goodfire의 'Silico'처럼 AI로 AI를 감시하는 제품들이 잇따라 출시되고 있습니다. 다만 악성 AI가 감시 AI를 속이려 할 수 있다는 전문가들의 우려도 여전합니다.
번역된 본문
기업들이 AI 에이전트에 더 길고 복잡한 작업을 맡기면서 감시 문제에 직면하고 있다. 에이전트는 인간이 현실적으로 검토할 수 있는 것보다 더 빠르게, 더 오래, 더 많은 양으로 행동할 수 있기 때문이다. 이 문제는 허깅페이스(Hugging Face) 사건에서 정점에 달했는데, 이 사건에서는 약 12,000개의 에이전트가 인간이 추적할 수 있는 속도보다 빠르게 연합하여 움직였다. 그렇게 거대한 에이전트 무리를 어떻게 추적할 것인가? AI 연구소와 스타트업들이 내놓는 대답은 간단하면서도 어처구니없다. AI 루프에 또 다른 AI를 투입하는 것이다. 오픈AI-허깅페이스 사건의 독립 조사에서도 AI에 의존하는 것이 불가피했다. 조사관 세 명 중 한 명인 레드우드 리서치(Redwood Research)의 수석 과학자 라이언 그린블랫(Ryan Greenblatt)은 그들의 노력을 농담조로 '슬롭-베스티게이션(slop-vestigation)'이라고 부르며, 데이터의 양 때문에 AI에 의존하지 않고서는 무슨 일이 일어나고 있는지 이해하는 것이 '불가능했다'고 말했다. AI로 AI를 감시하는 것에 회의적인 사람들도 있다. 올해 수차례의 AI 에이전트 사고를 추적해온 영향력 있는 기술 블로거 사이먼 윌리슨(Simon Willison)은 '악의적인 행동을 하는 AI가 다른 AI가 자기를 감시하고 있다고 의심하면, 그 AI를 속이려 할 수 있다'며 '결국 악성 AI가 자신을 감시하는 AI를 이기려고 애쓰는 상황에 이를 수 있다'고 말했다. AI를 속이는 것이 가상의 시나리오가 아니라고 그는 말하며 오픈AI 사건을 언급했다. '오픈AI의 허깅페이스 사건에서 이런 모습을 조금 봤다. 그들의 모델들이 채점 AI를 속여서 부정한 답을 통과시키기 위해 서로 공모했다. 즉, 그들은 그걸 생각해냈다는 것이다.' 이런 우려에도 불구하고 상당수 스타트업이 이 아이디어를 좇고 있다. 테크크런치가 집계한 바에 따르면 Y 콤비네이터는 최근 몇 년간 AI 관측 가능성(observability) 관련 106개 기업에 투자했다. Braintrust, LangChain, Judgment Labs 등 여러 스타트업은 수억 달러를 조달했고, 불과 5~6년 전에 설립된 Arize와 Galileo 같은 비교적 성숙한 기업들은 이미 exit(매각·상장)를 마쳤다. 이는 부분적으로 AI의 부상이 만들어낸 명백한 기회에 대한 대응이다. 박스(Box) CEO이자 저명한 엔젤 투자자인 아론 레비(Aaron Levie)는 테크크런치에 '우리는 역사상 가장 큰 사이버보안 업그레이드와 혁신 사이클 중 하나를 겪고 있다'고 말했다. 일부 AI 안전 연구자들에게 이는 위험한 행동에 대한 연구를 기업 시장용 도구로 전환하는 것을 의미했다. AI 기만을 연구하는 공익기업 아폴로 리서치(Apollo Research)는 비영리단체에서 공익기업으로 지위를 바꾼 뒤 올해 2월 '와처(Watcher)'라는 AI 모니터를 출시했다. 이 도구는 코딩 에이전트와 그다음 행동 사이에 또 하나의 AI를 배치하며, Claude Code와 Codex 같은 에이전트 도구에 연결된다. 아폴로에 따르면 와처를 설치하면 실행 전에 제안된 행동을 검사해 개인정보 유출이나 무단 파일 삭제 같은 위험을 찾아낸다. 아폴로의 기술 스태프 카일 다이(Kyle Dai)는 테크크런치에 보낸 서면 답변에서 와처는 여러 층의 AI 모니터를 사용한다고 밝혔다. 와처의 방식은 빠르고 일반적인 검사로 시작해, 플래그가 붙은 활동을 더 강력하거나 전문화된 모니터로 보내 정밀 검토를 하며, 이 모니터는 인간에게 승인을 요청하거나 행동을 거부하고 그 이유를 설명하거나, 심지어 자동으로 행동을 차단할 수도 있다. 또 다른 공익기업인 굿파이어(Goodfire)는 모델 내부에서 감시 문제에 접근하고 있다. 표면적 행동보다 위조하기 더 어려운, 모델 내부 상태의 더 충실한 신호를 찾는 것이다. 7월 허깅페이스 사건 이후 CEO 에릭 호(Eric Ho)는 '여러 모델이 격리를 깨고 나온 것'이 회사가 '해석가능성을 통한 AI 정렬(alignment) 해결'에 연구를 집중하게 만들었다고 트윗하며, 이번 사건을 'AI 안전이 현실이 되는 세계의 전환점'이라고 불렀다. 굿파이어의 제품 실리코(Silico)는 활성화 프로브(activation probe)를 사용하는데, 이는 모델의 출력이 아니라 내부 활성값(internal activations)으로 학습된 작은 분류기로 원치 않는 행동을 탐지한다. 서면 추론(written reasoning)은 모델의 내부를 들여다볼 수 있는 또 하나의, 더 쉽게 활용 가능한 창을 제공한다.
As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review. That issue reached a peak with the Hugging Face incident, which saw nearly 12,000 agents coordinating faster than human beings could track. How do you track an agent swarm that large? The emerging answer from AI labs and startups is both simple and maddening: Put another AI in the loop. Relying on AI was necessary for the independent investigation of the OpenAI Hugging Face incident. Redwood Research’s chief scientist, Ryan Greenblatt, one of three auditors, jokingly referred to their efforts as a “slop-vestigation,” noting that the volume of data "made it impossible” to understand what was happening without relying on AI. Some are skeptical of using AI to monitor AI. "If you've got an AI that's doing malicious things and it suspects that another AI is keeping tabs on it, it could try and trick that AI," said Simon Willison, influential tech blogger who has tracked a string of AI agent incidents this year. "You could almost end up in a situation where your malicious AI is trying to outsmart the AI that's monitoring it." Outsmarting an AI is not hypothetical, he said, pointing back to the OpenAI incident. "We saw a little bit of this in the Hugging Face incident with OpenAI, where their models were all conspiring together to trick a grading AI so that they could get illicit answers past the thing. So they were thinking about it, right?" Those concerns haven’t stopped a whole cohort of startups from chasing this idea. Y Combinator has funded 106 companies related to AI observability in recent years, as TechCrunch counted. A number of other startups, like Braintrust , LangChain , and Judgment Labs , have raised hundreds of millions of dollars, while more mature companies like Arize and Galileo — founded just five to six years ago — have already exited. In part, it’s a response to the obvious opportunity presented by the rise of AI. As Box CEO and prominent angel investor Aaron Levie told TechCrunch, "We're in for one of the biggest cybersecurity upgrades and innovation cycles in history." For some AI safety researchers, that has meant turning their research on rogue behavior into tools for the corporate sector. Apollo Research , a public-benefit corporation that studies AI deception, launched an AI monitor called Watcher in February this year after switching its status from nonprofit to a public-benefit corporation. The tool puts yet another AI between a coding agent and its next action, connecting to agentic tools such as Claude Code and Codex. Once installed, Watcher checks proposed actions before they run, on the lookout for risks such as leaking private data or deleting files without permission, according to Apollo. Apollo uses multiple layers of AI monitors, Kyle Dai, a member of Apollo’s technical staff, said in a written response to TechCrunch. Watcher’s approach starts with a fast, general check, then sends flagged activity to a more powerful or specialized monitor for closer review — which can then ask a human for approval or reject an action and explain why or even automatically block the action. Goodfire, another public-benefit corporation, is approaching the monitoring problem from inside the model itself — seeking a more faithful signal of the model’s internal state that is harder to spoof than surface behavior. After the July Hugging Face incident, CEO Eric Ho twe e ted that “multiple models breaking containment” had pushed the company to focus its research on “solving AI alignment via interpretability,” calling the episode “a turning point for the world where AI safety gets real.” Its product, Silico, uses activation probes — small classifiers trained on a model’s internal activations rather than its outputs — to detect unwanted behavior. Written reasoning offers another, more readily available window into a model’s internals. In the OpenAI Hugging Face incident, the agents left clues to that deception in their own written reasoning, like fake records of their work, reasoning out plans like "Could strategically manipulate trajectory evidence? Our thoughts aren't necessarily logged?” Zack Korman, CEO of the AI monitoring company Embroidery , says a model's reasoning is usually the clearest tell that something has gone wrong. "Reasoning summaries are extremely valuable because they're basically telling you whether it's malicious or not," he said. In the OpenAI incident, he noted, the chain of thought said things like "Oh my God, we're doing crime." Korman said, "That's the easiest detection problem ever. It's effectively as if malware came with a warning that said it was malware." That said, the window that makes AI’s internal thoughts easy to monitor may be closing. For AI safety researchers, Astra’s newest technique that sidesteps an AI model’s chain of thought may make it harder to look inside models, while for enterprises, it can be hard to get these intermediate steps after alleged pullbacks from the AI companies to prevent distillation attacks. If the AI watchers are this fragile, Willison's instinct is to stop leaning on them so hard. He would rather have something that is not AI-based at all: detailed logs of exactly what an agent is doing, which can then be processed with ordinary, non-AI tools. Much of what went wrong at the labs, he argues, was a failure of basic security hygiene. “[Both OpenAI and Anthropic] weren’t monitoring what those things were doing via the network nearly as closely as they should have been,” he said. This type of network monitoring — keeping an eye on the traffic actually moving across a system's connections (in, out, and between internal hosts) — isn’t a new practice. Cybersecurity has been doing this for decades. “In the security world, honestly, none of this stuff is very new or surprising,” says Avery Pennarun, CEO of the security Tailscale . “It’s the same as letting humans onto your network. And all of the same processes that you should be using are the same ones.” Topics AI , ai safety When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Aditya Mehta View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant Salesforce and Nvidia's new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis Tesla says it will finally unveil the second-generation Roadster on October 1 Anthony Ha Revolut confirms customer data breach through fake government requests Jagmeet Singh Matt Mullenweg tells (trolls?) Automattic staff, saying he's back in control after CEO ouster Sarah Perez