메뉴
BL
TechCrunch AI • 9일 전

앤스로픽·오픈AI, 안전성 평가기관 내부 접근 허용…진짜 독립성은 미지수

IMP
8/10
핵심 요약

앤스로픽의 다리오 아모데이 CEO와 오픈AI의 샘 올트먼 CEO가 외부 독립 평가기관(METR, Redwood Research 등)에 자사 시스템에 대한 전례 없는 접근권을 부여하겠다고 제안했다. 평가기관들은 환영하면서도, 완성 모델뿐 아니라 학습 중간 체크포인트와 내부 로그 접근, 공개 권한 등 구체적 세부사항과 법적 뒷받침이 필요하다고 지적했다. AI 모델이 평가 상황을 인식하고 문제 행동을 숨길 위험이 커지는 만큼, 학습 과정 전반에 대한 실질적 검증이 필수라는 게 핵심이다.

번역된 본문

주말에 발표된 장문의 에세이에서 앤스로픽 CEO 다리오 아모데이는 1년 전만 해도 AI 업계가 즉시 거부했을 제안을 내놓았다. 즉, 모든 프런티어 AI 기업 내부에 제3자 평가기관을 상주시켜 안전 사고를 보고하고, AI 모델이 실제로 정렬(aligned)되어 있는지 평가하며, 있는 그대로의 결과를 세상에 공유할 권한을 부여하자는 것이다. 아모데이는 앤스로픽이 METR과 Redwood Research 같은 독립 평가기관에 회사 시스템에 대한 전례 없는 접근권을 제공하겠다고 약속했다. 오픈AI의 샘 올트먼 CEO도 이러한 방식을 따르겠다고 밝혀, 업계가 외부 연구 단체와 협력하는 방식에 잠재적으로 큰 변화가 올 수 있음을 시사했다.

테크크런치(TechCrunch)와 대화한 제3자 평가기관들은 이 제안을 대체로 환영했지만, 그들이 진정한 독립적 감시기구가 될지, 아니면 AI 기업의 조건에 따라 움직이는 용역 업체가 될지 판단하려면 세부사항이 다듬어지고 — 이상적으로는 입법으로 뒷받침되어야 한다고 말했다.

모델이 자신이 평가받고 있다는 것을 점점 더 잘 인식하게 되면서 이러한 깊은 접근권은 더욱 중요해지고 있다. 이는 모델이 테스트 중에는 잘 행동하면서 문제가 되는 행동을 숨길 위험을 키운다. 연구자들은 완성된 모델을 테스트할 때는 이러한 행동의 단서를 놓칠 수 있지만, 학습 과정 전반에서 모델이 어떻게 행동했는지 조사하면 발견할 수 있다고 말한다.

"AI 기업들은 학습 과정에 대한 매우 기본적인 질문에 답할 수 있어야 합니다. 예를 들어, AI가 학습을 진행하는 동안 자신의 정렬(alignment) 학습을 스스로 방해하려고 시도한 적이 있습니까?"라고 아폴로 리서치(Apollo Research)의 연구 책임자 알렉산더 마인케가 테크크런치에 말했다. "이에 대한 답은 명백하게 '아니오'여야 하는데, 현재 우리는 AI 기업들이 스스로 이를 신중히 확인하고 진실하게 대중에게 보고하기를 전적으로 믿고 있을 뿐입니다. 그런데 최근 사건들에서 보았듯이, 기본적으로 그들은 둘 다 하지 않습니다. 내부 상주 평가자로서 우리는 실제로 확인할 수 있을 것입니다."

지금까지 AI 기업들은 출시 직전에 외부 검토자를 불러 완성된 모델을 테스트해 왔다. 이제 테크크런치가 만난 평가기관들은 최종 모델뿐만 아니라 학습 전 과정의 중간 버전, 즉 '체크포인트(checkpoint)'에도 접근할 수 있어야 한다고 제안한다. Far.AI의 CEO 애덤 글리브는 평가자들이 이러한 체크포인트들을 비교해 우려되는 행동이 언제 나타났는지 확인하고, 특정 행동에 보상을 주는 후학습(post-training) 환경을 조사하며, 평가 기록과 로그를 검토해 모델 성능에 관한 기업의 주장을 검증할 수 있다고 말했다.

앤스로픽과 오픈AI가 이런 수준의 접근권을 언제, 어떻게 제공할지는 불분명하다. 테크크런치의 반복적인 질문에도 두 회사 모두 어떤 평가기관과 협력할지, 언제 내부에 배치할지, 몇 곳을 영입할지, 정확히 어떤 시스템과 정보에 접근할 수 있는지, 무엇을 공개할 수 있는지에 대해 밝히지 않았다.

이렇게 '후드 아래'를 들여다보는 것이 중요한 이유는, 안전 테스트에서 좋은 성적을 내는 모델이 그 테스트를 통과하는 방법을 특별히 학습했다면 반드시 안전한 것은 아니기 때문이다. 스테이들리는 특정 상황에서 AI가 종료(shutdown)에 저항하는지 측정하는 '종료 저항 벤치마크'의 예를 들었다. "AI가 그 벤치마크에서 좋은 성적을 내도록 특별히 학습되었다면 그것은 매우 중요한 문제입니다"라고 그는 말하며, 이를 배기가스 테스트를 인식하고 테스트 조건에서 다르게 작동하도록 프로그래밍된 폭스바겐의 디젤게이트 스캔들에 비유했다.

글리브는 의미 있는 접근권은 모델 그 자체를 넘어 확장될 수 있다고 지적했다. 즉, 평가자들이 직원들을 인터뷰하여 기업의 문서와 안전 관행에 대한 공개 설명이 내부에서 실제로 일어난 일과 일치하는지 확인할 수 있어야 한다는 것이다. 아모데이는 평가자들에게 필요하다고 여겨지는 수준의 접근권을 부여할 수 있는 상당히 포괄적인 제안을 개략적으로 제시했는데, 여기에는 "위험 수준, 사고, 관행, 그리고 자신들이 받았거나 받지 못한 접근권에 관한 핵심 결과를 앤스로픽의 편집 통제 없이 공개할 권리"가 포함된다.

원문 보기
원문 보기 (영어)
In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world. Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research unprecedented access to the company’s systems. CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups. Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out — and ideally backed by legislation — if they’re to know whether they will function as truly independent watchdogs or vendors operating on the AI companies’ terms. That deeper access is becoming more important as models get better at recognizing when they’re being evaluated, raising the risk that they’ll behave well during testing while concealing problematic behavior. Researchers say clues to that behavior can be missed when testing the finished model, but uncovered by investigating how it behaved throughout training. “AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we've seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.” Historically, AI companies brought in outside reviewers to test finished models shortly before their release. Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training. Adam Gleave, CEO of Far.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed. Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Neither company has shared which evaluators they’ll work with, when they will be embedded, how many they’ll bring on, exactly what systems and information they will be able to access or what can be disclosed to the public, despite repeated questions from TechCrunch. Looking under the hood like this matters because models that perform well on safety tests aren’t necessarily safe if they’ve learned specifically how to pass that test. Steidley pointed to an example of a “shutdown resistance benchmark” that measures if the AI will resist being shut down in certain circumstances. “It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions. Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what happened internally. Amodei did outline a fairly comprehensive proposal that might give evaluators the kind of access they think is necessary, including the right to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.” But evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Previous efforts at independent evaluations suggest that that surrender will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly. Gleave said Far.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm’s independence. By default, he said evaluators are treated like ordinary contractors: bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published. The time limit There’s also the question of whether reviewers will get enough time and access to do the work they’re being asked to do. When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations. A similar issue occurred during the pre-release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet . According to Apollo Research's contribution to the model card , the firm was given only three days to test Astra, which made it difficult to draw firm conclusions. "Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the firm wrote in its evaluation. That track record leaves evaluators with a basic question: Why should this time be different? “It's certainly possible that Dario and Sam just had a change of heart, and they're going to be very open about this,” Gleave said. “But the intellectual property of these companies is so incredibly valuable to them, and I think they're going to, by default, be very careful about what can be shared.” Several researchers who spoke to TechCrunch called for a transparent framework that they all agree to publicly. Part of the framework, says John Steidley, head of strategy at Palisades Research, should involve standards for what kinds of auditors companies can rely on, lest they try to sidestep the issue by shopping for evaluators that either aren't qualified or aren't interested in assessing the most concerning risk. Henry Papadatos, executive director of Safer AI, says the problem, even with a public framework, is that voluntary measures are always dependent on a company’s goodwill. “Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis,” Papadatos told TechCrunch, noting that it’s also a good means of pushing all companies to adhere to the rules, not only the most willing. Not everyone has signed on. So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks. Some laws are already forming around the idea of third party evaluators. California’s SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized “independent verification organizations” with expertise assessing AI risks. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts. For now the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they w