메뉴
BL
Wired AI 48일 전

앤스로픽, AI 연구원들 몰래 성능 저하시킨 정책 철회

IMP
8/10
핵심 요약

앤스로픽(Anthropic)이 자사의 새로운 AI 모델을 사용해 경쟁사 AI를 개발하려는 연구자들에게 사용자가 모르게 성능을 저하시키려던 정책을 AI 연구계의 거센 비판을 받고 철회했습니다. 이 조치는 오픈소스 생태계 및 제3자 평가 기관의 연구를 방해하고, 소수 빅테크 기업만이 고급 AI 연구를 독점할 수 있다는 우려를 낳았기 때문에 매우 중요합니다. 이에 앤스로픽은 공식 사과하며, 앞으로는 AI 개발 관련 제한을 가할 경우 이를 사용자에게 투명하게 알리겠다고 밝혔습니다.

번역된 본문

앤스로픽(Anthropic)이 경쟁사들이 자사의 새로운 AI 모델인 '클로드 페이블 5(Claude Fable 5)'를 사용해 다른 AI 모델을 개발하는 것을 은밀하게 제한하려 했던 정책을 철회했습니다. 이 조치는 AI 연구 커뮤니티로부터 상당한 비판을 받은 후 수정되었습니다.

앤스로픽은 WIRED에 보낸 성명에서 "프론티어 LLM(대형 언어 모델) 개발에 대한 클로드 페이블 5의 안전장치를 사용자가 볼 수 있도록 투명하게 변경하겠다"며 "우리는 잘못된 트레이드오프를 선택했고, 균형을 맞추지 못한 점에 대해 사과한다"고 밝혔습니다.

앤스로픽은 이번 주 초 오용을 방지하기 위해 추가적인 안전 가드레일을 적용한 최신 AI 모델 버전인 클로드 페이블 5를 공개했습니다. 앤스로픽이 결정한 안전 조치 중 일부는 예상대로였습니다. 회사는 고도화된 AI를 사이버 공격에 악용하거나 생물무기를 제조하는 것을 방지하기 위해, 사이버 보안, 생물학 또는 화학에 대해 질문하는 사용자를 성능이 낮은 AI 모델로 우회(Reroute)시킨다고 발표했습니다.

하지만 클로드 페이블 5를 프론티어 AI 개발에 사용하려는 연구원들에 대해서는 다른 접근 방식을 취했습니다. 사용자가 알아채지 못하는 방식으로 의도적으로 모델의 성능을 저하시킨 것입니다. 이 조치는 서비스 약관에서 명시적으로 금지하고 있는, 클로드를 사용해 경쟁사 AI 모델을 학습시키려는 연구자들을 효과적으로 방해(Sabotage)하는 결과를 낳습니다.

현재 앤스로픽은 방침을 바꾸어, 클로드 페이블 5의 AI 개발 관련 안전장치가 사용자에게 명확하게 보일 것이라고 밝혔습니다. 회사가 사용자가 클로드를 사용해 고성능 AI를 구축하려 한다고 의심할 경우, 요청을 거부하거나 성능이 낮은 모델로 우회시킨다는 사실을 경고할 것입니다.

앤스로픽은 AI 연구 커뮤니티의 맹렬한 비판을 받은 후 이 정책을 번복했습니다. 앤스로픽은 이미 경쟁사가 클로드를 사용해 오픈소스 및 클로즈드 소스 AI 모델을 구축하는 것을 제한하는 조치를 취해왔지만, 특정 사용자의 모델 성능을 조용히 저하시키는 것은 너무 비난받을 만한 일이라는 비판이 나왔습니다. 클로드의 코딩 에이전트는 오픈소스 AI 연구 프로젝트를 진행하는 개발자들을 포함해 개발자들 사이에서 선호되는 도구가 되었습니다. 연구자들은 회사의 최신 정책이 소수의 선도적인 AI 연구소만이 고급 AI 연구를 수행할 수 있는 불안한 미래를 초래할 수 있다고 WIRED에 전했습니다.

미국혁신재단(Foundation for American Innovation) 선임 연구원이자 전 백악관 AI 자문관이었인 딘 발(Dean Ball)은 X(옛 트위터) 게시물에서 "사용자에게 알리지 않고 머신러닝(ML) 연구 성능을 저하시키는 것은 충격적으로 적대적이며 매우 좋지 않은 인상을 준다"고 지적했습니다. 그는 또 다른 게시물에서 이러한 '은밀한 방해(Secret sabotage)' 정책은 AI 연구자들이 AI 안전성에 대해 협력하는 것을 제한하기 때문에 앤스로픽의 전반적인 입장을 훼손한다고 덧붙였습니다.

오픈소스 AI 스타트업 프라임 인텔렉트(Prime Intellect)의 연구 책임자인 윌 브라운(Will Brown)은 "앤스로픽이 대중에게 '우리는 다른 누구도 AI 연구를 하는 것을 신뢰하지 않는다. 오직 우리만이 AI 연구를 해야 한다'고 말하는 것 같았다"며 "마치 그들이 뒤에 오는 사람을 위해 사다리를 걷어차고 올라가는 것과 같은 느낌을 준다"고 말했습니다.

브라운은 또한 앤스로픽의 안전장치가 작동할 때 이를 알리지 않기 때문에 이 정책으로 인해 개발자들은 자신이 앤스로픽의 규정을 위반했는지조차 모르는 상태가 될 것이라고 말했습니다. 그는 이러한 제한이 광범위한 파급 효과를 가져올 수 있다고 덧붙였습니다. 예를 들어, 그는 프론티어 모델의 안전성, 성능 및 신뢰성을 테스트하는 제3자 평가 기업 생태계가 점점 성장하고 있는데, 앤스로픽이 모델을 은밀하게 저하시켰다면 이러한 중요한 작업이 방해를 받았을 것이라고 지적했습니다.

앤스로픽은 클로드가 AI 연구를 가속화하는 데 점점 더 효과적으로 발전해 왔기 때문에 이러한 조치를 시행했다고 밝혔습니다. 회사는 최근 블로그 게시물에서 AI가 사회가 적응하는 속도보다 빠르게 자체 기능을 향상시킬 수 있다고 우려하고 있습니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Anthropic is backtracking on a policy that would have covertly limited competitors from using its new AI model, Claude Fable 5 , to develop other AI models. The company changed course after the move received significant backlash from the AI research community . “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible,” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Anthropic released Claude Fable 5, a version of its latest AI model with additional safety guardrails designed to prevent misuse, earlier this week. Some of the safeguards Anthropic decided on were unsurprising: The company said it would reroute users who asked questions about cybersecurity, biology, or chemistry to a less capable AI model to reduce the chances of someone using the advanced AI to carry out a cyberattack or build a bioweapon. But for researchers trying to use Claude Fable 5 for frontier AI development, Anthropic outlined a different approach. The firm would deliberately degrade the model’s performance in ways that were invisible to the user. The move would effectively sabotage researchers trying to use Claude to train competing AI models, which Anthropic explicitly bans in its terms of service . Got a Tip? Are you a current or former Anthropic employee who wants to talk about what's happening? We'd like to hear from you. Using a nonwork phone or computer, contact the reporter securely on Signal at mzeff.88. Anthropic now says it’s changing course, and that Claude Fable 5’s safeguards for AI development will be visible to users. If the company suspects a user is trying to use Claude to build a highly capable AI it will alert them that it’s either refusing the request, or rerouting the user to a less capable model. Anthropic reversed the policy after it received fierce backlash from the AI research community. Anthropic has already taken steps to limit competitors from using Claude to build closed and open source AI models, but critics say that quietly degrading the model’s performance for certain users went a step too far. Claude’s coding agent has become a favored tool among developers, including those working on open-source AI research projects, and researchers tell WIRED that the company’s latest policy could have led to a troubling future in which only a handful of leading AI labs could perform advanced AI research. Dean Ball, a senior fellow at the Foundation for American Innovation and a former advisor to the White House on AI, wrote in a post on X that “degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look.” He continued in another post that the “secret sabotage” policy undermines Anthropic’s overall stance, because it limits AI researchers from collaborating on AI safety. “It felt like Anthropic was saying to the public, ‘We don't trust anybody else to do AI research. We are the only ones who have to do AI research,” says Will Brown, research lead at the open source AI startup Prime Intellect. “It feels a bit like they’re starting to pull the ladder up behind them.” Brown said the policy would also have left developers in the dark about whether they were violating Anthropic’s rules, since the company wouldn’t alert them when its safeguards were triggered. He added that the restrictions could have had widespread consequences. For example, he pointed to the growing ecosystem of third-party evaluation firms that test frontier models for safety, performance, and reliability—work that could have been hindered if Anthropic secretly degraded its model. Anthropic said it implemented the measures because Claude has become increasingly effective at accelerating AI research. In a recent blog post , the company said it is concerned that AI could improve its capabilities faster than society can adapt to them. Anthropic argued that it would be “good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up.” “These safeguards prevent foreign adversaries from using our most capable models in ways that pose severe safety risks. The US and its allies hold an edge in frontier chips and the highly optimized software that runs them at full potential,” the company said in a statement to WIRED. “These safeguards ensure Claude isn't used to erode that advantage—by optimizing chips developed by those adversaries, for example […] In deciding whether to make them visible or invisible we faced a choice. A hidden safeguard is harder to probe and work around. This means the safeguards can be targeted much more narrowly.” Anthropic says that because this safeguard around AI development is now visible, it needs to cast a wider net, meaning more benign requests may trigger its safeguards. The company says it’s working to make its classifiers more precise as quickly as possible.