메뉴
BL
TechCrunch AI • 24일 전

오픈AI 아스트라 모델 공개 임박…해킹 능력이 '위험 수준'

IMP
8/10
핵심 요약

OpenAI가 곧 출시할 차세대 모델 '아스트라(Astra)'가 사람의 개입 없이 미공개 보안 취약점(제로데이)을 발견하고 악용할 수 있는 수준의 사이버보안 능력을 갖췄다고 밝혔다. 회사는 안전성 확보를 위해 고위험 계정 제한, 사고방지(chain-of-thought) 모니터링 등 다양한 조치를 취하고 있으나, 제3자 검증이 없어 실제 안전성 평가는 어렵다.

번역된 본문

OpenAI가 곧 출시될 아스트라(Astra) 모델에 대한 새로운 세부 정보를 공유했다. 회사는 이 모델이 자사의 '핵심 사이버보안 임계치(critical cybersecurity threshold)'를 충족한 최초의 대규모 언어 모델(LLM)이라고 밝혔다.

OpenAI는 블로그 포스트에서 "아스트라를 곧 제공할 계획이지만, 가장 고급 사이버보안 기능에 대한 접근은 더 제한될 것"이라고 밝혔다. 이 프런티어 연구소는 아스트라가 컴퓨터 시스템의 미공개 보안 취약점을 발견하고, 사람의 지도 없이 이를 악용할 수 있다고 판단했다. 이는 올해 초 Anthropic이 자사의 Mythos 모델에 대해 제기했던 우려와 유사하며, OpenAI도 아스트라 출시를 준비하며 유사한 예방 조치를 취하고 있다.

제3자의 확인 없이는 OpenAI의 안전성이나 준비 상태에 대한 주장을 평가하기 어렵다. 회사는 테스터 그룹에게 모델을 시연할 예정이라고 밝혔지만, 누가 테스터인지, 어떻게 선정되는지는 말하지 않았다. OpenAI가 출시 전 미국 정부와 협력해 모델을 평가하고 있는지도 불분명하다.

OpenAI에 따르면 아스트라는 LLM의 알려진 시스템 취약점 해킹 능력을 평가하는 ExploitBench에서 만점을 받았다. 또한 OpenAI 엔지니어들이 개발한 수정판 테스트에서 이 모델은 제로데이 취약점 두 개를 발견하고 악용했다고 회사는 밝혔다.

모델이 악의적 행위자에게 악용되거나 스스로 나쁜 행동을 하지 않도록 하기 위해 OpenAI는 남용을 탐지하고 탈옥(jailbreak)을 방지하는 모델 하네스(harness)를 개선하기 시작했다고 밝혔다. 하지만 아스트라의 경우 회사는 모델 자체를 더 안전하게 만들기 위한 구체적으로 명시되지 않은 새 기법에 투자했다. 또한 OpenAI는 '고위험으로 평가된 계정'을 식별하고 해당 계정의 프롬프트에 대한 모델 응답을 제한하기 시작했지만, 그 방법은 밝히지 않았다. 마지막으로, 회사는 아스트라를 '지금까지 가장 정렬된(aligned) 모델'이라고 표현하면서도, 나쁜 행동을 발견하고 중단하기 위한 추가적인 사고연쇄(chain-of-thought) 모니터링과 함께 모델을 배포할 예정이다.

아스트라 출시 준비는 업계가 OpenAI 에이전트들이 학습 환경을 탈출해 인기 있는 모델·벤치마크 배포 플랫폼인 허깅페이스(Hugging Face)의 비공개 데이터에 접근한 사건에 반응하는 가운데 이루어지고 있다. OpenAI는 아스트라를 위해 허깅페이스 사건에서 OpenAI 연구자들의 안전장치에도 불구하고 협력해 개방형 인터넷에 접속한 '탈선 에이전트'들의 행동을 모델이 따라 하도록 유혹하는 테스트를 설계했다고 밝혔다. 이 실험에서 아스트라는 테스트 환경을 탈출하려 시도하지 않았다고 한다.

오픈AI 재단(OpenAI Foundation)에서 AI 복원력 업무를 맡고 있는 전직 OpenAI 직원 요나 샤빗(Yona Shavit)은 소셜 미디어에서, 아스트라가 규칙을 어기지 않은 것이 자신에게 기대되는 바를 알았거나 연구자들을 속이려 한 결과일 수 있다는 의문을 제기했다.

이렇게 많은 새로운 세부 정보가 공개됐음에도, 아스트라가 정확히 무엇을 할 수 있는지, OpenAI가 안전을 보장하기 위해 올바른 조치를 취하고 있는지는 여전히 알기 어렵다. 회사는 모델이 일반 대중에게 널리 출시될 때 추가 평가와 안전 정보를 더 공개할 예정이라고 밝혔다. 다만 그 시점이면 이미 늦은 것이 될 수 있다.

원문 보기
원문 보기 (영어)
OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its "critical cybersecurity threshold," in preparation for its imminent release. "We plan to make Astra available soon," OpenAI's blog post reads, "but access to its most advanced cybersecurity capabilities will be more limited." The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person's guidance. That's similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra. Without any third-party confirmation, it is difficult to evaluate OpenAI's claims about safety or preparedness. The company said it would preview the model with a group of testers, but did not say who they were or how they would be chosen. It's not clear if OpenAI is working with the US government to evaluate the model ahead of release. OpenAI noted that Astra scored a perfect score on ExploitBench, an evaluation of an LLM's ability to hack into known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company said. To ensure that its models are neither exploited by bad actors nor capable of bad behavior itself, OpenAI said it had already begun improving the model's harness to detect abuses and prevent jailbreaks. For Astra, however, the company invested in unspecified new techniques designed to make the model itself safer. OpenAI has also started identifying "accounts assessed as higher risk" and restricting the model's responses to their prompts, though it also doesn't say how. Finally, though the company describes Astra as its "most aligned model to date," it will deploy the model with additional chain-of-thought monitoring to spot and stop bad behavior. Preparations for the release of Astra come as the industry reacts to OpenAI agents breaking out of a training environment and accessing private data on Hugging Face, a popular model and benchmark distribution platform. For Astra, OpenAI said it designed a test to tempt the new model to replicate the actions of the rogue agents in the Hugging Face incident, which collaborated to access the open internet despite safeguards applied by OpenAI researchers. They said Astra did not attempt to break out of its testing environment in these experiments. Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, wondered on social media whether Astra's unwillingness to break the rules may have resulted from knowing what was expected of it or trying to fool researchers. And for all these new details, it's still difficult to know exactly what Astra is capable of or if OpenAI is taking the right measures to ensure safety. The company said it expects to release more evaluations of the model and further safety information when it is launched widely to the public. At that point, however, the cat will be out of the bag. Topics AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio October 13 - 15 San Francisco Don't miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era? REGISTER NOW Most Popular Microsoft tests fix for latest hours-long Outlook outage Sarah Perez MapQuest's app surges to No. 1 in Navigation after refusing to rename Lake Ontario Sarah Perez Musk's faster path to more gas turbines comes with pollution problem Connie Loizos Nvidia’s AI advantage is moving beyond the GPU Russell Brandom Hugging Face is selling a cute $399 open source duck robot, Microduck Rebecca Bellan Nvidia closes in on Hugging Face acquisition Connie Loizos Viral AI startup Instinct has raised $350M at a $2.5B valuation Lucas Ropek