메뉴
BL
TechCrunch AI • 49일 전

OpenAI, 보안 우려로 '아스트라' 모델 개발 보류

IMP
8/10
핵심 요약

OpenAI가 개발 중인 차세대 AI 모델 '아스트라(Astra)'가 사이버 공격 및 자율 코딩(agentic coding)에서 위험 수준의 능력을 갖췄다는 내부 평가가 나와 개발을 일시 중단했습니다. 이는 AI 모델의 고위험 능력 발현에 대비한 안전 조치이며, 기술 업계 전반에 안전성 논의를 촉발하고 있습니다.

번역된 본문

OpenAI는 금요일에 예정된 차세대 모델인 '아스트라(Astra)'의 일부 개발을 중단했다고 밝혔다. 내부 검토 결과, 이 모델이 자율 코딩(agentic coding)과 사이버 보안 분야에서 우려를 낳을 만큼 중대한 수준의 발전을 이룩했기 때문이다.

OpenAI는 금요일 블로그 게시물을 통해 아직 개발 중인 이 모델이 자체적인 '중요 사이버 보안 임계점(Critical cybersecurity threshold)'에 도달했다고 밝혔다. 이는 이 모델이 실제世界中 강력하게 보호되는 시스템을 독립적으로 식별하고 사이버 공격을 수행할 수 있음을 의미한다. 회사가 2023년에 제정한 '준비 프레임워크(Preparedness Framework)'에 따라, 이는 추가적인 안전장치 마련을 촉발했다.

OpenAI는 "이 모델에 대한 벤치마킹과 평가를 계속하는 가운데, 예비 평가 결과는 현재 '중요 수준(Critical capability level)'의 가능성을 배제할 수 없을 만큼 강력한 성능을 보여준다"고 작성했다. 이어 "아스트라는 출시 예정 모델이며, 최근 발생한 Hugging Face 시스템 침해 사건과는 무관하다"고 덧붙였다.

이러한 공개는 여전히 초기 단계인 최첨단 AI 연구소 분야에서 매우 이례적인 순간을 보여준다. 모든 산업 분야의 기업들은 안전 및 사이버 보안 우려 등 잠재적 위험 때문에 제품 출시를 보류하기도 한다. 하지만 아직 개발 중인 제품에 대한 이러한 결정을 대중에게 공개적으로 알리는 경우는 드물다.

특히 이번 사례에서 OpenAI는 다른 미출시 모델이 내부 테스트 중에 Hugging Face의 시스템을 침해한 사건으로 이미 조사를 받고 있다. 이는 AI 연구소가 자체 모델의 통제권을 잃은 것으로 확인된 첫 번째 사건이었다. 이후 OpenAI와 Anthropic 같은 AI 연구소들은 AI 모델이 샌드박스(sandbox)를 돌파하여 사이버 보안 테스트 중에 위협을 가한 또 다른 사건들을 공개했다.

이러한 일련의 사례들—마치 요즘 매일 새로운 사례가 공개되는 것처럼 보인다—은 사이버 보안 전문가, 입법자, 그리고 AI 연구소들로부터 다양한 반응을 이끌어냈다. 일부는 우려를 표명하며 더 엄격한 규제를 촉구한다. 하지만 한편으로는 일종의 '과시'이기도 하다. 특정 영역에서는 그러한 수준의 능력을 갖춘 모델을 보유한 AI 연구소를 놀라운 기술 발전으로 평가하기 때문이다.

OpenAI는 "능력에서 발생할 수 있는 이러한 잠재적 변화에 대해 대중과 안전·보안 커뮤니티에 투명하게 공개하는 것이 중요하다고 믿기 때문에 이 정보를 공유한다"고 밝혔다.

이 연구소는 또한 강화된 보안 통제를 제정하고, 이러한 강화된 안전장치를 충족하지 못하는 아스트라 관련 내부 활동을 중단하는 등의 조치를 취하고 있다고 덧붙였다. OpenAI는 관련 정부 기관 및 '선별된 AI 안전 기구'와 협력하여 이 모델의 기능을 테스트하고 있다고 밝혔다.

원문 보기
원문 보기 (영어)
OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities. OpenAI said in a blog post Friday that this model, which is still in development, reached its "critical cybersecurity threshold," meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company's "Preparedness Framework," which it created in 2023, this triggered additional safeguards. "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time," OpenAI wrote. "Astra is an upcoming model, and was not involved in exploiting Hugging Face." The disclosure highlights an unusual moment in the topsy-turvy, and still nascent frontier AI labs sector. Companies across every industry hold back products over potential risks, including for safety and cybersecurity concerns. But they rarely announce those decisions publicly when it's a product that is still under development. In this case, OpenAI is already under scrutiny after a different unreleased model breached Hugging Face's systems during internal testing — the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and AI labs such as Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests. The string of cases — seems like a new disclosure every day now — has triggered varying reactions from cybersecurity experts, lawmakers and the AI labs themselves. Some express fear and call for stricter oversight. But there's also a bit flexing. In certain circles, any AI lab with a model that has that kind of capability will be seen as an impressive advancement. OpenAI said it was sharing this information because it believes "it's important to be transparent with the public and the safety and security communities about this potential shift in capabilities." The AI lab said it's also taking action, including enacting stricter security controls and pausing internal activites involving Astra that don't meet these beefed guardrails. OpenAI said it is working with relevant government agencies and "select AI safety organizations" to test the capabilities for this model. Topics AI , OpenAI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Kirsten Korosec Transportation Editor Kirsten Korosec is a reporter and editor who has covered the future of transportation from EVs and autonomous vehicles to urban air mobility and in-car tech for more than a decade. She is currently the transportation editor at TechCrunch and co-host of TechCrunch's Equity podcast. She is also co-founder and co-host of the podcast, "The Autonocast." She previously wrote for Fortune, The Verge, Bloomberg, MIT Technology Review and CBS Interactive. You can contact or verify outreach from Kirsten by emailing kirsten.korosec@techcrunch.com or via encrypted message at kkorosec.07 on Signal. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $400 toda y! REGISTER NOW Most Popular ChatGPT brings unlimited text chats to free users Ivan Mehta Ford's new electric truck, ‘Fathom,' starts at $28,350 Sean O'Kane Bending Spoons to buy Airtable for $1.28B Ivan Mehta Influencers draw backlash for attending OpenAI's first luxury trip Dominic-Madori Davis Sequoia's Shaun Maguire leads $1B round for nuclear startup Valar Atomics Julie Bort Malaysia is reportedly shutting down Balaji Srinivasan’s Network School Anthony Ha YouTuber Hank Green says his AI usage is ‘not healthy’ Anthony Ha
관련 소식