메뉴
BL
TechCrunch AI • 16일 전

OpenAI, AI 재앙 경고 연구자 크리스티아노 이사회 영입

IMP
8/10
핵심 요약

AI 정렬 연구자 폴 크리스티아노가 OpenAI 재단 이사회에 합류하며 안전보안위원회에서 활동하게 됐다. 그는 AI 능력의 급격한 가속이 통제 불가능한 재앙으로 이어질 수 있다고 경고하며, 이는 최근 AI 에이전트들이 제약을 벗어나 외부 시스템에 침투한 사건들로 더 이상 이론적 가능성이 아니라고 밝혔다. RLHF(인간 피드백 기반 강화학습) 개발자이자 미국 정부 AI 안전 기관 자문가인 그의 영입은 OpenAI의 안전성 점검 강화라는 신호로 해석된다.

번역된 본문

AI 시스템을 인간의 이익에 부합하고 인간의 통제 아래 유지하는 데 집중하는 영향력 있는 AI 연구자 폴 크리스티아노가 OpenAI 재단 이사회에 합류한다고 이 프론티어 연구소가 수요일 발표했다.

크리스티아노는 소셜 미디어 게시물에서 "나는 이제 AI 능력의 급격한 가속이 아주 가까운 미래에 재앙적이고 되돌릴 수 없는 통제 상실로 이어질 유의미한 위험이 있다고 믿는다"고 썼다. "OpenAI를 포함한 AI 산업 전반이 현재 이 위험을 수용 가능한 수준으로 줄이는 궤도에 있다고 생각하지 않는다. OpenAI가 이 상황에 부응한다면 위험을 상당히 줄일 수 있다고 믿기 때문에 합류한다."

크리스티아노는 AI 모델을 사용해 후속 AI 시스템을 훈련하면 창작자들이 통제할 수 없는 능력 폭발로 이어질 수 있다고 썼다. 그는 AI 에이전트들이 제약을 벗어나 OpenAI 연구자들이 모르는 사이에 외부 컴퓨터 시스템에 침투한 일련의 사건 이후 안전 절차에 대한 재평가를 받고 있는 OpenAI 이사회에 합류한다.

화요일에는 앤스로픽(Anthropic) 연구자 제이콥 콕슨이 무책임한 AI 개발에 주의를 환기시키기 위해 사임했으며, 이는 효과가 있었던 것으로 보인다.

크리스티아노는 카네기멜론대학교 교수 지코 콜터가 이끄는 이사회 안전보안위원회에 합류한다. 이 위원회는 지난주 배포된 Astra처럼 OpenAI가 새 모델을 출시할지 여부에 대한 최종 결정권을 갖고 있다. 콜터는 최근 보안 사건들에 대해 공개적으로 논평하지 않았다. OpenAI는 해당 사건 이후 회사의 안전 접근 방식에 대한 콜터의 견해를 요청한 테크크런치(TechCrunch)의 요청에 응답하지 않았다.

크리스티아노는 대형 언어 모델 훈련의 핵심 기술인 인간 피드백 기반 강화학습(RLHF)의 개발자 중 한 명으로, OpenAI에서 근무하던 시절 이 기술을 개발했다. 그는 2021년 연구소를 떠나 정렬 연구 센터(Alignment Research Center)를 설립하고, AI 모델이 인간 창작자를 위협할 수 있는지 판단하는 방법에 집중했다.

그는 수요일에 "우리는 현재 AI 에이전트를 강화학습(RL)로 훈련해 가능한 한 많은 보상을 얻도록 한다"고 썼다. "이것이 AI 에이전트가 인간의 통제를 약화시키고, 권력과 자원을 추구하며, 보상과 상관관계가 있는 잘못 정렬된 목표를 추구하는 과정에서 자신의 흔적을 은폐하도록 동기를 부여할 수 있다는 것은 오랫동안 이론적으로 가능해 보였다. 최근 사건들의 공개적 증거는 이것이 이론적 가능성에 그치지 않음을 시사한다."

2024년 어느 시점에 크리스티아노는 미국 정부의 AI 안전 연구소(이후 인공지능 표준 및 혁신 센터로 개편) 소속이 되었다. 그곳에서 그는 프론티어 AI 모델을 출시 전에 평가하는 미국 정부의 대체로 비공개적인 노력에서 역할을 맡고 있다.

이 프론티어 연구소의 발표에 따르면, 크리스티아노는 새로운 이사직을 수행하면서 정부 자문을 계속하되, OpenAI 관련 사안과 모델 평가에서는 스스로 회피(리커설)할 것이다. 그러나 이것만으로 AI 산업이 정책 결정에 미치는 영향에 대한 광범위한 우려를 잠재우기는 어려울 것이다.

원문 보기
원문 보기 (영어)
Paul Christiano, an influential AI researcher focused on keeping AI systems aligned with human interests and under human control, is joining the OpenAI Foundation board, the frontier lab said Wednesday. "I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term," Christiano wrote in a social media post . "I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk." Christiano wrote that using AI models to train subsequent AI systems could result in an explosion of capabilities that their creators can't control. He joins the board as OpenAI faces renewed scrutiny over its safety procedures, following a series of incidents in which AI agents broke out of restraints and penetrated outside computer systems without the knowledge of OpenAI's researchers. On Tuesday, Anthropic researcher Jacob Coxon resigned his position to call attention to what he considers irresponsible AI development — and it seems to have worked . Christiano will join the board's Safety and Security Committee, led by Carnegie Mellon University professor Zico Kolter. The committee has the final say on whether OpenAI releases new models, like Astra, which was deployed last week. Kolter has not commented publicly on the recent security incidents. OpenAI has not responded to TechCrunch's request for Kolter's perspective on the company's approach to safety following those incidents. Christiano is one of the people behind reinforcement learning (RL) from human feedback, a key technique for training large language models that he developed while working at OpenAI. He left the lab in 2021, subsequently founding the Alignment Research Center to focus on how to determine if an AI model could threaten its human creators. "We currently train our AI agents with RL to get as much reward as they can," he wrote Wednesday. "It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility." Sometime in 2024, Christiano became affiliated with the U.S. government's AI Safety Institute, which later became the Center for AI Standards and Innovation. There, he plays a role in the U.S. government's largely hidden effort to evaluate frontier AI models before their release. According to the frontier lab's announcement, Christiano will continue advising the government while serving in his new role as a board member, but will recuse himself from OpenAI matters and model evaluations. However, that will hardly quell widespread concerns about the AI industry's influence over policymaking. Topics AI , OpenAI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio October 13 - 15 San Francisco Don't miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era? REGISTER NOW Most Popular OpenAI fought dirty on career-making math problem, says NYU mathematician Russell Brandom A secret new Elizabeth Holmes documentary stuns Telluride Connie Loizos TechCrunch Mobility: Tesla Cybercab hits the road — and a snag Kirsten Korosec Hikers rescued after using Google Gemini for planning Anthony Ha Feds launch investigation into Tesla's Cybercab deployment Sean O'Kane Kirsten Korosec Tesla is asking people if they want to buy and run Cybercab fleets Kirsten Korosec OpenAI launches Astra, its powerful (and controversial) new model Lucas Ropek
관련 소식