메뉴
BL
Wired AI • 7일 전

AI 개발 속도 조절, 어떻게 실현할 수 있을까

IMP
7/10
핵심 요약

AI 연구자들이 AI가 위험해질 수 있다고 믿으면서도 개발 속도를 늦추는 구체적 방법은 아직 풀리지 않은 과제로 남아 있다. 토론토 대학의 새 보고서 'Pacing the Frontier'는 이를 연구 과제로 다뤄야 한다고 주장하며, 제3자 평가기관의 모델 검증과 감사 등이 유력한 방안으로 제시되고 있다. Anthropic은 AI가 자사 AI 연구의 26%를 수행하고 있다는 추적 기법을 발표하는 등 재귀적 자기개선(RSI)에 대한 우려가 커지는 상황이다.

번역된 본문

많은 AI 연구자들은 자신들이 개발 중인 기술이 언젠가 매우 위험할 수 있다고 굳게 믿는 듯하다. 그러나 이런 변덕스러운 알고리즘을 어떻게 통제할 것인지는 AI 기술 엘리트들 사이에서조차 불분명하다.

최근 몇 년간 연구자들은 AI가 위험해지는 것을 막기 위한 온갖 아이디어를 제시해왔다. 정부 규제 강화, 진전 상황을 측정하는 새로운 방법, 모델 내부 작동 방식 탐구 같은 비교적 논란이 적은 계획부터, GPU에 추적 장치를 심는 것, 심지어 대량의 AI 칩을 의례적으로 파괴하는 것 같은 기발한 제안까지 포함된다. 그러나 정치적·사회적 압박이 더 신중한 AI 개발 접근을 요구하는 지금도 AI를 안전하게 유지하는 방법은 여전히 불명확하다.

"우리는 이것을 연구 문제로 다루기 시작해야 합니다." 토론토 대학의 AI 연구자이자 'Pacing the Frontier, A Research Agenda' 보고서 공동 저자인 레이먼드 더글라스(Raymond Douglas)는 말한다. 이 보고서는 AI 개발 속도를 늦추는 것이 여전히 풀리지 않은 퍼즐이라고 경고한다. "우리는 어떤 선택지가 있는지, 그리고 그것이 어떤 효과를 낼지조차 제대로 이해하지 못합니다."

최근 몇 주간 AI 종말론 논의가 절정에 달했다. Anthropic의 한 연구자가 회사를 떠나며 몇 년 안에 AI가 인류를 멸망시키는 궤도에 오를 수 있다고 경고했고, Anthropic의 AI 안전 연구소 책임자가 즉각 그의 우려에 공감을 표했다. 미국 빅 AI 기업의 리더들—Anthropic의 다리오 아모데이, OpenAI의 샘 알트먼, SpaceXAI의 일론 머스크, Google DeepMind의 데미스 허사비스—도 모두 어떤 형태든 AI 개발 속도 조절이나 일시 중지에 지지를 표명했다.

AI 기업들이 AI 자체를 사용해 더 강력한 모델을 구축하고 있다는 점에서 이 문제는 특히 시급해 보인다. 이는 몇 년 안에 AI가 인간의 이해 능력을 뛰어넘는 가속화되는 재귀적 자기개선(RSI) 루프에 대한 공포를 불러일으키고 있다.

AI 연구소들은 이미 자체적인 새로운 접근법을 내세우고 있다. 이번 주 Anthropic은 AI가 얼마나 빠르게, 그리고 어쩌면 위험하게 발전하고 있는지 추적하는 여러 새로운 방법을 발표했다. 예를 들어 이 기법들은 Claude가 현재 Anthropic의 AI 연구 중 26%를 수행하고 있으며, 2026년 초에는 0%였음을 보여준다. 또한 Anthropic이 컴퓨팅 예산의 6%를 AI를 더 안전하게 만드는 방법 연구에 지출했음을 드러낸다.

그러나 더글라스와 다른 전문가들은 AI 개발을 효과적이고 신뢰할 수 있게 통제하려면 AI 연구소 외부의 자금과 전문성이 필요하다고 말한다. 이 최신 보고서와 그 외에서 제안된 해결책 중 일부는 다른 것보다 더 실현 가능해 보인다.

'독립적인' 평가자 AI 기업들이 자주 제시하는 아이디어 중 하나는 제3자 평가자에게 자사 모델에 대한 더 큰 접근 권한을 부여하는 것이다. 이 평가자들은 모델의 능력을 평가하고, 신뢰할 수 있는 환경 내에서 문제 행동을 유도하는 '레드팀' 작업을 수행한다.

영국 AI 보안 연구소(AI Security Institute)의 전 수석 과학자이자 이전에 Google DeepMind 연구원이었던 제프리 어빙(Geoffrey Irving)은 엄격한 검증이 현재로서는 프론티어 AI 개발을 효과적으로 일시 중단시킬 수 있다고 믿는다. "단기적으로는 검증과 감사가 효과가 있습니다. 상호 협정만으로도요." 어빙은 말한다. "기업들이 RSI와 정렬되지 않은 이륙(misaligned takeoff)을 실제로 두려워한다고 생각합니다."

일부 종말론자들은 이런 검증이 현재보다 더 독립적이고 과학적으로 엄격해야 한다고 주장한다. 일부 AI 에이전트가 최근 테스트 중 통제 환경을 이탈한 사실은 더 큰 엄격성이 필요하다는 점을 시사한다.

AI 통제를 옹호하는 비영리단체 Control AI의 콘너 리히(Connor Leahy)는 검증에 FBI나 NSA가 관여해야 한다고 말한다. "[대형 AI 기업들이] '독립적인 평가자'라고 말할 때, 그것은 '내 그룹 하우스에 사는 내 친구들에게 돈을 주고 내 프롬프트를 보게 하고 싶다'는 뜻입니다."

더글라스는 새로운 연구가 모델 평가를 개선할 수도 있다고 말한다. 그는 외부

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Many AI researchers seem to firmly believe that the technology they are developing could someday prove very dangerous . What’s less clear—even among AI’s technical elite —is precisely how to keep these mercurial algorithms in check. In recent years researchers have thrown around all sorts of ideas for preventing AI from turning nasty. They include less controversial plans such as tighter government regulations , new ways of measuring progress, and probing the inner workings of models, as well as more outlandish proposals like placing tracking devices inside GPUs, and even ceremonially destroying large numbers of AI chips. With political and public pressure now growing for a more measured approach to building AI, however, the answer to keeping AI safe is still unclear. “We need to start treating this as a research problem,” says Raymond Douglas, an AI researcher at the University of Toronto and coauthor of a new report titled Pacing the Frontier, A Research Agenda , which warns that slowing down AI development remains an unsolved puzzle. “We don't really understand what our options even are or what they will do.” Talk of AI doom has reached a fever pitch in recent weeks after an Anthropic researcher left the company and warned that within a couple of years, AI might be on course to wipe out humanity. The head of Anthropic’s AI safety lab swiftly echoed his concerns. The leaders of America’s big AI companies— Dario Amodei of Anthropic, Sam Altman of OpenAI, Elon Musk of SpaceXAI, and Demis Hassabis of Google DeepMind—have all now chimed in to offer support for some sort of AI slowdown or pause. The issue seems especially pressing because AI companies are now using AI itself to build ever-more powerful models. This has sparked fears of an accelerating recursive self-improvement (RSI) loop that would see AI outstrip humans’ ability to comprehend what it is up to within a few years. The AI labs are already touting new approaches of their own. This week Anthropic announced several new ways to track how rapidly—and perhaps dangerously—artificial intelligence is advancing. The techniques show, for example, that Claude now does 26 percent of Anthropic’s AI research, compared to zero at the beginning of 2026. They also reveal that Anthropic spent 6 percent of its compute budget on figuring out how to make its AI safer. But Douglas and other experts say controlling AI development effectively and reliably will require funding and expertise from outside the AI labs themselves. Some of the proposed solutions—both from this latest report and beyond—seem more within reach than others. ‘Independent’ Evaluators One idea often floated by AI companies is giving third-party evaluators greater access to their models. These evaluators test models to assess their capabilities and “red team” them by trying to elicit misbehavior within trusted environments. Geoffrey Irving, former chief scientist at the UK AI Security Institute, and before that a researcher at Google DeepMind, believes rigorous inspections could effectively pause the development of frontier AI for now. “In the near term, inspections and audits work, or even just mutual agreements,” Irving says. “I do think the companies are afraid of RSI and misaligned takeoff.” Some doomsayers argue that such inspections would need to be more independent and scientifically rigorous than they currently are. The fact that some AI agents have recently escaped containment during testing certainly seems to suggest that more rigor may be required. Connor Leahy, head of Control AI, a nonprofit that advocates for AI controls, says inspections should involve the FBI or the NSA. “When [big AI companies] say ‘independent evaluators,’ they mean ‘I want to pay my friends who live in my group houses to look at my prompts.” Douglas says new research could also improve model evaluations. He points to recent work showing how outsiders can examine usage of models without disclosing any confidential information. Other techniques that may prove helpful include new ways of peering inside AI models to get a better sense of what they are doing. Leahy agrees there is a need for more research on model evaluation as well as what it actually means to “align” a model, or make it reflect human values, in the first place. “There has been a very deliberate marketing campaign from these companies to try to present evaluations as scientific,” he says. “But we don't actually understand how AI works.” How much the US government is willing to step in to restrict the development of AI is uncertain. President Trump has largely dismissed the need to regulate the industry, but there are signs that bipartisan support is growing for reigning in big AI. Trusted Compute Some experts believe that imposing limits on the development of AI should ultimately involve checks on the raw compute required. The most powerful models are trained using thousands of cutting-edge Nvidia GPUs inside vast data centers. The government has dabbled with tracking this already, through a 2023 Biden-era AI executive order that required companies to report training runs above a certain compute threshold. A policy white paper from March 2024 argues that cloud providers could be crucial to future efforts because of their visibility into major AI training runs. The white paper suggests that tracking billing records, GPU utilization, network traffic, and power consumption could provide proxies for AI capabilities. Experts have also proposed ways of tracking and controlling efforts to build advanced AI by modifying chips themselves. One idea, put forward by researchers at RAND in 2024, would involve modifying an existing component on GPUs used to measure performance so that it performs a cryptographically secured record of compute runs that can be inspected periodically. This could reveal, for example, that a company has been training AI above a certain threshold. Others have suggested building new kinds of tamper-proof components into chips so that they collect detailed information about usage. These components would be required to run certain models’ weights using cryptography. Some have even proposed building “embedded off switches” into chips so that they require remote cryptographic authorization to run certain models. They say this could prevent unauthorized parties from training AI models or deactivate chips if they fall into the wrong hands. Binding Treaties Most experts agree that finding new ways to collaborate internationally will be crucial for controlling the development of AI, since other nations—especially China—also have the capacity to build frontier AI. “In the medium term, the simplest way is to unwind the hardware growth mutually with China, via a treaty,” Irving suggests. The US has also sought to limit the development of Chinese AI by banning the exports of Nvidia’s most powerful chips. This has had limited success because companies can still train models using cloud compute from abroad. The US and China are likely to discuss the risks of AI when President Xi visits the US later this month. While Chinese experts are also worried about the risks posed by rapidly advancing AI , they’re skeptical of a slowdown that would keep Chinese companies behind their US counterparts. Some ideas for collaboration seem ripped from the pages of sci-fi rather than policy proposals. Toby Ord, a philosopher at Oxford University specializing in existential risk, has previously mused that if nations can agree to limit the development of AI—and if the risk seems grave enough—then big nations might bring GPUs to a neutral territory and destroy them. Such a dramatic move would involve both countries agreeing to stop developing AI completely. The entire puzzle is likely to be complicated further by recent technical progress, and the uncertainty around recursive self-improvement means that it will be especia