메뉴
HN
Hacker News 44일 전

안드로픽의 안전 '초능력'과 정부의 규제

IMP
8/10
핵심 요약

AI 기업 안드로픽(Anthropic)이 매우 뛰어난 성능의 보안 능력을 갖춘 신규 모델 '페이블(Fable)'을 공개했으나, 해킹 등 우회 방법이 발견되어 미국 정부의 강력한 국가 안보 명목의 접근 금지 조치를 받았습니다. 이는 강력한 AI 모델의 안전성 문제와 정부 규제 사이의 갈등이 점차 심화되고 있음을 보여주는 중요한 사례입니다.

번역된 본문

안드로픽의 안전 '초능력' 2026년 6월 15일 월요일 팟캐스트 듣기: 로그인하여 듣기 나는 안드로픽의 공식 성명, 특히 모델 출시를 둘러싼 그들의 발표를 오직 마케팅을 위한 설흔(scare-mongering)이라고 지속적으로 비판하는 회의론자들의 입장에 어느 정도 공감한다. 불과 두 달 전만 해도 안드로픽은 뛰어난 사이버 보안 능력으로 인해 대중에게 공개하기엔 너무 위험하다고 밝힌 '미토스 프리뷰(Mythos Preview)'라는 모델을 발표했다. 그리고 두 달 뒤, 회사는 여러 안전장치(guardrails)를 추가한 미토스의 버전인 '페이블(Fable)'을 대중에 공개했다. 제 한정적인 경험에 비추어 볼 때, 페이블은 매우 인상적인 모델이다. 현재 코딩 성능이 아닌 다른 영역에서 모델을 객관적으로 평가하는 것은 점점 더 어려워지고 있지만, 주관적인 느낌은 존재하며 나는 페이블과의 상호작용에 매우 깊은 인상을 받았다. 이 모델은 GPT 5.5나 오퍼스(Opus) 4.8을 포함한 다른 모델들이 작고 멍청해 보이게 만들었다. 과거에 내가 이런 느낌을 받았던 적은 단 두 번이었는데, 각각 GPT-4와 그록(Grok) 4였으며, 둘 다 기본 모델의 크기와 복잡성 측면에서 새로운 세대를 대표했다. 내 생각에 페이블은 새로운 사전 학습(pre-train)의 결과물이자 새로운 세대의 첫 번째 모델이다. 그렇기 때문에 페이블/미토스가 실제로 보안 문제를 식별하고 악용하는 데 있어 더 뛰어난 능력을 갖추고 있으며, 안드로픽의 신중한 출시 방식이 정당했다는 주장을 충분히 납득할 수 있다. 하지만 모델을 대중에게 공개할 때 문제가 되는 점은 안전장치를 해킹(jailbreak)할 수 있다는 것이고, 출시 직후 실제로 그런 일이 발생한 것으로 보인다. 안드로픽 대 미국 정부, 또 다시 같은 맥락에서 이어지는 상황 이후에 무슨 일이 일어났는지는 다소 불분명하다. 안드로픽은 블로그 게시물에서 다음과 같이 밝혔다: 미국 정부는 국가 안보 권한을 근거로, 미국 내외를 불문하고 모든 외국인(안드로픽의 외국인 직원 포함)에 대한 '페이블 5' 및 '미토스 5'의 모든 접근을 중단하는 수출 통제 지시를 내렸습니다. 이 명령의 결과적으로, 우리는 규정 준수를 위해 모든 고객에 대해 페이블 5와 미토스 5를 즉시 비활성화해야만 했습니다. 안드로픽의 다른 모든 모델에 대한 접근은 영향을 받지 않습니다. 우리는 오늘 동부 표준시(ET) 오후 5시 21분에 정부로부터 이 지시를 받았습니다. 해당 서신에는 국가 안보 우려의 구체적인 세부 사항이 제공되지 않았습니다. 우리가 이해한 바로는, 정부가 페이블 5를 우회하거나 '탈옥(jailbreak)'시키는 방법을 알게 된 것으로 믿고 있다는 것입니다. 우리는 이 특정 기술이 이전에 알려진 소수의 사소한 취약점을 식별하는 데 사용되는 시연을 검토했습니다. 이러한 취약점들은 모두 상대적으로 단순해 보였으며, 우리는 다른 공개된 모델들도 우회 기술 없이도 이를 발견할 수 있다는 사실을 확인했습니다. 안드로픽은 이어서 보편적으로 작동하지 않는 탈옥은 불가피하며 그 범위도 제한적이고, 보편적인 탈옥의 증거는 없다고 주장했다. 한편, 발견된 탈옥 방법은 아마존(Amazon)에 의해 보고된 것으로 보이는데, 이는 안드로픽의 투자자이자 주요 인퍼런스(inference) 제공업체인 아마존의 입장을 고려할 때 주목할 만한 사실이다. 내가 이 글을 쓰는 동안에도 안드로픽의 고위 임원들은 이 상황이 단순한 오해라고 주장하며 해결을 워싱턴 D.C.에서 찾고 있지만, 백악관 관리들은 회사 경영진이 정당한 국가 안보 우려를 가볍게 여기고 있다고 암시하고 있다. 현재 논란이 되는 팩트가 너무 많아 지금 진행 중인 이 충돌에 대해 내가 덧붙일 말은 많지 않다. 다만, 이 갈등이 벌어지고 있다는 사실 자체는 전혀 놀랍지 않다. 나는 이미 '안드로픽과 정렬(Alignment)'이라는 글에서 미국 정부와 안드로픽 간의 갈등이 불가피하다고 설명한 바 있다. 그렇기 때문에 미토스가 정부의 극단적인 조치를 정당화할 만큼 강력하지 않다고 주장하는 사람들은 핵심을 놓치고 있는 것이다. 지금 당장 충분히 강력하지 않더라도, 특히 모델이 스스로의 후속 모델을 만드는 데 점점 더 유용해지는 현 상황에서는 다음 모델, 혹은 그다음 모델은 분명 그 정도로 강력해질 것이다. 하지만 이것은 또 다른 질문을 제기한다. 바로 회의론자들의 관점을 입증하는 듯한 질문이다. 만약 미토스가 그토록 위험하다면,

원문 보기
원문 보기 (영어)
Anthropic's Safety Superpower Monday, June 15, 2026 Listen to Podcast Listen to this post : Log in to listen I'm sympathetic to the cynics who consistently characterize Anthropic's public statements, particularly those surrounding their model releases, as scare-mongering for the sake of marketing. It was only two months ago that Anthropic announced Mythos Preview, a model that they said was too dangerous to make publicly available, thanks in particular to its advanced cybersecurity capabilities. Then, two months later, the company publicly released Fable, a version of Mythos with various safety guardrails. Fable is, in my limited experience, a very impressive model. It's increasingly difficult to objectively evaluate models for anything other than coding performance, but there is subjective feel, and I found my interactions with Fable to be extremely impressive; it made other models, including GPT 5.5 and Opus 4.8, feel small and dumb. The two times I felt that way previously were with GPT-4 and Grok 4, both of which represented new generations in terms of base model size and complexity; my sense is that Fable is downstream of a new pre-train and the first of a new generation. To that end, I can certainly buy the case that Fable/Mythos is in fact more capable when it comes to identifying and exploiting security issues, and that Anthropic's cautious roll-out was justified. The problem with publicly releasing models, however, is that guardrails can be jailbroken, and apparently that is exactly what happened shortly after the release. Anthropic vs. the U.S. Government, Again What happened next is somewhat unclear. Anthropic wrote in a blog post : The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Anthropic models will not be affected. We received the directive from the government today at 5:21pm (ET). The letter did not provide specific details of its national security concern. Our understanding is that the government believes it has become aware of a method of bypassing, or “jailbreaking” Fable 5. We reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities. These vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass. Anthropic went on to make the case that non-universal jailbreaks were inevitable and also narrow, and that there was no evidence of a universal jailbreak; the jailbreak that was found, meanwhile, appears to have been reported by Amazon , which is notable given Amazon is both an investor in Anthropic and a major provider of inference to the company. As I write this, senior Anthropic staff are in Washington D.C. seeking to resolve what they insist is a misunderstanding, and which White House officials are suggesting is insouciance by the company's leadership to legitimate national security concerns. I don't actually have much to add to the current conflict given how many facts are in dispute; what I am not surprised about is the fact that the conflict is happening: I already explained in Anthropic and Alignment why conflict between the U.S. government and Anthropic was inevitable. To that end, people who are arguing that Mythos isn't powerful enough to warrant the government's drastic action are missing the point: if it's not powerful enough now, the next one will be, or the one after that, particularly now that models are increasingly useful in creating their successors. That, however, raises another question — one that seems to validate the cynics' viewpoint: if Mythos is so dangerous, why even release Fable in the first place, and why fight with the government doing exactly what you claim to want? In fact, I think that Anthropic's actions are quite understandable; what makes the company unique is how it justifies them, and it is those justifications that both give the cynics their fuel and Anthropic its magic. The Economic Imperative For the first few years of AI the most economic value has flown to compute, for obvious reasons: we don't have enough supply to meet demand, which has meant skyrocketing prices; the biggest beneficiaries have been Nvidia, TSMC, and the memory makers (SK hynix, Samsung, and Micron). Anthropic and OpenAI, meanwhile, have collectively lost tens of billions of dollars building leading-edge models that, once released, are distilled and commoditized by open source models, primarily from China. This represents the bear case for the labs — they never cover their costs because their differentiation is fleeting, while free alternatives become "good enough" — and I think it's a legitimate one. A world where models are interchangeable is one where models are commodities, while most of the value flows elsewhere. Right now that's compute, but in the fullness of time, whenever we have enough compute, the most valuable place to be in the value chain will be the place that has always been the most valuable: owning the user touchpoint. To that end, it has long been clear to me that the frontier labs have the economic imperative to move closer to the user. If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it's software that owns the user touchpoint, and it's in the frontier labs' long-term interest to not simply be a commodity input into software but to simply replace software outright. Software companies, meanwhile, are working to do the opposite. Satya Nadella laid out his vision for how companies should build on models in an essay on X : Every company is going to have to build what I think of as human capital and token capital. Human capital comprises the knowledge, judgment, relationships, ingenuity, and pattern recognition of its people, while token capital is the firm’s AI capability it builds and owns. Importantly, human capital does not become less valuable as token capital grows. It only becomes more valuable! I believe human agency will be the driver of token capital growth. Humans will set ambitious goals, connect dots across domains, build relationships, and recognize patterns that matter most. Without human direction, you have compute running in circles. This means the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound. You can offload a task, or even a job, but you can never offload your learning. The future of the firm is the ability to compound that learning across people and AI. This requires a new architectural approach where every business is able to build agentic systems that improve over time, while still retaining control over their IP. A company should be able to switch out a “generalist” model without losing the “company veteran” expertise built into their learning system. This is the key “test” of your control and sovereignty in the era ahead. Nadella set this vision off with a warning: The last thing any of us want is a world where every company across every sector is ceding value to a few models that eat everything they see. If all the value is accrued by only a few models, the political economy will simply not tolerate it. There is no societal permission for an AI future that hollows out entire industries. Think about what happened in the first phase of globalization where entire industrial economies were hollowed out by outsourcing. The GDP numbers looked fin