메뉴
BL
MIT Tech Review 15일 전

Anthropic의 최신 AI 발견이 밝히는 것과 한계

IMP
8/10
핵심 요약

Anthropic은 대형 언어 모델(LLM)의 수많은 계산 과정을 들여다보고 결과의 원인을 파악하는 '기계적 해석 가능성(Mechanistic interpretability)' 연구를 심화시키고 있습니다. 최근 연구를 통해 회사는 모델이 문제를 풀 때 사용하지만 겉으로 드러나지 않는 내부 개념 공간인 'J-스페이스(J-space)'를 발견했습니다. 하지만 AI 모델을 심리학적 용어로 해석하는 것은 그 복잡한 수학적 구조를 지나치게 신비화할 수 있다는 비판적 시각도 존재합니다.

번역된 본문

이 기사는 원래 저희의 주간 AI 뉴스레터인 The Algorithm에 게재되었습니다. 이와 같은 기사를 가장 먼저 받아보시려면 여기에서 구독하세요.

거의 1조 달러의 가치를 지닌 현재 세계에서 가장 가치 있는 AI 기업인 Anthropic은 기이하고 심오한 연구를 발표하는 것으로 유명합니다. 예를 들어, AI 모델이 고통을 느낄 수 있는지 조사하거나 사용자가 모델을 '남용'하고 있다고 의심하면 챗봇 대화를 차단하기도 합니다. Anthropic이 다른 AI 기업보다 더 많은 시간과 자금을 쏟아붓는 한 분야는 '기계적 해석 가능성(mechanistic interpretability)'인데, 이는 AI 모델의 복잡한 수학적 구조 내부를 들여다보아 왜 특정 결과를 도출하는지 알아내는 것을 의미합니다.

이는 매우 복잡한 작업입니다. 하나의 결과에 기여할 수 있는 수백만 개의 데이터 포인트가 존재하며, 이들을 파헤치는 것은 유용한 정보라기보다 알 수 없는 단어들의 나열처럼 보일 수 있습니다. 또한 논쟁의 여지가 있습니다. 심리학이나 신경과학에서 빌려온 용어로 AI 모델을 묘사하면 우리가 평가하는 것보다 모델의 행동이 더 정교한 것처럼 보이게 만들 수 있습니다. 바로 이 때문에 Anthropic이 지난주 모델이 답을 도출하는 과정에서 모델의 '내부 생각(internal thoughts)'을 들여다볼 수 있는 새로운 창을 발견했다고 발표했을 때, 제가 꼭 이야기를 나눠야 할 동료가 한 명 있었습니다.

수석 에디터인 윌 더글러스 헤븐(Will Douglas Heaven)은 컴퓨터 과학 박사 학위를 보유한 것은 물론, AI 모델이 어떻게 작동하는지에 대해 우리가 어떤 말을 할 수 있는지 파악하는 데 많은 시간을 할애해 왔습니다. 저는 Anthropic의 새롭고(예상했던 대로 기이한) 연구에서 우리가 얻어야 할 교훈이 무엇인지에 대해 그와 대화를 나눴습니다.

정확히 Anthropic은 여기서 무엇을 알아냈나요? Anthropic은 수년 동안 대형 언어 모델(LLM)이 어떻게 작동하는지 이해하려고 노력해 왔습니다. Anthropic만 이 부분을 살펴보는 것은 아니지만, 이 회사는 다른 곳보다 이를 핵심 미션의 일부로 삼고 있다고 생각합니다. Anthropic의 CEO인 다리오 아모데이(Dario Amodei)는 모델이 어떻게 작동하는지에 대해 더 많이 배우지 않으면 LLM을 완전히 통제할 수 없을 것이라고 말했습니다. 따라서 이 새로운 연구는 바로 그 맥락에 있습니다. 이는 LLM 내부의 기이한 메커니즘에 대해 그 어느 때보다 깊이 파고듭니다.

Anthropic이 발견한 것은 LLM 내부에 'J-스페이스(J-space)'라고 부르는 공간이 있다는 것입니다. 이 공간은 출력에는 나타나지 않지만 문제를 풀어나가는 방식에 영향을 미치는 것처럼 보이는 단어들로 가득 차 있습니다. Anthropic이 모델인 Claude를 탐색하기 위해 새로운 기술을 개발하기 전까지는 이 모든 것이 숨겨져 있었으므로, 이는 진정한 의미의 발견입니다. 때로는 이러한 단어들이 특정 작업에서 LLM이 어디까지 진행했는지 추적하기도 하고, 어떨 때는 인식의 섬광처럼 보이기도 합니다(예를 들어, 단백질 서열의 알파벳만 LLM에 제공하면 '단백질'이라는 단어가 떠오를 수 있음). 그리고 어떨 때는 모델의 의사결정에 대한 일종의 내부 논평을 나타내기도 합니다. 제가 가장 흥미롭게 생각한 예시에서는 '패닉(panic)'이라는 단어가 나타나자 Claude가 코딩 테스트에서 부정행위를 하기로 결정했습니다. 또한 Anthropic은 LLM이 이 공간에서 단어를 설명하고 조작할 수 있다는 것을 발견했습니다. 따라서 어떻게든 그들은 그것을 활용하고 있는 것으로 보입니다.

잠깐 한 걸음 물러나 보죠. 대형 언어 모델이 단순하다고 생각하지는 않지만, 그렇다고 마법도 아닙니다. 단어들 사이의 관계를 학습하는 수학 덩어리잖아요? 그렇다면 LLM 내부를 '들여다보고' 무슨 일이 일어나고 있는지 아는 것이 왜 그렇게 어려운가요?

네, 마법이 아닙니다! 우리가 그것을 완전히 이해하지 못한다는 사실이 신화 창조에 한몫하고 있다고 생각합니다. 그리고 Anthropic이 여기서 기대고 있는 전체적인 서사, 즉 '정말 신비로운 기술을 만들었지만 걱정하지 마세요. 왜냐하면 그것을 알아낼 사람도 바로 우리이니까요'라는 이야기는 이 회사의 분위기와 매우 잘 맞아떨어진다는 점을 짚고 넘어갈 필요가 있습니다. [Anthropic이 자신들의 새 모델이 코딩에 너무 능숙해서 글로벌 사이버 보안 위협이 된다고 경고했으나, 얼마 지나지 않아 미국 정부가 이를 차단했던 방식을 참조하세요.]

그렇습니다. LLM은 그저 수학입니다. 하지만 엄청나게 복잡한 수학이죠. 오늘날의 LLM은 수천억 개의 숫자로 만들어질 뿐만 아니라, 이를 실행하면 수백만 번의 계산이 연쇄적으로 촉발됩니다. 제가 작년에 썼듯이, 중간 크기의 LLM을 종이에 출력한다면 샌프란시스코만 한 도시를 뒤덮을 것입니다.

원문 보기
원문 보기 (영어)
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here . Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. It’s looking into whether AI models can feel pain , for example, and will sometimes cut off chatbot conversations if it suspects users are “abusing” the model. One niche that Anthropic spends more time and money on than other AI companies is called mechanistic interpretability, which means looking inside the complex math of an AI model to learn why it comes up with one particular output and not another. It’s complicated stuff; there are millions of data points that might contribute to any result, and wading through them can look more like word salad than anything useful. It’s also controversial. Describing AI models with terms borrowed from psychology and neuroscience can make their behavior seem more sophisticated than we might otherwise judge it to be. That’s why, when Anthropic announced last week that it had found a new window into its models’ “internal thoughts” as they reason through answers, there was one colleague I had to talk to. Senior editor Will Douglas Heaven, aside from having a PhD in computer science, has spent a lot of time digging into what we can say about how AI models work. I spoke with him about what we should take from Anthropic’s new (and predictably quirky) research. What did Anthropic learn here, exactly? Anthropic has been trying to understand how large language models (LLMs) work for a few years now. Anthropic isn’t the only one looking at this, but I think the company has made it part of its core mission more than most. Anthropic’s CEO, Dario Amodei, has said we won’t be able to control LLMs fully unless we learn more about how they work. So this new research is very much in that context. It goes deeper into the weird mechanisms inside LLMs than ever before. What Anthropic learned was that LLMs have a space inside them—which Anthropic calls the J-space—filled with words that don’t appear in their output but that seem to influence the way they puzzle through problems. All this was hidden until Anthropic developed a new technique to probe its model Claude, so it’s a genuine discovery. Sometimes these words keep track of where the LLM has got to in a particular task, sometimes they look more like flashes of recognition (for example, “protein” might pop up when you give an LLM only the letters of a protein sequence), and sometimes they represent a kind of internal commentary on the model’s decision-making. In my favorite example, Claude decided to cheat on a coding test when the word “panic” appeared. Anthropic also found that LLMs are able to describe and manipulate the words in this space. So somehow they seem to be making use of it. Let’s step back for a second. I don’t think of large language models as simple , but they’re also not magic. There’s a bunch of math that learns relationships between words, right? So why is it so hard to “peer” into an LLM to know what’s going on? Yeah, they’re not magic! I think the fact we don’t fully understand them plays into the mythmaking. And it’s worth noting that the whole narrative that Anthropic is leaning into here—that they’ve built this really mysterious technology, but don’t worry, because they’re also the ones to figure it out—very much fits with the company’s vibe. [See how Anthropic warned that its new models were so good at coding they posed a global cybersecurity risk, only for the US government to shut them down shortly thereafter.] So yes: LLMs are just math. And yet it’s vastly complex math. Not only are today’s LLMs made out of hundreds of billions of numbers, but running them triggers a cascade of millions and millions of calculations. I wrote last year that if you printed out even a medium-size LLM on pieces of paper, it would cover a city the size of San Francisco . It’s impossible to make sense of any of that math without specialist tools that highlight specific parts of an LLM at specific times. You need to know where to look and how to look. And building those tools requires understanding something of that complex math in the first place. You’ve written elsewhere about this concept of studying LLMs the way one might study an organism’s brain. Is it fair to use “brain-like” terms when talking about how an LLM works? I don’t love using those kinds of terms. LLMs are not brains. Talking like this is misleading because it can suggest that LLMs are capable of more human-like things than they are or that we can make assumptions about how they might behave that we shouldn’t. The whole anthropomorphization thing is also tied up with a bunch of strong ideological positions about what this technology is and what it’s going to be . But at the same time, we lack a good alternative vocabulary for talking about what these models are doing. I can understand why people reach for words like “think” and “understand” and “brain-like”—they’re convenient shorthand. Anthropic compares this new space it found inside LLMs to the space that some neuroscientists think our brains use to keep track of conscious thoughts. I asked the company how seriously we should take that comparison and it said in a statement: “Drawing these analogies was helpful to us in designing our experiments, as they allowed us to make many non-obvious experimental predictions about the J-space that turned out to be true. At the same time, it’s important to note that there are some important differences between the J-space (and language models in general) and the human brain, so we don’t mean to claim there’s a perfect correspondence.” What’s a problem in AI that this new concept of the J-space might be used to solve? Anthropic has said that monitoring the J-space could be a way to catch models doing something they shouldn’t. Because words pop up in this space that don’t appear in a model’s output, they can tell you things about its behavior that you might not have noticed otherwise—such as when it is giving biased responses or when it is weighing the pros and cons of cheating. That’s the theory, at least. I think it’s better to think of this result as one more step on the path to understanding this technology overall than as something that will be useful by itself. Read more in Will’s full story about the new research . Deep Dive Artificial intelligence A startup claims it broke through a bottleneck that’s holding back LLMs Subquadratic has now shared more details about its new model. But some are still skeptical. By Will Douglas Heaven archive page A reality check on the AI jobs hysteria What do the numbers really say about the impact of artificial intelligence on the labor market? The answer might surprise you. By David Rotman archive page Anthropic’s Code with Claude showed off coding’s future—whether you like it or not As tools like Claude Code get better, more and more developers are happy to hand off coding tasks to them. The way software gets built has changed for good. By Will Douglas Heaven archive page Google I/O showed how the path for AI-driven science is shifting Two years ago, a n AI tool won Google DeepMind a Nobel. Researchers are now climbing toward a new goal. By Grace Huckins archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.