메뉴
BL
The Decoder • 7일 전

클로드, 앤스로픽 연구의 4분의 1 '주도'…하지만 '주도'의 의미가 다르다

IMP
7/10
핵심 요약

앤스로픽이 AI 자체 개발 속도 관련 지표를 공개하며, 클로드가 향후 모델 개발 작업의 26%를 '주도(AL4)'한다고 밝혔습니다. 그러나 이 수치는 인간 작업 시간 기준으로 측정된 것이고, 채점 자체도 클로드가 수행했으며, '주도'란 완전 자율이 아닌 인간이 과제와 방향을 정하는 수준입니다. 이 지표는 다리오 아모데이 CEO의 AI 개발 속도 조절 주장을 뒷받침하는 근거로도 활용될 수 있어 논란의 여지가 있습니다.

번역된 본문

앤스로픽은 클로드가 자사 연구의 4분의 1을 '주도(lead)'한다고 알리고 싶어 하지만, '주도'가 생각하는 의미가 아닐 수 있다는 점을 짚어야 한다.

앤스로픽이 자체 AI 개발 속도에 관한 지표를 공개했다. 수치에 따르면 클로드는 향후 모델 개발 작업의 26%를 '주도'한다. 하지만 기준이 되는 척도는 모호하고, 채점은 클로드 스스로 했으며, '주도'라는 말도 들리는 것만큼 큰 의미는 아니다.

앤스로픽은 게시물에서 세 가지 측정치를 제시했다. AI가 모델 개발에서 얼마나 스스로 처리하는지, AI 에이전트를 얼마나 잘 감시할 수 있는지, 그리고 안전 연구에 얼마나 많은 컴퓨팅이 투입되는지다. 회사는 각 항목을 자체 운영 데이터로 뒷받침한다.

배경에는 앤스로픽 CEO 다리오 아모데이가 AI 프론티어 개발을 조율된 방식으로 늦추자는 촉구가 있다. 이를 위해서는 대중에게 더 많은 정보가 필요하다고 앤스로픽은 말한다. 이 지표들은 모델이 어떻게 만들어지는지 보여주는 것으로, 모델이 무엇을 할 수 있는지 측정하는 능력 테스트를 보완한다.

이 '주도'가 무엇을 의미하고, 무엇을 의미하지 않는가

핵심은 앤스로픽의 모든 개발 작업을 Epoch AI의 척도에 따라 AL0(AI 없음)부터 AL5(완전 자율)까지 분류한 지수다. 2026년 8월 기준으로 26%의 작업이 AL4에 해당하며, 2월의 1% 미만에서 상승한 수치다. 90% 이상이 최소 AL3에 도달한다. 클로드가 AL5에 도달한 경우는 없다.

Epoch AI는 AL4를 'AI 주도(AI leads)'라고 부르며, 앤스로픽은 이를 헤드라인에 사용했다. 하지만 '주도'를 완전 자율과 동일시하면 안 된다. 그것은 AL5의 영역이다.

앤스로픽의 예시는 AL4가 어디에 위치하는지 보여준다. 엔지니어가 클로드에 버그 리포트를 건네면, 클로드는 질문 없이 분석·수정·테스트를 수행하지만 배포는 허용되지 않는다. 인간이 리포트를 읽고 결정한다. 과제와 방향은 여전히 인간에게서 나온다.

AL3('협업')와의 차이는 주로 클로드가 문제에 부딪혀도 더 이상 멈추지 않는다는 점이다.

클로드가 직접 채점했다

에이전트들이 Slack과 내부 문서에서 증거를 수집했고, 다른 클로드 모델이 등급을 매겼다. 앤스로픽은 이 '심판'이 검사 대상 시스템과 같은 실수를 할 수 있음을 인정한다.

'협업'이 끝나고 '주도'가 시작되는 경계는 앤스로픽 스스로에게도 명확하지 않다. 회사의 교차 검증이 이를 보여준다. 직원들에게 자신의 업무 영역이 얼마나 자동화되었는지 판단해달라고 요청했는데, 두 사람이 같은 영역을 평가할 때 같은 등급에 도달한 경우는 약 3분의 1에 불과했다. 공식 점수를 매긴 클로드 모델은 59%의 경우 인간 판단과 일치했다. 97%의 경우 인간과 모델의 평가가 최대 한 단계 차이였다. 하지만 한 단계 차이가 곧 '협업'에서 '주도'로의 도약이다. 어떤 과제가 26%에 포함되는지는 대체로 해석의 문제로 보인다.

그리고 그 26%가 무엇을 측정하는지의 문제도 있다. 앤스로픽은 7월 직원들의 작업 시간 사용 현황을 정리해 각 활동에 점수를 매겼으며, 시간을 많이 소요하는 작업에 더 큰 비중을 두었다. 즉, 클로드는 인간 작업 시간 기준으로 업무의 4분의 1을 '주도'하는 것이다. 이는 클로드가 얼마나 많은 결정을 내리는지, 연구 방향에 발언권이 있는지에 대해서는 아무것도 말해주지 않는다.

물론 이 숫자는 회사에 두 가지 이점을 준다. 제품 성능 주장으로서, 그리고 아모데이의 속도 조절 촉구를 위한 논거로서다. AI 기업들이 안전 경고를 활용해 자신들에게 유리한 규제를 만들려 한다는 비판을 하는 사람들은 '주도'가 정의되는 방식에서 새로운 소재를 찾을 수 있을 것이다.

3만 개 에이전트에 대한 미숙한 감시

보고서에 따르면 앤스로픽의 가장 많이 사용되는 내부 플랫폼에서는 약 3만 개의 에이전트가 동시에 실행된다. 실시간 모니터가 모든 행동을 검사하며 모델 가중치 복사 같은 행위를 차단하는 역할을 한다. 8월의 10억 건이 넘는 결정 중 0.002%를 차단했다. 두 번째 모니터는 매주 약 10만 건의 로그를 사후에 플래그하며, 그중 약 50건만 인간에게 전달된다.

앤스로픽은 이러한 감시를 몇 개월만 운영해왔으며 모든 행동 패턴을 포착하는지 확신하지 못한다고 인정한다. 7월의 표본 주간 동안…

원문 보기
원문 보기 (영어)
Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 18, 2026 Nano Banana Pro prompted by THE DECODER Anthropic is disclosing metrics on how fast it's building its own AI. According to the numbers, Claude "leads" 26 percent of the work on future models. But the underlying scale is fuzzy, the scoring comes from Claude itself, and "lead" means less than it sounds. In a post , Anthropic laid out three measures. How much of model development the AI handles on its own, how well AI agents can be monitored, and how much compute goes into safety work. The company backs each one with figures from its own operations. The context is a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way . For that, the public needs more insight, Anthropic says. The metrics are meant to show how models get built, and they complement capability tests that measure what models can do. What "leads" means here, and what it doesn't At the center is an index that sorts all development work at Anthropic onto a scale from Epoch AI , running from AL0 (no AI) to AL5 (fully autonomous). As of August 2026, 26 percent of the work sits at AL4, up from under one percent in February. More than 90 percent reaches at least AL3. Claude hits AL5 nowhere. Epoch AI calls AL4 "AI leads," and Anthropic put that in its headline. But anyone who ties "lead" to fully autonomy has it wrong. That's what AL5 is for. An example from Anthropic shows where AL4 sits: An engineer hands Claude a bug report, and Claude analyzes, fixes, and tests it without asking questions, but it isn't allowed to ship. A human reads the report and decides. The task and the direction still come from the human. The difference from AL3 ("collaborates") is mainly that Claude no longer stalls when it runs into a snag. Claude did the scoring itself. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels. Anthropic admits this "judge" could make the same mistakes as the system it's checking. Where "collaborates" ends and "leads" begins isn't clear even to Anthropic. A cross-check by the company shows this: Employees were asked to judge how automated their own work area is. When two people rated the same area, they landed on the same level only about a third of the time. The Claude model that assigned the official scores matched the human judgment 59 percent of the time. Almost always, in 97 percent of cases, human and model were at most one level apart. But one level is the jump from "collaborates" to "leads." Whether a task counts toward the 26 percent is often a matter of interpretation, it seems. Then there's what the 26 percent even measures. Anthropic listed how its employees spent their work time in July and scored each activity, with tasks that eat up a lot of time counting for more. So Claude leads a quarter of the work as measured by the human hours it takes. That says nothing about how many decisions Claude makes or whether it has a say in the research direction. The number helps the company two ways, of course: As a performance claim for the product, and as an argument for Amodei's call to slow down. Anyone who accuses AI firms of using safety warnings to shape regulation in their own favor will find fresh material in how "leads" is defined. Young oversight for 30,000 agents On Anthropic's most-used internal platform, about 30,000 agents run at the same time, according to the report. A real-time monitor checks every action and is meant to block things like copying model weights. Out of more than a billion decisions in August, it stopped 0.002 percent. A second monitor flags about 100,000 logs a week after the fact, and roughly 50 reach humans. Anthropic admits it has only run this monitoring for a few months and isn't sure it catches every behavior pattern. During a sample week in July, about six percent of the compute for AI research went into safety work. Anthropic cares about this metric mainly because compute is the input that outsiders can check most easily. If the industry agreed to slow down, it would be a possible lever, for instance through voluntary commitments on the share going to safety research. The low number doesn't necessarily mean little safety work, Anthropic says. Safety work is mostly about designing experiments, which costs researcher time but few chips, while a training run for a new model burns enormous capacity. Anthropic also didn't count work that advances safety and capabilities equally as safety. The company warns that the line to capability research is blurry, and every vendor is tempted to draw it generously. The burden of proof, it says, should sit with the developer. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->