메뉴
HN
Hacker News 54일 전

AI가 스스로를 발전시킬 때: 재귀적 자기 개선을 향한 여정

IMP
9/10
핵심 요약

AI가 직접 후속 모델을 설계하고 개발하는 '재귀적 자기 개선' 시대가 현실로 다가오고 있습니다. Anthropic의 내부 데이터와 벤치마크에 따르면 AI 엔지니어링 생산성이 폭발적으로 증가하고 있으며, 에이전트가 수행할 수 있는 작업의 범위와 시간도 기하급수적으로 확장되고 있습니다. AI가 스스로를 개선하는 수준에 도달할 경우 과학, 의료 등 인류에 막대한 혜택을 주겠지만, 동시에 인간의 AI 통제력 상실 리스크도 커진다는 점에서 그 안전성과 통제가 매우 중요해집니다.

번역된 본문

AI 역사의 대부분 동안 인간은 AI 개발 주기의 모든 단계를 주도했습니다. 하지만 Anthropic에서는 AI 시스템 자체에 AI 개발의 비중을 점점 더 많이 위임하고 있으며, 이는 우리의 업무 속도를 높이고 있습니다. 이러한 추세가 충분히 극대화되고 충분한 컴퓨팅 능력이 주어진다면, AI 시스템이 완전히 자율적으로 자신의 후속 시스템을 설계하고 개발할 수 있는 수준에 도달할 것입니다. 이를 '재귀적 자기 개선(recursive self-improvement)'이라고 부릅니다. 아직 우리는 그 단계에 이르지 않았으며, 재귀적 자기 개선이 필연적인 것도 아닙니다. 하지만 이는 대부분의 기관이 준비하는 것보다 더 빨리 찾아올 수 있습니다. Anthropic Institute는 공개 벤치마크와 Anthropic 내부의 이전에 보고되지 않은 데이터를 활용하여, AI가 이미 AI 시스템의 개발을 가속화하고 있음을 보여주고 있습니다. 한 가지 예를 들자면, 오늘날 Anthropic 엔지니어들이 분기당 배포(Ship)하는 코드의 양은 2021~2025년 평균에 비해 8배나 많습니다. 이 글에서 논의된 기술적 추세는 향후 수년 내에 AI 시스템이 훨씬 더 능력치를 갖추게 될 것임을 시사합니다. 이러한 추세는 엄청난 파급력을 갖습니다. 스스로를 구축할 수 있는 AI는 기술 역사상 중대한 발전이 될 것이며, 과학, 의료 및 그 이외의 분야에서 세상에 막대한 선을 가져올 수 있습니다. 하지만 완전한 재귀적 자기 개선은 인간이 AI 시스템에 대한 통제력을 잃을 위험을 증가시킬 수도 있습니다. 시스템이 완전히 자신의 후속 모델을 구축할 수 있게 되면, 우리가 그것을 안전하게 보호하고 모니터링하며 그 행동을 올바르게 형성하는 방식이 모두 훨씬 더 중요해집니다.

2021–2023: 첫 번째 Claude 구축 초기 시절, Anthropic의 업무는 다른 모든 기술 회사와 비슷했습니다. 즉, 사람들이 노트북으로 코드와 문서를 직접 작성했습니다.

2023–2025: 챗봇 사람들은 초기 챗봇을 활용해 짧은 코드 스니펫을 생성하고 그 출력물을 텍스트 편집기에 복사하는 등 업무 과정의 일부를 보조받았습니다.

2025–2026: 코딩 에이전트 에이전트들이 더 유능해지면서, 때때로 전체 파일을 포함하여 코드를 직접 작성하고 편집할 수 있게 되었습니다.

현재: 자율 에이전트 이제 에이전트는 스스로 코드를 실행하고, 다른 에이전트에게 몇 시간씩 걸리는 업무를 위임할 수 있습니다.

20XX?: 루프(Loop) 닫기 미래에는 에이전트가 스스로 모델을 구축하고 훈련시킬 수 있을 만큼 능력이 발전할 수 있습니다. 이 일이 현실이 되면, 미래의 Claude 버전은 Claude 자체에 의해 지속적으로 개선될 수 있을 것입니다.

외부 세계의 증거 AI 모델이 발전하는 속도가 가속화되고 있습니다. AI가 자율적으로 안정적으로 완료할 수 있는 작업의 길이는 이전의 7개월마다 두 배로 증가하던 추세에서 대략 4개월마다 두 배로 증가하고 있습니다. 2024년 3월, Claude Opus 3은 인간이 완료하는 데 약 4분이 걸리는 소프트웨어 작업을 해결할 수 있었습니다. 1년 후인 Claude Sonnet 3.7은 약 1시간 반 정도 걸리는 작업을 수행했습니다. 그리고 그 1년 후인 Claude Opus 4.6은 12시간짜리 작업을 해결했습니다. 이 추세가 계속 유지된다면, 올해 안으로 숙련된 사람이 며칠씩 걸려 하던 작업을 AI가 해결할 수 있는 범위에 들어올 수 있습니다. 2027년에는 한 사람이 몇 주가 걸릴 작업을 AI 시스템이 수행할 수 있게 될 것입니다.

동일한 패턴이 코딩 및 연구 벤치마크에서도 나타납니다. 벤치마크는 특정 도메인에서 모델의 성능을 측정하며, 모델이 100%에 가까운 성능을 달성하면 해당 벤치마크가 '포화(Saturated)'되었다고 평가합니다. SWE-bench는 실제 소프트웨어 엔지니어링을 평가하는 표준 테스트입니다. 이 테스트는 모델에게 실제 오픈소스 코드베이스와 실제 버그 리포트를 제공하고, 문제를 해결하며 프로젝트 자체의 테스트를 통과하는 코드 변경을 작성하도록 요구합니다. 모델들의 성적은 한 자릿수에 머물던 수준에서 시작해 불과 2년 만에 해당 벤치마크를 포화시키는 수준에 도달했습니다.

CORE-Bench은 모델이 기존 연구를 재현할 수 있는지 테스트하는데, 이는 모델이 독창적인 연구를 수행하기 위한 필수 조건입니다. 이 벤치마크는 AI 모델에게 출판된 논문의 코드와 데이터를 제공한 뒤 모든 것을 재실행하여 논문의 결과를 복제할 수 있는지 확인하도록 요구합니다. AI 시스템은 2024년에 약 20%의 성공률을 보였으나, 15개월 후에는 벤치마크를 포화시키는 수준까지 올라갔습니다.

장기 작업을 완료하는 모델의 능력을 측정하는 벤치마크를 운영하는 METR은 Claude Mythos Preview가 '최소' 16시간 동안 작업할 수 있었으며, 이는 "[METR]이 현재 측정할 수 있는 범위의 상한선"에 해당한다고 발표했습니다.

원문 보기
원문 보기 (영어)
For most of AI’s history, humans drove every step in its development cycle. But at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work. Taken far enough, and given enough compute, that trend points to an AI system capable of fully autonomously designing and developing its own successor. This is called recursive self-improvement . We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for. Using public benchmarks and previously unreported data from within Anthropic, The Anthropic Institute is showing that AI is already accelerating the development of AI systems. To take just one example: today, Anthropic engineers on average ship 8x as much code per quarter as they did from 2021-2025. The technical trends discussed in this piece suggest that AI systems are going to become much more capable in coming years. These trends have huge implications. AI that can build itself would be a major development in the history of technology—one that could bring enormous good for the world in science, healthcare, and beyond. But full recursive self-improvement also might increase the risks of humans losing control over AI systems. If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important. 2021–2023 Building the first Claude In the early days, work at Anthropic looked like work at any other tech company: people writing code and docs on laptops. 2023–2025 Chatbots People used early chatbots to help with parts of the process, like generating short code snippets and copying the output into text editors. 2025–2026 Coding agents As the agents became more capable, they were able to write and edit code on their own, sometimes entire files. Today Autonomous agents Agents can now run code themselves and delegate hours of work to other agents. 20XX? Closing the loop In the future, agents could become capable enough to build and train models themselves. If this happens, future versions of Claude could be continuously improved by Claude itself. Evidence from the outside world The rate at which AI models improve is accelerating. The length of tasks that they can reliably complete on their own has been doubling roughly every four months, up from an earlier trend of doubling every seven months. In March 2024, Claude Opus 3 could complete software tasks that take humans about four minutes to complete. A year later, Claude Sonnet 3.7 managed tasks that took about an hour and a half. A year after that, Claude Opus 4.6 managed 12-hour tasks. 1 If this trend holds, tasks that take a skilled person days could come into range this year. In 2027, AI systems could be capable of tasks that take a person weeks. The same pattern appears on coding and research benchmarks. Benchmarks measure the performance of models in a given domain, and they’re “saturated” when models achieve close to 100% performance. 2 SWE-bench is a standard test of real-world software engineering: it hands a model an actual open-source codebase and a real bug report, and asks it to write a code change that fixes the issue and passes the project’s own tests. Models have gone from scoring in the low single digits to saturating the benchmark in two years. CORE-Bench tests whether a model can reproduce existing research, a prerequisite for them to conduct original research. It gives an AI model the code and data behind a published paper, and asks it to rerun everything and confirm it can replicate the paper’s results. AI systems went from succeeding at reproducing the results roughly 20% of the time in 2024 to saturating the benchmark fifteen months later. METR, which runs the benchmark measuring how well models can complete long-duration tasks, found that Claude Mythos Preview could work for “at least” 16 hours and was “at the upper end of what [METR] can measure without new tasks.” Public benchmarks say a lot about the capabilities of these systems. But they can’t reveal the impact AI systems are having on speeding up AI development itself. For that, we need direct evidence from within AI companies like Anthropic. Evidence from within Anthropic Building a frontier model takes two broad categories of work. There is engineering : writing the code, standing up the infrastructure, and overseeing the model training. And there is research : deciding what experiments to run, interpreting what comes back, and figuring out which ideas to try next. Across both engineering and research, the picture is consistent. In engineering, Claude can be handed an underspecified problem and figure out how to solve it; humans supply the goal, but they no longer need to supply the method. In research, Claude can already match or outperform skilled humans at executing a well-specified experiment. However, large performance gaps persist when it comes to Claude exercising judgement in choosing goals in both engineering and research. That’s the gap between AI today and a future system that could autonomously design its own successor. It’s common for employees at Anthropic to receive more open-ended and important tasks as they gain more experience. Early on, they execute a task someone else specified, like, “The export button isn’t working, please fix it.” With experience, they’re handed a goal and design the approach themselves, such as, “Investigate why the network slows down under heavy load.” At the most senior levels, they are deciding which problems are worth working on at all: “What should the team build next quarter?” We can use internal Anthropic data to see how far Claude has come in being able to handle these different kinds of tasks. Claude writes a significant proportion of Anthropic’s code. As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude. 3 Before Claude Code launched in research preview in February 2025, this number was in the low single digits. That shift also shows up in the amount of output per engineer. Lines of code merged per engineer per day stayed constant through Anthropic’s first four years (2021-2024), then began to climb upward in 2025 when Claude began to run code rather than just suggesting it for an engineer to copy and paste. The slope steepened again in 2026 when models began to work autonomously over longer time horizons. These two inflection points are shown in the chart below. In the second quarter of 2026, the typical engineer was merging 8× as much code per day as they were in 2024. 4 This is because much of the code is written by Claude, with the engineer directing and reviewing, rather than typing it themselves. A caveat: Lines of code is an imperfect measure, as it measures quantity over quality. So 8 × lines of code/engineer/day in the second quarter of 2026 is almost certainly an overstatement of the true productivity gain. Nonetheless, it indicates an acceleration. At Anthropic, we don’t reward people for how many lines of code they write; rather, team members are producing more code simply because they’re using AI systems to write more code. The increase in lines of code written lines up with subjective impressions of large productivity increases. In a March 2026 poll of 130 employees from across Anthropic research teams, the median respondent estimated that they produced around 4x as much output with Mythos Preview as they would have without access to any AI models, on the kinds of projects they would have been working on regardless. 5 We expect that the true degree of uplift in March was somewhat lower. 6 Nevertheless, we find the overall claim plausible, and in line with our other observations: a significant fraction of Anthropic technical staff is accomplishing their core work multiple times faster than they could without AI assistance. We also see evidence that people at Anthropic are using Claude to do work that sim