메뉴
BL
Wired AI 20일 전

스스로 발전하는 AI, 우리도 직접 만들 수 있다

IMP
7/10
핵심 요약

저자는 최신 AI 기술을 활용해 뉴스레터 작성에 필요한 반복 작업을 자동화하는 '스스로 개선되는 AI 모델'을 직접 구축한 과정을 공유합니다. Claude와 Prime Intellect 등의 도구를 활용해 특정 업무에 특화된 맞춤형 모델을 만들며, 이러한 자가 개선 기술이 거대 기업의 전유물이 아닌 개인과 일반 기업에도 개방될 수 있음을 보여줍니다. 이는 중앙화된 거대 AI 기업에 의존하지 않고도, 누구나 맞춤형 AI를 구축해 업무 효율성을 크게 높일 수 있음을 시사합니다.

번역된 본문

요즘 최고 수준의 AI 연구소들은 모두 '스스로 개선되는 모델(Self-improving models)'을 만들기 위해 경쟁하고 있습니다. AI가 무한히 자신을 발전시키는 궁극의 루프에 돌입하면 결국 인간의 이해 능력을 초월하게 될 것이라는 생각에, 이를 초지능(Superintelligence)으로 가는 확실한 길이라 믿는 사람들도 있습니다. (어쩌면 통제력을 넘어설 수도 있죠). 그건 그렇다 치고, 저는 당장 제 뉴스레터를 만들어야 했습니다. 그래서 이런 '재귀적 자가 개선(Recursive self-improvement)' 기술이 저 개인에게도 유용하게 쓰일 수 있을지 궁금해졌습니다. 제 뉴스레터 작업 중 귀찮은 단순 반복 작업을 자동화하면서 스스로 계속 똑똑해지는 AI 모델을 만들 수 있을까요? 일주일쯤 실험해 본 결과, 그 대답은 아주 명쾌하고도 놀랍게도 '그 bet(할 수 있다)'였습니다. 게다가 자가 개선 모델을 직접 만들어보면서, 소수의 기업이 전체 산업을 통제하는 데 집중하지 않는 AI의 또 다른 발전 방향을 볼 수 있었습니다. 저는 가장 단순한 자가 개선 루프를 시도하며 시작했습니다. 첫발을 내딛기 위해 처음부터 작은 언어 모델을 학습시키는 실험을 했습니다. (물론 힘든 작업은 모두 Claude에게 떠넘겼습니다.) 저는 일반 AI 모델이 더 작은 모델을 직접 만들고 개선하도록 도와주는 'AutoResearch'를 설치했습니다. 이 도구는 OpenAI 공동 설립자이자 테슬라의 AI 책임자를 거쳐 최근 앤스로픽(Anthropic)에 합류한 슈퍼스타 AI 연구자인 안드레이 카파시(Andrej Karpathy)의 작품입니다. 저는 Claude를 실행시키고 권장 지시어를 입력했습니다. "안녕, program.md 파일을 확인하고 새로운 실험을 시작해 보자!" 어려운 작업은 Claude가 처리하는 동안, 저는 실험용 '실리콘'(AI 실험을 위해 설계된 데스크톱 수퍼컴퓨터인 엔비디아 DGX)과 전력(며칠 동안 과부하를 견디며 작동), 그리고 모델이 제 역할을 하도록 모든 일반적인 권한 확인 단계를 건너뛰게 해주는 다소 무모한 의지(알아서 하라고 놔두기!)를 제공했습니다. 저는 몇 시간마다 AutoResearch 프로젝트를 확인하며 감탄했습니다. Claude가 스스로 파라미터와 학습 체계를 조정하고, 이것이 작은 모델의 출력에 어떤 변화를 일으키는지 확인한 뒤 더욱 정교하게 다듬는 과정을 지켜보았습니다. 초기 버전의 작은 언어 모델에게 "태초에(In the beginning)…"라는 문구를 완성해 보라고 프롬프트를 주자 이런 결과가 나왔습니다. "태초에... 시작의 시작 끝의 끝의 끝의 끝의 끝의 끝 끝의 끝 끝의 시작 끝 끝..." (그다지 훌륭하지 않았습니다.) 하지만 Claude가 자율적으로 개선한 이후의 모델들은 훨씬 말이 되게 출력했고 미친 듯한 무한 반복도 줄어들었습니다. GPT-5 급은 결코 아니었지만, 지속적인 개선으로 나아가는 유망한 길을 보여주었습니다. 그리고 저는 더 복잡하고 유용한 무언가를 향해 여정을 계속했습니다. 저는 이미 주목할 만한 연구 논문을 찾는 데 도움을 주는 Claude 기반의 에이전트를 사용 중이었기에, 이를 뛰어넘는 무언가를 구축할 수 있을지 알아보기로 했습니다. 저는 AI를 사용해 특정 작업을 위한 맞춤형 모델을 학습시키는 스타트업 'Prime Intellect'의 도구를 활용했습니다. 저는 이전 뉴스레터에 포함되었던 'AI 최전선의 다른 소식들' 섹션 약 100개 분량을 수집했습니다. 그다음 Prime Intellect 학습 환경을 만들고 Claude에게 흥미로운 논문을 찾고 요약하는 모델을 만들어 달라고 요청했고, 모델의 이름은 'Frontier_Paper_Curator'로 지어졌습니다. Claude는 더 많은 논문을 찾고 학습에 도움이 되는 다량의 합성 데이터(Synthetic data)를 생성했습니다. 그런 다음 학습 환경에서 강화 학습(Reinforcement learning)을 통해 모델을 개선하는 동시에, 또 다른 모델을 활용해 이 모델의 출력물을 평가했습니다. 최근 1,500만 달러(약 200억 원)의 자금 조달을 받은 Prime Intellect의 빈센트 와이서(Vincent Weisser) 최고경영자(CEO)는 자사의 목표가 이러한 재귀적 자가 개선 기술을 최고 수준의 연구소뿐만 아니라 누구나 사용할 수 있도록 만드는 것이라고 밝혔습니다. 빅테크 연구소들이 만든 모델은 훌륭할 수 있지만, 이러한 방식의 AI 학습을 민주화하면 그에 못지않게 유능한 전문 모델을 만들어낼 수 있다고 그는 말합니다. 와이서 CEO는 이렇게 말합니다. "모든 기업에 최고 수준의 학습 인프라에 대한 접근 권한을 주면, 시장의 집단적 창의성이 소수 연구소가 이뤄낼 수 있는 것보다 훨씬 더 많은 가치를 창출해 냅니다. 우리는 중앙 집중화된, 거의 신(God)과 같은 하나의 지능을 원하지 않습니다."

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story These days, the frontier AI labs are all racing to build self-improving models . Some believe it’s the surest route to superintelligence —as AI improves itself in a mind-melting loop, the thinking goes, it will eventually surpass human comprehension (and perhaps even control). That’s all well and good, but I have a newsletter to produce. I wondered if recursive self-improvement might also be useful for me. Could I use AI to train and continually improve a model that automates some of this newsletter’s busywork? After a week or so of experimenting, the answer appears to be a resounding—and surprising—hell yes. What’s more, dabbling with self-improving models shows a different vision for how AI might unfold—one that doesn’t center on a handful of companies that control the whole industry. I started by trying out a simple self-improving loop To get my feet wet, I experimented with training a small language model from scratch—by which I mean I dumped all the hard work on Claude’s plate. I installed AutoResearch , which helps an off-the-shelf AI model build and improve a smaller model. AutoResearch is the brainchild of Andrej Karpathy , a superstar AI researcher who helped found OpenAI, led AI work at Tesla, and recently joined Anthropic. I fired up Claude and gave it the recommended instruction: “Hi, have a look at program.md and let's kick off a new experiment!” While Claude did the hard stuff, I provided silicon (an Nvidia DGX, a desktop “supercomputer” designed for AI experimentation), the electricity (running hot for a few days straight), and a possibly ill-advised willingness to let the model skip all the usual permission checks in order to do its thing (let him cook!) I checked in on the AutoResearch project every few hours and marveled as Claude adjusted parameters and training regimes, looked at how this changed the smaller model’s output, and went on refining it further. Here’s what an early version of that smaller language model produced when I prompted it to complete the phrase “ In the beginning …” “In the beginning of the beginning of the end of the end of the end end of end end end end end end end end beginning end end end end…” Not so brilliant. But later models, improved autonomously by Claude, got more coherent and less prone to insane, endless repetition. It’s hardly GPT-5, but it showed a promising path toward continual improvement. My journey continued with something more complex—and useful I already use an agent that relies on Claude to help me find noteworthy research papers, so I decided to see whether it was possible to build something that went beyond that. I turned to a tool from a startup called Prime Intellect , which uses AI to train a custom model for a specific task. I collected 100 or so previous “Elsewhere on the frontier of AI” entries—the bits and bobs of research that follow the main essay in my newsletter . Then, I created a Prime Intellect training environment and asked Claude to help me build my own model, which it dubbed Frontier_Paper_Curator, to find and summarize interesting papers. Claude found more papers and generated a bunch of synthetic data to help with training. It then tapped yet another model to assess Frontier_Paper_Curator’s output, while the training environment also improved the model with reinforcement learning. Vincent Weisser, CEO of Prime Intellect, which recently received $15 million in funding, tells me that his company aims to make recursive self-improvement accessible to everyone—not just frontier labs. The models made by frontier labs might be brilliant, but democratizing this kind of AI training could produce just as capable specialized models, he says. “Give every company access to frontier training infrastructure, and the collective creativity of the market unlocks far more than any handful of labs can,” Weisser says. “We don't want one centralized, almost godlike intelligence, we want a billion intelligences that go into all the niches that create beautiful things.” Prime Intellect isn’t the only company that sees the future this way. Adaption, another startup, offers a tool called AutoScientist , which automates AI model training. CEO Sara Hooker says Adaption is working with several large companies that are burning through tokens and don’t have in-house AI experts. When Anthropic decided to block certain requests to its latest model Fable 5, it exposed the risk of relying too heavily on one frontier model. And some executives, like Palantir’s Alex Karp, have warned that using frontier labs also means handing over your own data and control over your technology. The ultimate goal for recursive self-improvement is for AI to apply novel ideas to a model and come up with its own insights. The tools available to the rest of us are more limited, but still impressive. After less than a day of cooking with Prime Intellect, I was able to create a surprisingly good model for finding and summarizing research. Here’s one example entry it created for me: Researchers at iFLYTEK have developed iFLYTEK-Embodied-Omni, a unified multimodal AI model that integrates vision, language, and action generation into a single framework. Unlike prior embodied agents which treat visual understanding, future state prediction, and action generation separately, their model uses shared multimodal self-attention to enable close coordination—analogous to a brain-cerebellum collaboration—between a vision-language "high-level brain" and an action-generating "low-level cerebellum." This approach reduces error compounding and interface bottlenecks common in cascaded pipelines. By training on a large diverse dataset including human and robot-annotated embodied videos and image-text data, and using a staged training strategy, they demonstrate a general-purpose embodied agent capable of joint reasoning, prediction, and control. This contributes a novel architectural and training paradigm toward more integrated, versatile robotic AI systems. Not bad for a first try. The new model is still a bit overeager, choosing too many papers that I would skip, and its summaries are a tad generic. But it’s a promising start. Here’s hoping I can one day use it to free me from the chains of busywork. This is an edition of Will Knight’s AI Lab newsletter . Read previous newsletters here.