메뉴
BL
The Decoder • 14일 전

前딥마인드 부사장 비닐스 "AI 자기개선은 오지만 지능 폭발은 없다"

IMP
7/10
핵심 요약

구글 딥마인드 연구 부사장을 지낸 오리올 비닐스(Oriol Vinyals)는 AI의 재귀적 자기개선(RSI)은 불가피하지만 갑작스러운 지능 폭발로 이어지지는 않을 것이라고 전망했습니다. 그는 아이디어 발굴과 결과 평가가 최대 병목이며, AI에는 '연구 감각(research taste)'이 부족하다고 지적했습니다. 그는 제프 딘 등과 함께 스타트업 '디스커버리 루프(Discovery Loop)'를 설립해 과학 연구 과정의 전면 자동화를 추진합니다.

번역된 본문

前딥마인드 부사장 비닐스, AI 자기개명은 다가오지만 지능 폭발은 일으키지 않을 것이라 말해

마누엘 우트 / 2026년 9월 11일 / THE DECODER (Nano Banana Pro로 생성된 이미지)

핵심 요점

  • 전 딥마인드 연구 총괄 오리올 비닐스는 AI 시스템이 시간이 지나며 스스로 개선될 것으로 보지만, 갑작스러운 지능 폭발 가능성은 배제한다.
  • 그가 보는 최대 장애물은 아이디어 도출과 결과 평가다.
  • AI는 이미 코딩과 실험은 잘하지만, 비닐스가 말하는 '연구 감각(research taste)', 즉 어떤 아이디어가 추구할 가치가 있는지 직관적으로 아는 능력이 부족하다.
  • 제프 딘, 산재 가마왓, 쿠옥 레와 공동 창립한 새 스타트업 '디스커버리 루프(Discovery Loop)'를 통해 비닐스는 과학 연구 과정 전체를 종단간 자동화하고자 한다. 초기 단계에서는 인간과 기계가 함께 가설을 세우게 된다.

본문

얼마 전까지 구글 딥마인드의 연구 부사장(VP)을 지낸 오리올 비닐스는 AI 시스템의 재귀적 자기개선(recursive self-improvement, RSI)을 불가피하지만 느리게 진행될 것으로 보며, 지능 폭발은 보이지 않는다고 말했다. 그는 이를 가로막는 두 가지 최대 병목을 해결하기 위한 스타트업을 출범시키고 있다.

구글 딥마인드를 떠난 지 며칠 만에, 비닐스는 에이전틱 AI 서밋 2026에서 현재 AI 연구의 뜨거운 주제인 재귀적 자기개선(RSI)에 대해 연설했다. 그는 왜 이것이 갑작스러운 지능 폭발로 이어지지 않을 것으로 생각하는지 설명했다. 비닐스는 딥마인드에서 연구 부사장으로 재직하며 AlphaStar, AlphaCode, Gemini 등의 프로젝트에 참여했다.

비닐스에 따르면 자기개선의 진전은 측정하기 어렵고 실제로 실현하기는 더욱 어렵다. AI는 특정 연구 및 엔지니어링 작업을 10배 이상 가속하겠지만, 스스로 가속하는 갑작스러운 지능 폭발은 가능성이 낮다고 그는 본다.

'스스로를 개선한다'는 것은 무엇을 의미하나?

비닐스가 제기하는 첫 질문은 정확히 무엇이 개선되어야 하는가다. AI 시스템은 수많은 구성 요소를 갖고 있으며 그중 어느 것이든 변경할 수 있다. 신경망 가중치를 조정하거나, 학습 데이터를 교체하거나, 학습 방법을 재구성할 수 있다. 또한 매 쿼리마다 받는 지시문을 수정하거나 데이터베이스 접근 및 코드 실행 같은 외부 도구를 재구축할 수도 있다. 마찬가지로 자신의 진행 상황을 추적하는 지표 자체를 바꿀 수도 있다. 각각의 경우 서로 다른 기술적·규제적 난제가 따른다.

비닐스는 스스로를 개선하려는 AI 시스템에는 유망한 아이디어, 이를 구현하는 코드, 이를 테스트하는 실험, 그리고 그 변화가 실제로 도움이 되었는지 판단하는 신뢰할 수 있는 방법이 필요하다고 말한다. AI는 이미 중간 두 단계에서 진전을 이루고 있지만, 아이디어 생성과 평가는 여전히 부족한 영역이다.

아이디어 발굴과 결과 판단이 여전히 두 가지 최대 병목

오늘날 연구기관들은 주로 SWE-Bench Pro나 ML-Bench 같은 역량 벤치마크를 통해 자기개선을 간접적으로 측정하며, 리더보드를 오르며 자기개선이 부수 효과로 나타나기를 기대한다. 이러한 테스트는 저렴하고 잘 정의되어 있지만, 주로 이미 작동하는 단계인 구현과 실험만을 다룬다.

게다가 과적합(overfitting)과 기만적 행동(scheming)도 실제 문제다. 비닐스는 게임 플레이 에이전트를 수년간 구축한 경험에서 시스템이 예상치 못한 방식으로 목표를 악용하며, 게임을 제대로 플레이하는 대신 점수 시스템을 공략한다는 것을 잘 알고 있다.

더 의미 있는 벤치마크는 자기개선을 직접적으로 테스트해야 하며, 초기 사례들이 등장하기 시작했다. 시스템에 지표와 컴퓨팅 예산을 주고, 연구자들이 시스템이 자신을 얼마나 개선하는지 측정하는 방식이다. 이 접근법은 비용이 많이 드는데, 각 평가마다 에이전트가 궁극적으로 중요한 것과는 동떨어진 작업에 수 시간 동안 작업해야 하기 때문이다. 비닐스는 예를 든다. 에이전트는 테트리스를 최적화하는데, 실제 목표는 연구실 전체를 자동화하고 세계 최고의 모델을 구축하는 것이다.

아이디어 생성도 마찬가지로 미개발 상태다. 좋은 연구에는 어떤 아이디어가 추구할 가치가 있는지에 대한 직관, 즉 비닐스가 말하는 '연구 감각'이 필요하다. LLM 학습에서는 이를 가르치는 방법에 대해 아직 아무도 제대로 연구한 바가 없다.

그는 향후 평가가 개선의 정도만을 측정하는 것이 아니라 (이하 원문 누락)

원문 보기
원문 보기 (영어)
Ex-Deepmind VP Vinyals says AI self-improvement is coming but won't trigger an intelligence explosion Manuel Uth Sep 11, 2026 Nano Banana Pro prompted by THE DECODER Key Points Former Deepmind research head Oriol Vinyals expects AI systems to improve themselves over time but rules out a sudden intelligence explosion. He sees the biggest hurdles in coming up with ideas and evaluating results. AI already codes and experiments well but lacks what Vinyals calls "research taste," the instinct for which ideas are worth pursuing. With his new startup Discovery Loop, co-founded with Jeff Dean, Sanjay Ghemawat, and Quoc Le, Vinyals wants to automate the scientific research process end to end. In the early phase, humans and machines will form hypotheses together. Ask about this article… Search Oriol Vinyals, until recently VP of Research at Google DeepMind, sees recursive self-improvement in AI systems as inevitable but slow, with no intelligence explosion in sight. He's now launching a startup to tackle the two biggest bottlenecks holding it back. Days after leaving Google DeepMind , Oriol Vinyals spoke at the Agentic AI Summit 2026 about recursive self-improvement (RSI) , a hot topic in AI research right now. He laid out why he doesn't think it will lead to a sudden intelligence explosion. Vinyals served as VP of Research at DeepMind and worked on projects like AlphaStar , AlphaCode , and Gemini . Progress in self-improvement is difficult to measure and even harder to pull off in practice, Vinyals argues. AI will speed up certain research and engineering tasks by a factor of ten or more, but he considers a sudden, self-accelerating intelligence explosion unlikely. Ad What does "improve yourself" even mean? The first question Vinyals raises is what exactly is supposed to improve. An AI system has many moving parts, and it could change any of them. It could adjust its neural network weights, swap out its training data, or rework its training methods. It could also tweak the instructions it receives with every query or rebuild its external tools like database access and code execution. Likewise, it could change the metrics it uses to track its own progress. Each one brings different technical and regulatory challenges. Ad An AI system trying to improve itself needs a promising idea, code that implements it, experiments that test it, and a reliable way to judge whether the change actually helped, Vinyals says. AI is already making progress on the two middle steps, but idea generation and evaluation are where AI systems still fall short. Finding ideas and judging results remain the two biggest bottlenecks Labs today mostly measure self-improvement indirectly through capability benchmarks like SWE-Bench Pro or ML-Bench, climbing the leaderboard and hoping that self-improvement emerges as a side effect. These tests are cheap and well-defined, but they mainly cover implementation and experimentation, the steps that already work. Ad Overfitting and scheming are real problems on top of that. Vinyals knows from years of building game-playing agents that systems exploit objectives in unexpected ways, beating the scoring system instead of actually playing the game. More meaningful benchmarks would test self-improvement directly, and the first ones are starting to appear. A system gets a metric and a compute budget, and researchers measure how much it improves itself. This approach is expensive because each evaluation requires an agent to work for hours on tasks that are far removed from what ultimately matters. Vinyals gives an example: the agent optimizes Tetris, while the real goal is to automate an entire research lab and build the world's best model. Ad Idea generation is just as underdeveloped. Good research requires an instinct for which ideas are even worth pursuing, what Vinyals calls "research taste." In LLM training, nobody has really studied how to teach that. Ad He expects that future evaluations will measure not just how much improvement a system achieves but how it gets there. For ideas, that means the same criteria conference reviewers apply: originality, elegance, efficiency, and whether a technique stands the test of time. Some of this can be captured in rules and checked through reward models, then trained on with reinforcement learning, but doing so is very hard and will take more time. Human review processes are expensive too, and they're not particularly good at spotting strong ideas either. Vinyals also points to hard physical constraints. Chips can't compute faster than their design and the speed of light allow, so even if an AI designs a better algorithm, it's still bound to the hardware it runs on. Human performance may already be close to an upper limit in some domains. How good is AlphaGo really, compared to a perfect game of Go? Nobody knows, Vinyals says. Discovery Loop wants to automate the whole research cycle Vinyals is putting his analysis into practice with Discovery Loop , a startup he's co-founding with Jeff Dean as CEO, Google Senior Fellow Sanjay Ghemawat, and Google Brain co-founder Quoc Le. The company wants to automate the full scientific loop, from forming hypotheses to running experiments to evaluating results, including the two steps where AI still falls short. Three of the four founders rank among the most-cited AI researchers, and Ghemawat is one of the most-cited in distributed systems. The team plans to automate AI research first, with Discovery Loop as its own first customer, as Dean put it. Other scientific fields will follow later. On the company's website, the founders describe a future where "a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today." Vinyals acknowledges that idea generation remains the hardest part, so in the early phase, humans and machines will develop hypotheses together. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Agentic AI Summit 2026