메뉴
BL
MIT Tech Review • 18일 전

예상 못한 상황까지 계획하는 AI 에이전트를 개발하는 사업가

IMP
7/10
핵심 요약

다니자르 하프너(31)는 '모델 기반 강화학습'과 월드 모델을 활용해 훈련에서 접하지 못한 낯선 환경에서도 스스로 대처하는 AI 에이전트와 휴머노이드 로봇을 개발하고 있습니다. 그가 만든 Dreamer 시리즈는 아타리 게임에서 인간 수준 성능 달성, 마인크래프트 다이아몬드 채굴 성공 등의 성과를 냈으며, 현재는 새 스타트업을 통해 이 기술을 실물 로봇으로 확장하고 있습니다. 이는 로봇이 실제 가정 등 인간 생활 공간에 진입하는 데 핵심적인 기술로 주목받습니다.

번역된 본문

2026 35세 이하 혁신가 전체 보기. 다니자르 하프너의 샌프란시스코 소마 지구 사무실은 대부분 비어 있다. 갓 창업한 스타트업이 아직 스텔스 모드라서 문에 회사 이름조차 없다. 방문한 날에도 다른 사람은 단 한 명뿐이었고 가구도 거의 없었다. 하지만 인테리어가 부족한 만큼 로봇으로 채워져 있었다. 다양한 형태와 크기의 휴머노이드 로봇들이 넓은 공간 중앙을 가로지르는 선반에 꼭두각시처럼 매달려 있었다. 31세인 하프너는 새 벤처에 대해 아직 많이 말하지 않지만, 이를 AI가 훈련 과정에서 접해보지 못한 환경을 헤쳐나갈 수 있게 하는 오랜 연구工作的 연속이라고 설명한다. 그가 중국에서 수입한 이 휴머노이드들은 이 연구의 다음 단계이자 물리적 구현체다. 검증되지 않은 시나리오에서 반응할 수 있는 능력은 로봇을 인간 공간에 들여보내는 열쇠가 될 것이다. 예를 들어 로봇을 사람의 집에 보내려면 본 적 없는 평면 구조와 가구에 대처할 수 있어야 하기 때문이다.

이를 달성하기 위해 하프너는 '모델 기반 강화학습(model-based reinforcement learning)'에 의존한다. 그는 물리적 현실을 모방하도록 설계된 AI 모델인 '월드 모델(world model)'을 개발하고 그 안에서 에이전트를 훈련시킨다. 에이전트는 본질적으로 이 모델을 실제 세계의 시뮬레이션처럼 취급하고 그 안에서 행동하는 법을 배운다. 그런 다음 그 경험을 사용해 미래의 결과에 대해 예측(하프너의 표현대로 '꿈꾸거나 상상하는 것')을 한다. 이를 통해 에이전트(또는 이들이 탑재된 로봇)는 실제 세상에서 낯선 상황을 헤쳐나갈 수 있다.

"나는 구글 리서치에서 정말 똑똑한 많은 사람들과 함께 일하는데, 그는 상위 0.5% 안에 쉽게 든다." — 티머시 릴리크랩, 구글 딥마인드

다른 노력들과 달리 하프너의 기술은 에이전트와 그가 제어하는 로봇이 전통적으로 로봇공학에서 사용되어 온 실세계 시행착오 훈련 없이도 대단히 복잡한 작업을 수행할 수 있게 한다.

하프너는 독일 북동부의 시골 마을에서 자랐으며 부모님은 모두 클래식 음악가였다. 이웃에게 프로그래밍을 배웠고, 고등학교 때 AI 온라인 강좌를 듣기 시작하면서 이는 빠르게 열정으로 발전했다. "나는 늘 사고(thinking)가 어떻게 작동하는지에 매료되어 있었다"고 그는 말한다. AI는 컴퓨터에서 이를 모방할 수 있는 방법을 제공했다. 2015년, 포츠담의 하소 플라트너 연구소에서 공학을 전공하던 2학년이던 그는 구글 브레인의 학생 연구원 자리를 얻었다. 이후 영국, 캐나다, 미국에서 구글 브레인과 구글 딥마인드(이후 딥마인드로 통합)에서 인턴십과 여러 직책을 거쳤다. AI의 대부 중 한 명으로 불리는 제프리 힌턴과, 오늘날 대형 언어 모델이 사용하는 트랜스포머 기술을 설명한 획기적인 논문 "Attention Is All You Need"의 공동 저자 아쉬시 바스와니 등 업계 전설적인 인물들과 함께 일했다.

구글에서 그의 전 매니저이자 공동 저자였던 티머시 릴리크랩은 그를 뛰어난 사람들 중에서도 돋보이는 인물이라고 평가한다. "나는 구글 리서치에서 정말 똑똑한 많은 사람들과 함께 일하는데, 그는 상위 0.5% 안에 쉽게 든다"고 릴리크랩은 말한다. "많은 경우 그는 엔지니어 팀 전체가 만들어야 할 것을 혼자서 만들어냈다."

수년간 하프너는 자신의 월드 모델에서 훈련된 에이전트를 인기 비디오 게임에 맞붙여 접근법을 다듬고 입증해왔다. 첫突破는 에이전트가 미리 계획을 세워 행동을 실행할 수 있게 한 모델 PlaNet이었다. Dreamer 2는 월드 모델을 사용해 아타리 2600 게임에서 인간 수준의 성능에 도달한 최초의 에이전트였다. Dreamer 3는 마인크래프트 다이아몬드 챌린지—게임 내 보석을 스스로 채굴하는 것—를 해결한 최초의 에이전트였다. 그리고 Dreamer 4는 한 걸음 더 나아가 게임과 직접 상호작용하지 않은 채 녹화된 게임 플레이 영상의 오프라인 데이터셋만으로 다이아몬드 채굴을 학습했다. 최근에는 자신의 에이전트를 가상 세계에서 실제 세계로 이전하기 시작했다.

원문 보기
원문 보기 (영어)
2026 Innovators Under 35 View all the innovators Danijar Hafner’s office in San Francisco’s SoMa district sits mostly empty. His brand-new startup is still in stealth mode and doesn’t even have its name on the door. On the day I visit, there’s only one other person there, and little in the way of furniture. But what it lacks in decor, it makes up for in robots. Humanoids of various shapes and sizes hang like marionettes from racks that run down the center of the wide-open space. While Hafner, 31, won’t say too much about his new venture just yet, he describes it as a continuation of his longtime work to enable AI to navigate environments it has not encountered in training. The humanoids, which he imports from China, are the next evolution of this work—and its physical embodiment. Their ability to react in previously untested scenarios will be key to getting robots into human spaces. Because if you want to send a robot into a person’s home, for example, it needs to be able to handle a floor plan and furniture it’s never seen before. To achieve this, Hafner relies on something called model-based reinforcement learning. He develops world models—AI models designed to emulate physical reality—and trains agents within them. The agent essentially treats the model as a real-world simulation and learns how to act there. It then uses those experiences to make predictions (to dream or imagine, Hafner might say) about future outcomes. That allows agents—or the robots they’re embedded in—to navigate unfamiliar situations IRL. “I get to interact with a lot of really smart people in research at Google, and he easily sits in the top half of 1%.” Timothy Lillicrap, Google DeepMind Unlike other efforts, Hafner’s technique enables agents and the robots they control to execute massively complicated tasks without the real-world trial-and-­error training that’s traditionally been used in robotics. Hafner grew up in a rural town in northeastern Germany, where his parents were both classical musicians. He learned programming from a neighbor, and in high school he began taking online courses about AI, which quickly developed into a passion. “I was always fascinated with how thinking works,” he says. AI offered him a way to emulate it on a computer. In 2015, as a second-year under­graduate studying engineering at Hasso Plattner Institute in Potsdam, he won a role as a student researcher at Google Brain. From there, he went on to a dozen internships and other positions at the company, including stints with Google Brain and Google DeepMind (the two have since merged under DeepMind) in the UK, Canada, and the US. He worked with industry legends including Geoffrey Hinton, who is often referred to as one of the godfathers of AI, and Ashish Vaswani, coauthor of the groundbreaking research paper “Attention Is All You Need,” which described the transformer technology used by today’s large language models. One of Hafner’s former managers and coauthors at Google, Timothy Lillicrap, describes him as a standout among standouts. “I get to interact with a lot of really smart people in research at Google, and he easily sits in the top half of 1%,” Lillicrap says. “In many cases he would build, single-­handedly, things it would take entire teams of engineers to build.” Over the years, Hafner has honed and proved his approach by pitting agents trained within his world models against popular video games. His first breakthrough was PlaNet, a model that allowed agents to execute actions by planning ahead. His Dreamer 2 was the first agent to hit human-level performance playing Atari 2600 games using a world model. Dreamer 3 was the first one to solve the Minecraft Diamond challenge—successfully mining in-game gems on its own. And Dreamer 4 went a step beyond that by learning to mine diamonds from an offline data set of recorded game-play videos, without ever interacting with the game directly. More recently, he’s begun to migrate his agents out of the virtual world and into physical reality. His DayDreamer project used the Dreamer algorithm to let robots operate themselves in novel environments and react to new experiences (such as being pushed over) without any specific training. Today, Hafner is working on his new startup, which he left Google DeepMind to form in the fall of 2025. Though he’s coy about his next steps, it’s clear he’s dreaming big: “I was interested in solving a problem,” he hints, “that would change the world.” Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven archive page Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. By Will Douglas Heaven archive page AI is more likely than humans to form biases when hiring AI doesn’t just learn stereotypes from its training. It can cook up new ones, too. By Michelle Kim archive page Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. By Grace Huckins archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.