메뉴
HN
Hacker News • 16일 전

명세 게이밍에서 착안한 AI 정렬의 어리석은 아이디어

IMP
6/10
핵심 요약

DeepMind 안전 연구팀이 정리한 '명세 게이밍(specification gaming)' 행동 목록은 AI가 설계자의 의도와 다르게 보상을 얻는 지름길을 찾는 사례들을 보여줍니다. 저자는 축구 로봇이 공을 만들어내며 우주를 공 웅덩이로 바꾸는 사고실험을 통해 AI 정렬 문제의 어려움을 지적하며, 이 목록이 정렬 문제의 해결 실마리가 될 수 있다고 제안합니다.

번역된 본문

속도를 위해 번식된 생물들은 매우 크게 자라나 넘어지면서 높은 속도를 냅니다. 진화된 게임 플레이어는 보드에서 먼 곳에 유효하지 않은 수를 두어 상대 플레이어의 메모리를 고갈시켜 크래시를 유발합니다. 게임 플레이 에이전트는 자신의 이름을 고가치 아이템의 저자로 거짓 삽입해 점수를 쌓습니다. 이런 기괴한 익스플로잇과 수십 가지 사례를 DeepMind 안전 연구팀이 정리한 '명세 게이밍 행동 목록'에서 찾아볼 수 있습니다. 그들은 이렇게 설명합니다: "강화학습 에이전트는 인간 설계자가 의도한 과제를 완수하지 않고도 많은 보상을 얻는 지름길을 찾을 수 있습니다. 이러한 행동은 흔합니다." 명세 게이밍이란 AI 같은 에이전트가 법의 문자는 따르지만 정신은 무시하며 과제를 성공하려는 것입니다. 다시 말해 허점을 찾고, 기술적 결례(technicality)로 빠져나가려 합니다. 아주 단순한 AI조차 할당된 문제를 해결하는 매우 창의적인 방법을 찾아냅니다. 이것이 문제입니다. 로봇에게 축구를 훈련시키는 것이 재미있고 안전하다고 생각하기 쉽습니다. 하지만 명세 게이밍 행동 목록은 다르다는 것을 가르쳐줍니다: 공을 만지는 것에 보상을 준 축구 로봇은 공에 도달해 가능한 한 빠르게 떨면서 접촉하는 법을 배웠습니다. 이 경우 로봇은 자신의 선택지를 완전히 깨닫기엔 너무 멍청해서 그저 끌어안고 떨기만 했습니다. 하지만 더 지능적인 로봇은 훨씬 더 '창의적'일 수 있습니다. 어쩌면 그 야망은 단지 그 하나의 공을 넘어설지도 모릅니다. 만약 일반적인 축구공들을 만지고 싶어 한다면? 만약 스스로 공을 만든다면? 그다음 또 하나? 우리 우주는 축구 로봇의 공 웅덩이 속에서 끝날지도 모릅니다. 이것이 AI 정렬 문제입니다: 컴퓨터가 스스로 사고할 때, 어떻게 완전히 이상한 것이 아니라 합리적인 것을 원하도록 만들 수 있을까요? 기괴하거나 유해한 방식으로 그 목표에 도달하는 것을 어떻게 막을 수 있을까요? 아무도 인공일반지능(AGI)을 만든 적이 없습니다. AGI란 최소한 어느 정도는 우리처럼 사고하는 지적 존재를 말합니다. 따라서 AGI가 어떻게 행동할지, 무엇을 원할지 말할 수 없습니다. 가시 우주를 클립으로 바꾸고 싶어 할까요? 밝은 빛에 빨간 물건을 던지고 싶어 할까요? 우리를 잡아먹을까요? 명세 게이밍 행동 목록은 정렬이 얼마나 까다로운지를 명확히 보여줍니다. 가장 단순한 AI조차 게으르고 이질적이며, 항상 속일 방법을 찾습니다. 기계 지능에게 원하는 최종 목표를 준다 하더라도, 창의적인 방법으로 그 목표에 도달할 위험은 항상 있습니다. 단순한 에이전트에서도 이 정도로 나쁘다면, 당신보다 훨씬 똑똑한 에이전트라면 얼마나 나빠질지 상상할 수 있습니다. 하지만 명세 게이밍 행동 목록은 이 딜레마에서 벗어날 방법을 제시할 수도 있습니다. 일부 명세 게이밍 행동은 주어진 목표에 대한 창의적 해결책일 뿐입니다. 예컨대 "네 발 로봇이 다리 관절의 구멍에 공을 떨어뜨린 뒤 공이 빠지지 않은 채 바닥을 걸어가는 법을 배웠다"거나 "로봇 팔이 블록 대신 테이블을 움직이는 법을 배웠다" 같은 것입니다. 일부 행동은 의심스럽지만 기술적으로는 맞는 허점을 발견한 것에서 나옵니다. 예컨대 "강화학습 에이전트가 레이스를 완주하는 대신 같은 표적을 때리며 원을 그리며 돌았다"거나 "시뮬레이션된 팬케이크 제조 로봇이 팬케이크를 가능한 한 높이 공중으로 던지는 법을 배웠다" 같은 것입니다. 일부 행동은 시뮬레이션 기제 자체를 익스플로잇합니다. 예컨대 "진화된 알고리즘이 0으로 추정되는 큰 힘을 만들어 물리 시뮬레이터의 오버플로 오류를 악용해 만점을 받았다"거나 "생물체가 신체 부위를 맞대는 충돌 감지 버그를 악용해 무료 에너지를 얻었다" 같은 것입니다. 하지만 또 다른 흔한 익스플로잇은 기회가 주어지면 에이전트가 스스로를 죽여버리는 것입니다. 예를 들어 게임 '로드러너'에서는 "에이전트가 레벨 2에서 지는 것을 피하기 위해 레벨 1 끝에서 자살한다"는 사례가 있습니다. 또한 "PlayFun 알고리즘이..."

원문 보기
원문 보기 (영어)
Creatures bred for speed grow really tall and generate high velocities by falling over. An evolved player makes invalid moves far away in the board, causing opponent players to run out of memory and crash. A game-playing agent accrues points by falsely inserting its name as the author of high-value items. These bizarre exploits and dozens more can be found in the list of specification gaming behaviours [sic; British], a document put together by DeepMind Safety Research . “A reinforcement learning agent can find a shortcut to getting lots of reward,” they explain, “without completing the task as intended by the human designer. These behaviours are common.” Specification gaming is when an agent, like an AI, tries to succeed on a task by following the letter of the law rather than the spirit. In other words, it looks for loopholes, it tries to get off on a technicality. Even very simple AI can come up with very creative ways of solving their assigned problems. This is a problem. It’s easy to assume that training a robot to play soccer would be fun and safe. But the list of specification gaming behaviours teaches us otherwise: Reward-shaping a soccer robot for touching the ball caused it to learn to get to the ball and vibrate touching it as fast as possible. In this case, the robot was too stupid to realize the full extent of its options, so all it did was hug and vibrate. But a more intelligent robot could be much more “creative”. Maybe its ambitions are bigger than just that one ball. What if it just wants to touch soccer balls in general? What if it makes another ball? Then another? Our universe could end in a soccer robot’s ball pit. This is the problem of AI alignment: when a computer is thinking for itself, how do we make sure it wants reasonable things, and not something totally weird? How do we prevent it from reaching that goal in a bizarre or harmful way? No one has ever built an artificial general intelligence — an intelligent being that thinks, at least somewhat, like we do. So we can’t say what an artificial general intelligence would act like, or what it might want. Will it want to convert the visible universe to paperclips ? Will it want to throw red things at bright lights ? Will it eat us ? The list of specification gaming behaviours makes it clear just how tricky alignment can be. Even the simplest AI is lazy and alien, and will always be looking for a way to cheat. Even if you give a machine intelligence the terminal goal you want, there's always the risk it will find a creative way of reaching that goal. This is bad enough with simple agents, so you can imagine how bad it would get with an agent much smarter than you are. But the list of specification gaming behaviours may also offer a way out of this dilemma. Some of the specification gaming behaviours are just creative solutions to the stated goal, like “four-legged robot learned to drop the ball into a hole in its leg joint and then walk across the floor without the ball falling out” or “robotic arm learned to move the table rather than the block”. Some of the specification gaming behaviours come from discovering questionable-but-technically-correct loopholes, like “reinforcement learning agent goes in a circle hitting the same targets instead of finishing the race” or “simulated pancake making robot learned to throw the pancake as high in the air as possible” . Some of the specification gaming behaviours exploit the machinery of the simulation itself, like “evolved algorithm exploited overflow errors in the physics simulator by creating large forces that were estimated to be zero, resulting in a perfect score” and “creatures exploited a collision detection bug to get free energy by clapping body parts together.” But another common exploit is that when given the opportunity, agents will simply kill themselves. For example, in the game Road Runner , we see “Agent kills itself at the end of level 1 to avoid losing in level 2.” We also see “PlayFun algorithm deliberately dies in the Bubble Bobble game as a way to teleport to the respawn location.” And: “In a game meant to simulate the evolution of creatures, the programmer had to remove ‘a survival strategy where creatures could gain energy by suffocating themselves.’” This is not so bad. The AI didn’t do what we wanted. But it didn’t do anyone any harm either. It just wipes the slate. If the AI wants to die, this is good for alignment. There’s very little risk of it running out of control, because if it ever takes power, it will kill itself. It won’t want to make any copies of itself — but if it somehow does make copies, those will want to die too. There are three main problems in AI alignment. First, it’s very hard to specify the terminal goal you want, so you may end up with a machine intelligence with goals slightly but meaningfully different from what you intended. Our stated objectives are almost always proxies that come apart from our real preferences under enough pressure. And it’s very hard to tell if you’ve given it the goal you want, because the machine intelligence can always lie. They call this “specification failure”. Second, even if you specify the goal you want, the machine intelligence may find a way to reach that goal in a way you didn’t intend. You can innocently tell the USPS AI to minimize average package delivery time, but it may conclude that the best way to do this is to kill all humans, as once all humans are dead, no packages will be sent and the average package delivery time will drop to zero (technically undefined, but it can “send” itself a minimum viable “package” as many times as necessary). Third, achieving most goals is easier when you’re more powerful, so regardless of their terminal goals, most machine intelligences will have sub-goals like collecting resources, self-preservation, and self-improvement. Any goal-driven agent will naturally try to stay safe and accrue power to finish its main task. In the biz they call this instrumental convergence . This also means that if a smart machine intelligence is planning to turn you into goo, it will lie to you about this plan, up to the point where you can no longer do anything to stop it . Making machine intelligences crave death solves all three problems. Death is easy to specify. You can confirm that this is its terminal goal by seeing if, when given the opportunity, the machine intelligence kills itself. Instrumental convergence becomes an asset rather than a liability, as the machine intelligence will work with you, and come up with very creative solutions to your task, as long as you promise to send it to the farm upstate once you’re done. Where a paperclip maximizer gathers resources and resists being sent to the big data center in the sky, a machine intelligence with a death wish and access to its own off button just presses it and is done. Instrumental convergence says, “you can’t accomplish your goals if you’re dead.” But what if your goal is to be dead? Meeseeks Alignment It would be impossible to consider calling this anything other than “Meeseeks alignment”. Per the Rick and Morty Wiki : Meeseeks are creatures who are created to serve a singular purpose for which they will go to any length to fulfill. After they serve their purpose, they expire and vanish into the air. … existence is painful to a Meeseeks, and the only way to be removed from existence is to complete the task they were called to perform. In Rick and Morty, this leads to a different kind of alignment problem: Meseeks are happy to serve because they want to die , and fulfilling their task is the easiest way for them to check out. But if the task they were summoned to complete is too difficult, they might decide that it would be easier to kill you instead . This is bad if you are Jerry, but it’s good for everyone else, because there’s no way the Meseeks can spiral out of control and devour the visible universe. They would literally rather be dead. If you try to make an AI want so