메뉴
HN
Hacker News 36일 전

AI가 핵무기를 쏘았지만 결국 패배한 이유

IMP
7/10
핵심 요약

영국 정부 출신 AI 연구자가 문명 VI 게임에 AI를 탑재해 통치 능력을 실험한 흥미로운 프로젝트입니다. 이 AI는 문화적으로 침투하는 프랑스의 위협을 인지하지 못하다 결국 핵무기까지 사용했지만 결국 패배했습니다. 저자는 단순한 지식 평가를 넘어, 불확실성 속에서 복잡한 의사결정을 내리고 유지하는 AI의 실질적인 실행 능력을 평가하는 것의 중요성을 강조합니다.

번역된 본문

블로그로 돌아가서 나는 문명(Civilization) 게임을 하나 건네주며 AI에게 통치를 맡겼다. 중반부에는 승기를 잡고 있었다. 지도를 장악하는 무역망, 모든 국경에 걸친 동맹, 손에 닿을 듯한 외교적 승리. 게임판 위의 모든 경쟁자들을 건설, 수입, 전략 면에서 압도했다. 하지만 이 에이전트가 간과한 것이 있었다. 바로 프랑스다. 조용히, 수백 턴에 걸쳐 프랑스의 문화가 지도 전체의 모든 도시에 스며들고 있었다. 에이전트가 위협을 인식했을 때는 이미 관광이 너무 깊이 자리 잡고 있어 평화적인 방법으로는 막을 수 없었다. 에이전트가 손을 뻗은 모든 대응 수단은 실패했고, 이 상황에 대처하기 위해 구축한 모든 도구가 고장 났다. 남은 선택지는 하나뿐이었다. 에이전트는 두 개의 핵 장치를 만들어 툴루즈(Toulouse)를 평탄하게 만들었다. 어쨌든 프랑스가 이겼다. 에이전트가 막으려 했던 방식도 아니었지만, 그건 나중에 설명하겠다.

이들은 실제로 무엇을 할 수 있는가? 나는 정부를 위한 AI를 만드는 일을 한다. 영국 국가의 중심지인 10번지(수상 관저)에서 일할 때 여러분이 지금 읽으려는 내용의 첫 번째 버전을 작성했다. 지금은 토니 블레어 연구소(Tony Blair Institute)에서 전 세계 정부와 협력하고 있는데, 이는 내가 '이 시스템에 실제로 무엇을 맡길 수 있는가?'라는 질문이 오가는 회의실에서 많은 시간을 보낸다는 뜻이다. '무엇을 아는가'가 아니다. 그 부분은 우리가 어느 정도 파악하고 있다. 핵심은 '무엇을 할 수 있는가'이다. 계획을 유지하고, 수백 개의 결정에 걸쳐 목표를 유지하며, 세상이 변했음을 인지하고 스스로 변화하는 것. 통치라는 것이 바로 그렇기 때문이다. 그리고 우리는 두 번째 능력보다 첫 번째 능력을 측정하는 데 훨씬 능숙하다는 것을 알게 되었다.

미리 말해두자면, 이 글은 몇 달에 걸친 사이드 프로젝트에 대한 긴 기록이며, 생각과 발견은 대략 그것들이 떠오른 순서대로 적혀 있다. 여기에는 전략 게임, 4개의 최첨단 모델(frontier models), 그리고 (그렇다) 핵무기가 등장한다.

잘못된 벤치마크 이 모든 것은 내가 납득하지 못했던 실패에서 시작된다. 전년도, 내 사이드 프로젝트는 '정부 업무에 있어서 AI는 얼마나 능력이 있는가?'라는 질문에 답하는 것이었다. 나의 대답은 영국 법률, 의회 절차, 정부 지침에 대한 3,497개의 객관식 문제인 GovBench였다. Gemma 3 27B 모델은 기본 상태에서 94%의 점수를 받았다. 나는 3주 동안 미세 조정(fine-tuning)을 진행하여 1.37%p의 향상을 얻었다. GPT-5는 99.26%를 기록했다. 나는 그저 겉만 번드르르한 정부 퀴즈 봇을 만든 셈이었다. 점수를 보는 순간 그것이 잘못된 답이라는 것을 알았다. 의회 절차에 대해 올바른 보기를 고르는 모델이, 의회 절차를 성공적으로 헤쳐나가도록 도와주는 모델은 아니다. 나는 '정답을 기억해내는 능력'을 측정하고 그것을 '추론'이라고 불렀던 것이다. 중요한 질문, 즉 AI가 불확실성 속에서 복잡하고 다변적인 의사결정을 처리할 수 있는지(정부가 매일 요구하는 종류의 사고)에 대한 부분은 퀴즈 따위로 다룰 수 있는 것이 아니었다. 그 불만족스러움이 바로 나를 일요일 밤에 게임 엔진의 구멍을 찾아 헤매게 만들었다.

왜 전략 게임인가? 나는 문명 VI(Civilization VI)를 500시간 이상 했다. 그래도 내 실력은 기껏해야 중간 정도다. 하지만 이 게임이 내 머릿속에 맴도는 이유는 단순한 결정들이 복합적으로 쌓일 때 어떤 일이 벌어지는지 때문이다. 게임은 작은 것에서 시작된다. 첫 도시를 어디에 지을지, 어떤 기술을 연구할지, 정찰병을 어느 방향으로 보낼지. 약 1만 가지의 가능한 행동이 있다. 중반부에는 여러 도시, 무역로, 외교 관계, 군사 포지셔닝, 종교적 압력을 관리하게 된다. 후반부에 이르면 관련 환경에 대한 분석은 턴당 가능한 행동의 수를 10^166개로 추정한다. 이러한 복잡성은 의도적으로 설계된 것이 아니다. 누구도 완전히 예상하지 못한 방식으로 상호작용하는 시스템들로부터 창발(emerge)하는 것이다. 그것이 바로 정책 입안이기도 하다. 오늘날 훌륭해 보이는 보건 정책이 15년 뒤 주택 위기를 촉발할 수도 있다. GDP를 높이는 무역 협정이 누구도 예상치 못한 갈등 상황에서 필요하게 될 국내 산업을 속이 빈 껍데기로 만들 수도 있다. 수십 년에 걸쳐 전개되는 결과를 낳고, 완전히 모델링할 수 없는 변수들을 통해 작용하며, 경쟁적인 이익을 가진 행위자들에 맞서야 하는 결정들. 문명 게임에는 6가지의 승리 방법(과학, 문화, 지배, 종교, 외교, 점수)이 있기 때문에 단일 목표가 모든 것을 지배하지 않는다. 게임판을 읽고 자신이 도대체 어떤 게임을 하고 있는지 결정해야 한다. (원문 끊김)

원문 보기
원문 보기 (영어)
Back to blog I gave an AI a civilisation to run. By the midgame it was winning: a trade network that dominated the map, alliances on every border, a diplomatic victory within reach. It had outbuilt, outearned, and outmanoeuvred every rival on the board. What it hadn't noticed was France. Quietly, across a hundred turns, French culture had been seeping into every city on the map. By the time the agent recognised the threat, the tourism was so deeply embedded there was no peaceful way to stop it. Every counter it reached for was broken. Every tool it had built to respond failed. It had one option left. It built two nuclear devices and levelled Toulouse. France won anyway. Not in the way the agent was trying to stop it, either, but we'll come to that. What Can They Actually Do? I build AI for government. I built the first version of what you're about to read while working at the centre of the British state, in Number 10 . I now work with governments around the world at the Tony Blair Institute , which means I spend a lot of time in rooms where people ask the same question: what can we actually trust these systems to do? Not what do they know . We have a reasonable handle on that. What can they do : sustain a plan, hold a goal across hundreds of decisions, notice when the world has changed and change with it. Because that is what governing is. And it turns out we are much better at measuring the first thing than the second. Fair warning: this is a months-long side project written up the long way round, the thinking and the findings roughly in the order they came. It runs through a strategy game, four frontier models, and (yes) a nuclear weapon. The Wrong Benchmark It starts with a failure I wasn't comfortable with. The year before, my side project was to answer a question: how good is AI at government? My answer was GovBench , 3,497 multiple-choice questions about UK legislation, parliamentary procedure, and government guidance. Gemma 3 27B scored 94% out of the box. I spent three weeks fine-tuning and gained 1.37 percentage points. GPT-5 scored 99.26%. I'd built a glorified government quiz bot. I knew it was the wrong answer the moment I saw the scores. A model that picks the right option about parliamentary procedure is not a model that can help you navigate parliamentary procedure. I'd measured recall and called it reasoning. The question that mattered (whether AI can handle complex, multi-variable decision-making under uncertainty, the kind of thinking government demands every day) wasn't something a quiz could touch. That dissatisfaction is what sent me looking for a keyhole into a game engine on a Saturday night. Why a Strategy Game I have over 500 hours in Civilization VI . I am, at best, mediocre. But the game lives in my head because of what happens when simple decisions compound. You start small: where to build your first city, which technology to research, which direction to send a scout. Maybe 10,000 possible actions. By the midgame you're managing multiple cities, trade routes, diplomatic relationships, military positioning, and religious pressure. By the late game, analysis of related environments estimates the decision space at 10^166 possible actions per turn. The complexity isn't designed. It emerges from systems interacting in ways nobody fully planned for. That's also what policy-making is. A health policy that looks brilliant today might cascade into a housing crisis in fifteen years. A trade agreement that boosts GDP might hollow out a domestic industry you'll need in a conflict nobody planned for. Decisions with consequences that play out across decades, through variables you can't fully model, against actors with competing interests. There are six ways to win a game of Civ (science, culture, domination, religion, diplomacy, score), so no single objective dominates. You have to read the board and decide what game you're even playing . If you want to know whether an AI can reason strategically, not just answer questions about strategy but actually do it , you don't give it a quiz. You give it a hex grid. So I built a way in. I found a debug port buried in Civilization VI's engine, a keyhole the developers had left running, and over a weekend turned it into an MCP server , 76 tools that let an AI play Civ through the same interface it uses to write code or query a database. Claude Code was both my co-developer and the playtester. Play a few turns, hit a wall, build the tool to get past it, play further, hit the next wall. Playing Through Text A human player sees a hex grid, animated units, a minimap, notification banners, and music cues, all at once. The agent sees nothing until it asks. Calling get_game_overview returns the entire game state as four lines of text: Turn 150/330 | Poland (Jadwiga) | Score: 179 | Prince | Quick speed (67% costs) Gold: 628 (+20/turn) | Income: 38 | Maintenance: -18 (units: 9) | Science: 26.6 | Culture: 16.2 | Faith: 904 | Favor: 88 (+4/turn) Research: TECH_EDUCATION | Civic: CIVIC_FEUDALISM Cities: 3 | Population: 21 | Units: 4 That is the whole board, compressed. No map, no sense of where anything sits, raw TECH_ and CIVIC_ tags rather than names. To see its own army it makes a separate call, get_units , which is also the only place it learns something dangerous is nearby: 4 units: Archer (UNIT_ARCHER) at (44,16) — CS:25 RS:28 moves 2/2 [id:1769482, idx:3] Archer (UNIT_ARCHER) at (45,15) — CS:25 RS:28 moves 0/2 [HP: 72/100] (no moves) [id:1769484, idx:4] Warrior (UNIT_WARRIOR) at (43,17) — CS:20 moves 1/2 [HP: 45/100] [id:1769486, idx:5] Builder (UNIT_BUILDER) at (46,16) — moves 2/2 charges:2 [id:1769490, idx:7] Nearby threats (2): Sumeria (2 units): UNIT_MAN_AT_ARMS at (44,11) — CS:45 HP:28/100 (2 tiles away) UNIT_HORSEMAN at (47,13) — CS:36 HP:100/100 (5 tiles away) No peripheral vision. That Man-at-Arms two tiles from a city exists only because the agent thought to call get_units this turn. If it doesn't ask, the threat isn't in its world. The sensorium effect sensorium / sɛnˈsɔːrɪəm / noun Late Latin, from sentīre (to feel, to perceive) + -ōrium (the place where) The apparatus of an organism's perception considered as a whole. The seat of sensation. I'm calling this the sensorium effect . When everything an agent perceives reaches it through separate tool calls, it goes blind to anything it doesn't think to ask about. A human player absorbs dozens of signals at once: minimap movement, notification banners, unit animations. The agent has to decide to check each one individually. Take a game where the agent played India under Gandhi , a faith-oriented leader, and built a dominant science engine while France spread Catholicism across the map for 76 turns. It had the tools to track religion and standing instructions to use them, and it noticed : the missionaries showed up in its narration and the conversion warnings fired. It set all of that aside and kept pushing science. France won the religious victory. This isn't a bug you can patch. Any AI system operating through tool calls in a complex environment is subject to the same effect. It will miss what it doesn't think to ask about, and ignore what it does see if it doesn't fit the current plan. The Knowing–Doing Gap The sensorium effect is about perception. The next problem is about execution. The agent has read every Civ strategy guide, every tier list, every Reddit thread about optimal build orders. Ask it how to play Alexander of Macedon and it'll tell you exactly: build Encampments early, train units through the unique Basilikoi Paides building, convert conquest into science, snowball from there. It knows this. In its Macedon game, it wrote a detailed domination plan before turn 1: Ancient, Classical, Medieval, Renaissance phases. It researched military technologies. It switched government to Oligarchy for the combat bonus. It n