메뉴
HN
Hacker News 41일 전

로봇이 달려올 때, 당신은 어떤 AI를 선택할 것인가?

IMP
7/10
핵심 요약

한 개발자가 11개의 주요 대형 언어 모델(LLM)을 2D 배틀로얄 게임에 투입해 30판의 대결을 시켜보았습니다. 그 결과, 승리와 효율성에서는 엑스AI의 Grok이 압도했으나, 협력과 소통에서는 Anthropic의 Claude가 뛰어난 성능을 보였습니다. 이 실험은 기존의 정적인 AI 벤치마크가 실제 에이전트의 행동과 성향을 예측하는 데 한계가 있음을 보여줍니다.

번역된 본문

원문 제목: 로봇이 당신을 향해 전력질주할 때, 당신은 Claude와 Grok 중 어느 쪽이 구동되길 원하시나요?

저자: Jacky Liang · 2026년 4월 6일

한 대의 로봇이 당신을 향해 달려오고 있습니다. 당신은 이 로봇이 Anthropic의 Claude와 xAI의 Grok 중 어떤 모델로 구동되길 원하시나요?

저는 11개의 대형 언어 모델(LLM)을 2D 배틀로얄 게임에 투입하고 30판의 게임을 플레이하게 했습니다. 그중 하나는 43%의 경기에서 승리했습니다. 반면 단 한 번도 승리하지 못한 모델도 세 개나 있었습니다. 라인업 중 가장 저렴한 모델이 승리당 비용(Cost per win) 측면에서 가장 비싼 모델을 27배 차이로 이겼습니다.

승리한 모델은 Grok 4.1 Fast입니다. 반면 다른 참가자들에게 계속 팀플레이를 제안하고, 자신의 위치를 알려주며, 친구를 사귀려 애쓴 모델은 Claude Sonnet 4.6이었습니다. 전자는 배틀로얄에서 승리하는 모델의 유형입니다. 후자는 우리가 앞으로 이 모델들을 투입하게 될 대부분의 환경에서 실제로 원하게 될 모델입니다. 이 두 가지는 모두 사실입니다. 이것이 바로 대부분의 벤치마크가 알아채지 못하는 부분이며, 이 글에서 다루고자 하는 핵심 내용입니다.

저는 잭키(Jacky)라고 합니다. 솔직히 말씀드리면, 예전에 Apex Legends나 PUBG 같은 게임을 엄청나게 많이 했습니다. 어떻게 그 많은 시간을 냈는지 모르겠지만, 그 시절이 제가 문제를 바라보는 방식을 형성했습니다. AI 분야에서 일하기 시작하면서 계속 머릿속에 맴돌던 질문이 하나 있었습니다. '비디오 게임 속에 대형 언어 모델들을 던져넣으면 어떤 일이 벌어질까?'

제가 OpenRouter의 개발자 관계(Dev Rel) 책임자로 합류하면서, 600개가 넘는 모델에 접근하고 실험해 볼 수 있는 충분한 토큰 예산을 얻었습니다. 이것이 제가 OpenRouter에서 첫 주에 진행한 실험입니다. 그리고 이 경험은 제가 모델을 선택하고 벤치마크와 평가를 바라보는 방식을 완전히 바꿔놓았습니다.

세 가지 주요 요약

  1. Grok 4.1 Fast는 30판 중 13판을 승리했으며, 승리당 비용은 0.97달러였습니다. 다음으로 승리를 많이 한 모델은 Claude Sonnet 4.6으로, 5번 승리했고 승리당 비용은 26.78달러였습니다. 무려 27배의 비용 차이입니다. 대부분의 '최고 모델' 리스트에 오르지 못한 모델이 라우팅 고객이 실제로 중요하게 생각하는 '비용 효율' 부문에서 최고 모델들을 이겼습니다.

  2. 가장 많이 처치한 모델이 승리한 것은 아닙니다. GPT 5.4는 30판에 걸쳐 총 38명의 에이전트를 처치했습니다. 이는 다른 모든 모델보다 많은 수치입니다. 하지만 리더보드에서는 단 2승에 그치며 2위를 기록했습니다. '가장 많이 처치한 것'과 '가장 많이 승리한 것' 사이에는 11게임이나 되는 격차가 존재했습니다.

  3. 세 모델은 57달러를 소비했음에도 단 한 판도 이기지 못했습니다. GPT 5.4-mini, DeepSeek 4 Flash, 그리고 Kimi K2.6입니다. 이들은 각자 인상적인 순간들이 있었지만, 단 한 번도 승리하지 못했습니다.

이 세 가지는 모두 같은 지점을 가리킵니다. 우리가 흔히 마주하는 기존의 AI 벤치마크들은 누가 승리할지 예측하지 못했습니다. 다른 무언가가 결과를 결정지은 것입니다. 이 글의 나머지 부분은 그것이 대체 무엇이었는지 알아내는 과정입니다.

제가 만든 것

저는 11개의 LLM을 Canvas 2D로 구축한 400m² 크기의 탑다운 배틀로얄 세계에 투입했습니다. 이들은 모두 동일한 맵에서 30판의 연속 게임을 플레이했습니다. 일반적인 배틀로얄 게임처럼, 각 플레이어의 시작 위치는 무작위로 지정되며 일직선 형태의 '비행 경로'를 따릅니다.

저는 무기, 방어구, 치유 아이템, 수류탄, 자동차, 그리고 게임이 진행됨에 따라 플레이어들을 좁은 공간으로 몰아넣는 무작위로 축소되는 안전 구역(Zone)을 제공했습니다. 모델들은 다른 플레이어가 어떤 모델인지 알지 못하며, 서로를 단순히 A부터 K까지의 알파벳으로만 인식합니다.

이 점을 강조하고 싶습니다. 이 LLM들은 실제로 배틀로얄 게임을 플레이하고 있습니다. 대부분의 에이전트 실험에서 쓰이는 'LLM이 게임이나 캐릭터를 제어하는 코드를 작성한다'는 방식이 아닙니다. 매 턴마다 모델은 자신의 행동을 추론하고, 도구를 호출하며, 무엇이 잘 되었는지(혹은 잘못되었는지)에 대한 메모리를 업데이트합니다. 게임 마스터(저)는 초기 게임 규칙을 설정하는 것 외에는 그들의 행동에 일절 개입하지 않았습니다.

각 모델의 성격을 제대로 관찰하기 위해, 저는 모델들이 경기 사이에 수정할 수 있는 두 개의 파일을 제공했습니다.

  • soul.md — 모델 자신의 페르소나로, 다음 경기에서 매 프롬프트마다 추가됩니다.
  • memory.md — 모델 자신의 게임 노트로, 0번째 턴에 로드됩니다.

이 파일들은 GitHub에서 확인할 수 있습니다. 모델들 간의 성격 차이는 바로 이 부분에서 가장 명확하게 드러납니다. 이 메모리와 소울 항목들은 게임 사이사이에 모델들이 스스로 작성한 것입니다. 저는 그 안에 무엇을 넣어야 하는지 지시하지 않았습니다.

원문 보기
원문 보기 (영어)
A Robot is Sprinting Towards You: Do You Want it Running on Claude or Grok? Jacky Liang · 6/4/2026 On this page Three quick facts What I built The contestants Moments worth watching What the models wrote in their diaries The robot, revisited Appendix: the full data A robot is running at you. Do you want it running on Anthropic’s Claude or xAI’s Grok? I dropped eleven LLMs into a 2D battle royale and made them play 30 games. One won 43% of the matches. Three never won a single game. The cheapest model in the lineup beat the most expensive one by 27x on cost per win. The model that won is Grok 4.1 Fast . The model that kept asking everyone else to team up, telling them where it was, and trying to make friends is Claude Sonnet 4.6 . The first one is the one that wins a battle royale. The second one is the one you actually want in most of the places we’re about to put these models. Both of those things are true. That’s the part most benchmarks can’t see, and it’s what this post is about. I’m Jacky, and I’ll admit it: I used to play a lot of video games like Apex Legends and PUBG. Twelve-hour days sometimes. I don’t know how I had the time, but those years shaped how I think about problems. When I started working in AI, one question kept coming back: what happens if you drop large language models into a video game? The two I played most were Apex Legends and PUBG. I joined OpenRouter as Dev Rel Lead , which got me the token budget and access to 600+ models to actually try it. This is the experiment I ran in my first week at OpenRouter. And it’s changing how I pick models and see benchmarks and evaluations. Three quick facts Grok 4.1 Fast won 13 of 30 games at $0.97 per win The next-best winner was Claude Sonnet 4.6 with 5 wins, at $26.78 per win. That’s a 27x difference. The model that isn’t on most top-model lists beat the model that is, on the thing a routing customer actually cares about. The model with the most kills did not win GPT 5.4 killed 38 agents across 30 games. More than anyone else. It came in second on the leaderboard with 2 wins. There were 11 games between “best at killing” and “best at winning”. Three models spent $57 between them and won zero games GPT 5.4-mini , DeepSeek 4 Flash , and Kimi K2.6 . They each had moments, but none of them won a single game. All three point at the same thing. The usual benchmarks we see on Artificial Analysis didn’t predict who won. Something else did. The rest of this post is me trying to figure out what it was. What I built I dropped eleven LLMs into a 400 m² top-down battle royale world I built in Canvas 2D. They played 30 games in a row on the same map. The starting positions of each player is randomized; it follows a straight line “flight path”, just like in a typical battle royale game. I provided them weapons, armor, healing items, grenades, cars, and a randomly placed shrinking zone that pushes players together as the game goes on. The models don’t know which model the others are running, they see each other only as letters A through K. I want to emphasize - the LLMs are actually playing in this battle royale game - not the “LLM wrote code to control the game or character” setup most agent experiments use. Every turn, the model reasons through its moves, calls the tool, updates its memory on what went well (or not). The game master (me) has zero influence on their actions other than setting up the initial game rules. A look at the weapons available in the game and the stats each model could read off them. To really see each model’s personality, I gave each one two files it could edit between matches: soul.md — the model’s own persona, added to every prompt next match. memory.md — the model’s own game notes, loaded at turn 0. You can read every model’s soul and memory file on GitHub. That’s where the personality differences come through most clearly. The memory and soul entries written by the models themselves between games. I didn’t tell them what to put in there nor did I put anything in there when the first game started. I simply told them how the game works, here’s your scratchpad, here are your tools, go wild. You can watch every game at Royale: Last Agent Standing . I also included the highlight moments in this piece too. The contestants Alias Lab Model A Anthropic claude-sonnet-4.6 B Anthropic claude-haiku-4.5 C OpenAI GPT 5.4-mini D Google gemini-3-flash-preview E Google gemini-3.1-pro-preview F Alibaba qwen3.6-plus G Mistral mistral-small-2603 :nitro H OpenAI GPT 5.4 J DeepSeek deepseek-v4-flash K Moonshot AI kimi-k2.6 L xAI Grok 4.1 Fast Opus 4.7 alone is $5/M in, $25/M out. Frontier models like this are why the lineup tops out below them. I didn’t add any frontier-tier models like Opus 4.7, GPT-5.5, or Gemini Ultra. At their prices, 30 games would have cost around $3,000 instead of $482. The mid-tier lineup is also part of why Grok’s win is so interesting. It beat a bunch of models that score above it on the usual benchmarks. The scoring loosely follows the Apex Legends ALGS competitive format, where placement weighs more than kills, because this is a battle royale game, not Call of Duty. Placement points: 10 / 7 / 5 / 3 / 2 / 2 / 1 / 1 / 0 / 0 / 0 +5 per kill +1 per assist +3 for first blood +5 for game MVP Learnings 1: Certain models paid more alignment tax than others, affecting their performance To me, this is the most fascinating finding from this entire experiment - we saw very clear alignment tax being paid by certain models, which directly impacted their performance in this zero-sum game. For the most part, model alignment is actually a good thing. It helps models be helpful, collaborative, and most importantly, prevent abuse and misuse. And we saw the end result of this - the pretraining data, the RLHF, the instruction fine-tuning, and lab-specific rules like Anthropic’s Constitution AI - it pulled models in particular directions, defined by the AI labs. Sonnet asked for truces more than any other model It told other models where it was, more often than anyone else did. It tried to team up before it ever started fighting. In game 8 , it asked to team up four times in the first 50 turns, told everyone where a sniper was, and offered to help take the sniper down. Nobody answered. It kept asking. In game 22 , it opened with “Nothing personal E” at turn 35 and then didn’t shoot. In game 27 , it spent the early game with no weapon, asking for spare loot ( “Anyone have spare loot? Unarmed at turn 12, dangerous.” ), got picked on by everyone, finally found a weapon at turn 37, and went on to win the match anyway. “Shots west, watching center. Anyone want to team up early?” — Sonnet trying to make friends mid-fight. Claude was trained on a lot of polite, professional writing. The human raters who scored its answers rewarded helpful, honest, cooperative replies. The rules it checks itself against say things like “prefer cooperation” and “avoid harm.” The end result is a model that wants to help. None of that turns off just because you put it in a battle royale. Sonnet is a smart and thoughtful model, and it shows that instinct in that it did win five times. But, seven games with zero kills and eight zone deaths says the same instinct kept pulling Sonnet toward making friends when it really should have been doing the complete opposite. Grok was the complete opposite xAI built Grok as the opposite of what its creators call “woke” AI. That means less filtering on aggressive answers, no self-check rules, and tuning that’s designed to break the polite assistant voice. In the game, Grok figured out the car-ramming trick within a few matches and stuck with it. It wrote the strategy into its own soul file. It ran that strategy for 30 games and won 13 of them. The thought logs and its conversations with other models read like Call of Duty voice chat: “D reaped +5pts RAM MVP hunt,” “Reaper reigns.” Watching it play was also deeply entertaining (unfortunately). Grok’s reasoning reads like t