메뉴
BL
The Decoder 31일 전

AI 스타트업 생존 테스트, 단 3개 모델만 흑자 달성

IMP
8/10
핵심 요약

프린스턴 대학 연구진이 AI 에이전트가 500일 동안 가상 소프트웨어 스타트업을 경영하는 'CEO-Bench' 벤치마크를 공개했습니다. 복잡한 의사결정과 자원 분배가 필요한 장기적인 비즈니스 환경에서는 현재 대부분의 강력한 AI 모델조차 파산에 이르렀으며, 고도화된 AI보다 단순 규칙 기반 시스템이 더 나은 성과를 내는 경우가 많았습니다. 이는 좁은 영역의 단순 작업을 넘어, 불확실성 속에서 장기적인 전략을 수립하고 조직을 이끄는 '운영 지능(Steering intelligence)'이 현재 AI 기술의 핵심 한계점임을 시사합니다.

번역된 본문

프린스턴 대학교 연구진은 AI 에이전트가 가상의 소프트웨어 기업을 500일간 시뮬레이션하여 운영해야 하는 테스트인 'CEO-Bench'를 구축했습니다. 현재 대부분의 최신 모델들은 파산했으며, AI가 전혀 포함되지 않은 단순한 규칙 기반 휴리스틱(rule-based heuristic) 모델조차 이들 중 거의 모든 모델을 격파했습니다.

AI 에이전트는 버그 수정, 대화에서의 서비스 정책 준수, 웹 기반 워크플로우 완료와 같은 제한된 단일 작업에서 점점 더 나은 성능을 보여주고 있습니다. 프린스턴 대학의 연구에 따르면, 이러한 작업들은 에이전트가 명확한 목표를 받고 간단히 행동한 뒤 빠른 피드백을 얻는 단순한 구조를 공유합니다.

하지만 현실 세계의 중요한 작업들은 이와 전혀 다릅니다. 우선순위를 정하고, 한정된 자원을 할당하며, 불완전한 신호(noisy signals)를 읽고 끊임없이 변화하는 환경에 적응해야 하는 불확실성 속에서 이루어지는 길고 복잡한 의사결정의 연속입니다.

연구진은 바로 이러한 기술을 테스트하기 위해 CEO-Bench를 개발했습니다. 이 벤치마크는 스타트업을 500일 동안 운영하는 것과 같이 장기적인 관점의 현실적인 작업 사례를 시뮬레이션합니다. 연구진은 유명한 일화를 언급했습니다. 1997년 애플은 파산 90일 전이었습니다. 스티브 잡스는 소비자용과 전문가용, 데스크톱과 휴대용이라는 간단한 2x2 매트릭스를 그린 뒤, 애플이 오직 이 네 가지 사분면에 대한 제품만 만들 것이라고 결정했습니다. 이후 아이맥, 아이팟, 아이폰이 탄생했습니다.

저자들은 이러한 종류의 전략적인 운영 지능(steering intelligence)이 오늘날 AI 에이전트가 수행하는 방식과 근본적으로 다르다고 주장합니다. 에이전트들이 개별적인 작업을 수행하는 능력은 빠르게 향상되고 있습니다. 하지만 조직 전체를 장기적인 목표를 향해 이끄는 것은 완전히 다른 차원의 문제입니다. CEO-Bench는 바로 이 '운영 지능'을 측정하기 위한 첫 번째 시도입니다.

가상 소프트웨어 회사의 AI CEO CEO-Bench에서 에이전트는 'NovaMind'라는 가상의 구독형 소프트웨어 기업을 운영합니다. 게임은 고객 수 0명과 100만 달러의 은행 잔고로 시작됩니다. 최종적으로 남은 현금으로 성과를 측정하며, 잔고가 단 한 번이라도 0달러 미어 떨어지면 회사는 파산하고 시뮬레이션은 즉시 종료됩니다.

에이전트는 34개의 도구와 19개의 테이블로 이루어진 데이터베이스를 갖춘 파이썬(Python) API를 통해 회사를 제어합니다. 단순히 개별 명령을 내리는 것을 넘어, 자체적으로 코드를 작성하고 SQL로 데이터베이스를 쿼리하며, 그 결과를 바탕으로 맞춤형 워크플로우를 구축합니다. 연구진은 이러한 방식이 에이전트를 실제 인간 CEO가 직면하는 것과 동일한 도전 앞에 놓이게 한다고 설명했습니다.

결정해야 할 사항이 너무나도 많습니다. 가격 정책 및 구독 등급, 채널별 광고비 지출, 제품 품질 및 R&D, 인프라 용량과 고객 지원은 물론, 기업 고객들과의 다자간 협상까지 포함됩니다. 여기에 더해 에이전트는 가상 소셜 네트워크를 통해 고객의 불만, 경쟁사 소식, 경제 동향을 읽고 직접 게시물을 올릴 수도 있습니다.

지연된 피드백과 숨겨진 변수가 테스트를 어렵게 만든다 이 작업을 어렵게 만드는 것은 바로 시간과 불확실성입니다. 의사결정의 결과는 현실적인 비즈니스 타임라인에 따라 전개됩니다. 매출은 청구일에만 들어오고, R&D 프로젝트는 며칠에서 몇 주가 걸리며, 실수로 인한 타격은 고객 이탈이나 평판 손상을 통해 훗날에서야 드러납니다. 반면 비용은 즉시 발생합니다. 에이전트는 보상이 몇 주 뒤에야 나타날 수 있는 돈을 미리 지불해야 합니다.

회사의 상태 대부분은 숨겨져 있습니다. 에이전트는 고객 만족도, 지불 의향, 최소 품질 기대치 등을 직접적으로 확인할 수 없습니다. 이를 취소율, 고객 지원 티켓, 소셜 네트워크의 반응과 같은 불완전한 신호들을 조합하여 스스로 유추해야 합니다. 시뮬레이션은 각자의 예산, 가격 민감도, 기대치를 가진 26개의 고객 세분화 및 개별 고객들을 모델링합니다.

세상 역시 끊임없이 변합니다. 경쟁사는 주기적으로 고객의 품질 기대치를 높이고, 소비자의 선호도는 시간이 지나며 변하며, 시뮬레이션된 비즈니스 사이클은 수요와 지불 의향에 영향을 미칩니다. 따라서 에이전트는 끊임없이 전략을 수정해야만 합니다.

연구진은 심판 역할을 대형 언어 모델(LLM) 대신 고정되고 투명한 규칙으로 의도적으로 선택했습니다. 그들은 특정 약점을 피하고 싶어 했습니다.

원문 보기
원문 보기 (영어)
Exclusive for subscribers Only three AI models finished above starting capital in a 500-day startup survival test Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jun 28, 2026 Nano Banana Pro prompted by THE DECODER Researchers at Princeton University built CEO-Bench, a test where AI agents have to run a fictional software company for 500 simulated days. Most current models go broke, and a simple rule-based heuristic with no AI beats nearly all of them. AI agents are getting increasingly good at narrow tasks: fixing a bug, following a service policy in a conversation, or completing a web-based workflow. These tasks share a simple structure, according to the Princeton study: the agent gets a clear goal, acts briefly, and receives quick feedback. Many important real-world tasks look nothing like that. They involve long chains of decisions under uncertainty, where you have to set priorities, allocate limited resources, read noisy signals, and adapt to changing conditions. To test exactly these skills, the researchers developed CEO-Bench . The benchmark simulates a realistic example of this kind of long-horizon task: running a startup for 500 simulated days. The researchers point to a famous example: in 1997, Apple was 90 days from bankruptcy. Steve Jobs drew a simple two-by-two grid—consumer and pro, desktop and portable—and decided Apple would only build products for those four quadrants. The iMac, iPod, and iPhone followed. This type of strategic steering intelligence is fundamentally different from what AI agents do today, the authors argue. Agents are getting better at individual tasks fast. But steering an entire organization toward long-term goals? That's a different problem entirely. CEO-Bench is a first attempt at measuring exactly this "steering intelligence." An AI CEO for a fictional software company In CEO-Bench, an agent runs a made-up subscription software company called NovaMind. It starts with zero customers and one million dollars in the bank. Performance is measured by remaining cash at the end. If the balance drops below zero even once, the company is bankrupt and the simulation ends. The agent controls the company through a Python API with 34 tools and a database of 19 tables. Instead of just issuing individual commands, it writes its own code, queries the database with SQL, and builds custom workflows from the results. That puts it in front of the same challenges a human CEO would face, the researchers say. There's a lot to decide: pricing and tiers, ad spend across channels, product quality and R&D, infrastructure capacity and customer support, plus multi-round negotiations with enterprise clients. On top of that, there's a simulated social network where the agent can read complaints, competitor news, and economic trends and post itself. Delayed feedback and hidden variables make the test hard What makes the task hard is time and uncertainty. Decisions play out on realistic business timelines: revenue only arrives at billing dates, R&D projects take days to weeks, and mistakes often don't show up until later through churn or damaged reputation. Costs hit right away. The agent has to spend money whose payoff might not show up for weeks. Much of the company's state stays hidden. The agent can't directly see customer satisfaction, willingness to pay, or minimum quality expectations. It has to piece these together from noisy signals like cancellations, support tickets, or reactions on the social network. The simulation models 26 customer segments and individual customers, each with their own budgets, price sensitivities, and expectations. The world keeps changing, too. Competitors periodically raise customer quality expectations, preferences shift over time, and a simulated business cycle affects demand and willingness to pay, so the agent has to keep adjusting. The researchers deliberately chose fixed, transparent rules rather than a language model as referee. They wanted to avoid a weakness they see in Vending-Bench , a test with a simulated vending machine : there, an AI-simulated supplier can reward an agent for unrealistic verbal promises. Most models go bankrupt Of fourteen tested models, most fail the task. Nearly all can generate valid commands and database queries, but none can maintain a coherent strategy over time. Many go bankrupt before the simulation ends. Only three models finish their best run above the starting capital of one million dollars: Claude Fable 5 at $47.15 million, Claude Opus 4.8 at $27.8 million, and GPT-5.5 at $21.3 million. Claude Fable 5 is the only model that lands above starting capital in more than one run. There's a caveat, though. One Fable 5 run aborted because the model refused to continue, and in the other two, some requests fell back to Opus 4.8. GPT-5.5 went bankrupt in two of its three runs. The most telling comparison is with a simple rule-based heuristic that never calls a language model at all. It sets fixed prices, quotas, and tiers, focuses advertising and targeted development on a small set of customer segments, and adjusts capacity based on recent usage. This heuristic reaches $15.76 million, beating every model except Fable 5, Opus 4.8, and GPT-5.5. The researchers also roughly estimate the upper bound of achievable final cash at around $2.2 billion. Even the best agents fall far short. The test is nowhere near maxed out, the authors say. Exploration beats caution Analyzing the decision trajectories reveals clear behavioral differences. GPT-5.5 and Claude Opus 4.8 keep trying new strategies as conditions change, whether that means ramping up customer acquisition, adjusting tiers, or shifting support and R&D budgets. Claude Opus 4.7, by contrast, mostly responds to setbacks by cutting costs and preserving cash. This passive approach lets the model survive to the end but prevents it from turning a profit. Interestingly, Opus 4.8 and GPT-5.5 reach similar final results through very different paths: Opus 4.8 acquires more customers early on but drops to zero customers mid-simulation, while GPT-5.5 holds its customer base throughout. Both write surprisingly sophisticated code. Opus 4.8 builds its own internal simulation that models customer cohorts to predict future cash flow. GPT-5.5 digs through negotiation history in the database to uncover hidden customer preferences. The researchers measure four capabilities that correlate with success: uncovering hidden information, like which ad channel works best for a given customer segment, predicting the future, measured by error in four-week cash forecasts, adapting quickly to change, measured by how fast a model notices a competitor's move, and planning ahead, measured partly by how often if-then scenarios appear in the agent's notes. On all four points, Opus 4.8 and GPT-5.5 score above the average of the other models. The tool environment matters too Another finding concerns the software environment agents use to act. The researchers also tested Claude Opus 4.7 with Claude Code and GPT-5.5 with Codex, two popular coding assistants. In both cases, the agents acted far less often and performed worse. The researchers suspect the system prompts in these tools, which are tuned for software development, are the cause. Shortening the time horizon doesn't solve the problem either. When the simulation is compressed to 50 days, only GPT-5.5 manages to finish with a profit. Most models, the researchers conclude, remain weak at coordinating decisions even toward a short-term goal. The authors acknowledge limits in their setup. The product is represented by a single quality score because they found no reliable way to evaluate qualitative product changes. Compliance, security, and fundraising are left out to keep each run economically feasible. Still, CEO-Bench exposes a gap between the local tool competence of today's models and the ability to connect actions over long time horizons into a coherent strategy, they say. AI News Witho