메뉴
HN
Hacker News 36일 전

AI 평가 스타트업이 실패하는 이유 (2025)

IMP
8/10
핵심 요약

AI 평가(eval) 전문 스타트업들이 실패하는 핵심적인 이유는 우수 인재의 유출, 제한된 고객층, 그리고 모델 개발사로부터 가해지는 최적화 압력 때문입니다. 평가 인력은 더 큰 경제적 보상과 영향력을 얻을 수 있는 포스트트레이닝(post-training) 등 다른 분야로 빠져나가며, 타겟 고객층이 모호하여 비즈니스 모델을 유지하기 어렵습니다. 결국 독립적인 평가 스타트업은 생존하기 힘들다는 분석입니다.

번역된 본문

AI 평가 스타트업이 실패하는 이유 (2025년 5월 8일)

왜 독립적인 AI 평가(eval) 스타트업은 이렇게 적을까요? 에이전트, 음성, 또는 음성 에이전트와 같은 새로운 AI 트렌드가 나타날 때마다, 개발자들은 수많은 선택지에 직면합니다. 그 중 일부는 최고의 모델을 식별하고 그 지식을 다른 개발자들에게 판매하는 것, 즉 '평가(eval)'를 파는 것이 비즈니스 기회라고 확신합니다. 저는 우리가 이를 생성형 AI라고 부르기 이전부터, 생성형 AI의 모든 물결 속에서 이런 현상을 보아왔습니다. 안전성 평가(safety evals)라는 틈새 시장을 제외하고는 성공한 곳을 보지 못했습니다.

저는 독립적인 평가 스타트업이 왜 도산하는지에 대해 몇 가지 이론을 가지고 있습니다. 첫째, 훌륭한 평가를 설계하고 실행할 수 있는 사람들은 모델 개발 스택의 다른 부분에서 더 많은 돈을 벌고 더 큰 영향력을 행사할 수 있기 때문에 인재가 유출됩니다. 둘째, 평가 스타트업은 고객을 찾기가 매우 어렵습니다. 왜냐하면 고객은 API를 사용해 구축하려는 기술적 역량이 있는 개발자이면서도, 동시에 자체적인 평가를 실행할 만큼 기술력이 뛰어나지는 않은 사람들이어야 하기 때문입니다. 셋째, 평가 스타트업은 일반적인 모델 개선 과정이나 모델 개발사들로부터 가해지는 압력으로 인해 그들이 만든 평가 도구가 무용지물이 되어버리는 엄청난 최적화 압력에 직면합니다.

평가 인재는 다른 곳에서 더 잘 쓰입니다 훌륭한 평가 인재는 스택의 다른 부분으로 이동합니다. 훌륭한 평가를 위해 필요한 기술이 포스트트레이닝(post-training) 및 애플리케이션 개발에도 유용하기 때문입니다. 이 분야들은 더 많은 가치를 창출하고(즉, 더 많은 돈을 벌고), 모델 개발에 더 직접적인 영향을 미칩니다(즉, 더 명예롭고 흥미롭습니다).

예를 들어, 좋은 평가를 구축하려면 인간 피드백 파이프라인을 운영하거나 합성 데이터(synthetic data)를 통해 고품질 데이터를 수집해야 합니다. 고품질 데이터 수집은 포스트트레이닝의 주요 병목 현상입니다. 데이터 포인트당 가치가 같다고 가정할 때, 평가에 사용되는 데이터의 양은 포스트트레이닝을 위해 수집되는 데이터의 양보다 항상 자릿수가 적습니다. 따라서 엄밀히 말해, 평가용 데이터를 수집하여 얻는 가치는 포스트트레이닝용 데이터를 수집하여 얻는 가치에 비해 한계가 있습니다. 게다가, 좋은 포스트트레이닝의 재정적 수익은 잠재적으로 수억에서 수십억 달러에 달할 수 있지만, 평가의 재정적 수익은 귀하의 가장 큰 평가 계약 규모로 제한되며, 이는 그것과는 비교할 수도 없을 정도로 적습니다. 기회비용의 개념을 우연히 이해하고 있는 똑똑한 젊은 연구원들에게는 이러한 역학이 명백하게 다가옵니다. 이를 보여주는 예시로, 에이전트 평가 업무를 하던 에포크 AI(Epoch AI)의 세 명의 연구원이 직장을 그만두고 대신 에이전트를 위한 포스트트레이닝 도구를 구축하는 스타트업을 설립한 사례가 있습니다.

평가 고객이 부족합니다 평가 스타트업이 인재를 유지한다 하더라도 여전히 고객을 찾기는 어렵습니다. 왜냐하면 '모델 API를 기반으로 구축함'과 '모델을 평가할 능력이 없음'이라는 두 개의 원이 교차하는 벤 다이어그램의 영역은 무시해도 될 정도로 작기 때문입니다.

시장 조사 기업인 가트너(Gartner)가 공급업체들을 비교하는 차트를 보면, X축은 환상적이고 Y축은 허구적입니다. 즉, 차트는 차트가 인쇄되어 전달되는 기업 임원들과 기술적 수준이 비슷한 유아들이 해석할 수 있도록 만들어졌습니다. 제가 과장하고 있다고 생각하면 구글에 'Gartner Magic Quadrant AI'를 검색한 다음, 차트 범죄 부서에 신고해 보시길 권합니다. 이와 같은 늪이 AI 평가 스타트업을 빠뜨린 곳입니다.

포스트트레이닝 모델을 다루는 모든 고객은 확실히 그들 스스로 평가를 구축하고 있습니다. 도구 사용 없이, Best of N으로 계산된 AIME 2024에서 10% 개선의 의미와 영향을 이해하는 개발자는 그저 직접 평가를 실행하는 것과 멀지 않은 사람입니다. 만약 그들이 GPT 4o와 GPT 4.1의 차이를 이해하지 못한다면, 그들은 기능이나 분석이 아닌 '솔루션'을 원하는 고객일 것입니다(당연히 ELO 방식에 대한 설명도 원하지 않을 것입니다). 가트너는 클라우드 제공업체와 대규모 계약을 맺는 임원들을 위해 수준을 낮출 수는 있지만, 평가 스타트업은 항상 개발자들에게 판매하려는 경향이 있습니다. 따라서 AI 서비스에 대한 수요가 증가하더라도 평가 스타트업을 위한 시장은 그리 크지 않을 것이라고 저는 회의적으로 생각합니다.

원문 보기
원문 보기 (영어)
Why eval startups fail May 8 th , 2025 Why are there so few independent eval startups? Whenever there's a new AI trend, like agents, or voice, or voice agents, developers are faced with a flurry of options, and a subset of them are convinced that there's a business opportunity in identifying the best models and selling that knowledge to other developers—that is, selling evals. I've seen this in every wave of generative AI, since before we were calling it generative AI. I haven't seen any succeed, outside the safety evals niche. I have a few theories why independent eval startups die. First, people who can design and run good evals can make more money and have more influence in other parts of the model development stack, so talent attrits. Second, eval startups have a hard time finding customers, because clients have to be technical developers who want to build with APIs, but also not technical enough to run their own evals. And third, eval startups face immense optimization pressure that renders their evals useless, both from garden-variety hill climbing and through pressure from model developers. Eval talent is better used elsewhere Good eval talent moves to other parts of the stack because the same skills that are needed for good evals are useful for post-training and for application development, and these areas capture more value, i.e. make more money, and have more direct influence on model development, i.e. are more prestigious and interesting. For example, building a good eval requires collecting high-quality data, whether from operating a human feedback pipeline or through synthetic data. Collecting high-quality data is a major bottleneck for post-training. The amount of data in an eval is always smaller than the amount of data collected for post-training, by orders of magnitude, so in a real sense the value you generate from collecting data for evals is capped compared to the amount of data you generate from collecting data for post-training, assuming the value per datapoint is equal. Additionally, the financial return on a good post-train is potentially very high, up to a few hundred million or billions of dollars, whereas the financial return on an eval is capped at the size of your largest eval contract, which is nowhere close. This dynamic is readily apparent to smart young researchers who incidentally understand the notion of opportunity cost. An illustrative example is provided by three researchers who quit their jobs at Epoch AI evaluating agents to instead start a startup building post-training tools for agents [0] . Not enough eval customers Even if an eval startup retains talent, it still has a hard time finding customers, because the Venn diagram intersection of the two circles "building on model API" and "unable to evaluate models" has negligible area. When you look at charts comparing vendors by Gartner, a market research firm, the X-axes are fantastical and the Y-axes are fictional; in short, the charts are made to be interpreted by toddlers, who have technical caliber comparable to the corporate executives those charts are printed for. If you think I'm exaggerating I encourage you to Google "Gartner Magic Quadrant AI" then report them to the Department of Chart Crimes. This same quagmire ensnares AI eval startups. Any customer that is post-training models is definitely building evals themselves. A developer who understands the meaning and implication of a 10% improvement on AIME 2024, without tool use, computed with best of N, is not far from just running that eval themselves. If they don't understand the difference between GPT 4o and GPT 4.1 they're the kind of customer that wants solutions, not features, and certainly not an explanation of ELO. Gartner can dumb down for execs, who are deciding on large contracts with cloud providers, but eval startups seem always to want to sell to developers. Thus I am skeptical the market for eval startups is very large, even as the demand for AI services grows. Big labs Goodhart evals An eval startup that overcomes these two hurdles now has to face down the big labs themselves, who are highly incentivized to climb the public eval and apply pressure and tricks to improve their numbers. Once benchmarks are targeted models can improve rapidly, whether that's from benign adjustments like including more diverse data to outright training on test data, which Meta did for Llama 1 [1] and is rumored to have done for Llama 4 [2] . So eval startups have to be wary about a potentially adversarial relationship with big labs, who don't want to lose their own customers and will play their unfair advantages. Other kinds of tricks big labs employ include asking employees to vote for their own models on public leaderboards, poaching employees from eval startups, dangling free compute in return for better results, asking for private insights about model performance; the list of shenanigans is long. A principled team can resist these gambits, but the pallor of suspicion is hard to dispel. For two years every researcher has asked themselves — why is every new model release always at the top of the LMSys Chatbot Arena leaderboard? A new report led by Cohere suggests the cause is systematic gaming, claiming that Meta tested twenty-seven unique model variants before releasing Llama 4 [3] . Meta, by the way, advertised that its tiny Llama 4 Maverick model outperformed GPT-4.5, before revealing that the result was achieved with a version optimized specifically for Chatbot Arena, and not the released version, which ranked abysmally. Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. And all eval startups have to sell are measures. Safety evals are an exception I believe eval startups can work when they're targeting safety benchmarks specifically. Researchers who want to work on safety evals tend to be ideologically opposed to working on capabilities, which means they don't migrate to post-training or applications due to monetary incentives. (This is how the internal safety eval divisions of the big labs retain talent.) They can provide services to technical clients who are capable of replicating those services, because it's specifically important for safety evals that those services are provided by an external vendor and not only done internally. They can also sell to policymakers, or have business assured by regulation if proposals for external model audits are passed. Safety eval startups would still be vulnerable to Goodharting, but if labs are Goodharting safety evals, there are other things to be worried about. So safety evals have particular characteristics that make them more amenable than other evals. I've presented three reasons why it's hard for eval startups to survive. The most pernicious of these is the first, which is that there are better opportunities available for any company or engineer who is good at evals, but the other two pose serious headwinds as well. I have nothing against eval startups, and I am rooting for them, but I am not counting on them. ❖ ❖ ❖ Additional comments The above is for application-focused evals, i.e. evals for developers who want to build on top of model APIs. There are also startups that want to sell research evals to big labs. These will fail, because the primary point of research evals is to set research directions, and big labs will never outsource setting their research agenda. Also, outsourcing research evals adds a ton of latency to model iteration, and velocity is everything. Added: May 21 st , 2025 . There's a difference between selling evals and selling evals tooling. In the same way that selling human labels is different from selling tooling to collect human labels - one is an ops business with ops margins, the other is a SaaS business with SaaS margins - selling evals and selling evals tooling have two very different economics. LM Arena, the organization behind Chatbot Arena, today announced a $100M seed round [4] . That's a very