메뉴
BL
TechCrunch AI • 6일 전

안드리센 호로위츠 투자받은 Vals, AI 벤치마킹 표준 노린다

IMP
7/10
핵심 요약

AI 벤치마킹 스타트업 Vals가 안드리센 호로위츠(Andreessen Horowitz)가 주도한 시리즈 A에서 4,000만 달러를 유치했습니다. Vals는 공개된 시험 문제에 모델을 학습시켜 성적을 부풀리는 기존 벤치마크의 한계를 지적하며, 문제를 비공개로 유지하고 법률·금융·코딩 등 실제 산업별 과제 수행 능력을 평가하는 방식으로 차별화합니다. 매출이 작년 대비 8배 성장하는 등 빠르게 확장 중이며, AI 모델 도입 결정의 핵심 판단 기준으로 자리 잡고 있습니다.

번역된 본문

벤치마킹은 AI 기업들이 자사 모델의 역량을 검증하고, 지표가 자신에게 유리하게 흐를 때 경쟁사와 차별화하며 우위를 홍보하는 산업 표준이 되었다. 다시 말해, 좋은 벤치마크 결과는 거의 언제나 좋은 홍보로 이어진다. 안타깝게도 기업들은 기존 벤치마킹 시스템을 이해하는 방법을 찾아냈는데, 이러한 시스템 상당수는 오래되었으며 최신 모델의 역량을 측정하도록 설계되지 않았다.

2024년에 설립된 스타트업 Vals는 바로 이 불완전한 시스템을 고치는 것을 사명으로 삼고 있다고 말한다. 2년도 채 되지 않는 기간 동안 이 회사는 기술 산업에서 주목할 만한 존재로 자리 잡았으며, 작년에는 8VC와 블룸버그 베타(Bloomberg Beta)가 주도하는 시드 라운드 투자를 확보했다. 그리고 지난달에는 급성장 기간을 거쳐 안드리센 호로위츠(Andreessen Horowitz)가 주도하는 4,000만 달러 규모의 시리즈 A 투자를 유치했다.

공동 창업자인 라얀 크리슈난(Rayan Krishnan, 25세)은 이전에 팔란티어(Palantir)에서 인턴으로 일했으며, 스탠퍼드대 학부 재학 중 마이크로소프트와 학교의 유명한 인공지능 연구소에서 일했다. 크리슈난은 Vals가 자신이 목격한, 즉 측정 대상 산업의 발전 속도에 벤치마킹이 뒤처지고 있다는 관찰에서 탄생했다고 말한다.

"우리는 매우 유능한 새로운 모델들이 빠르게 시장에 나오는 것을 목격했는데, 학계의 벤치마크는 그 최전선의 발전 속도를 따라가지 못하고 있었습니다"라고 크리슈난은 말한다. AI가 사회의 모든 부분에 통합되면서, 벤치마크는 기업들이 광고하는 바를 모델이 실제로 수행할 수 있는지 검증하는 역할을 해야 한다고 그는 덧붙였다.

지난주, 이 젊은 창업자는 샌프란시스코 폴섬 스트리트에 있는 회사의 2층짜리 사무실을 안내해 주었다. 이 오래된 벽돌 건물은 한 세기 전 대형 맥주 양조장 자리였으며, 지금은 맥주라는 산업 생산물 대신 기술 산업의 미래를 만들어내려는 여러 스타트업의 본거지가 되었다.

"역사적으로 평가는 매우 추상적인 방식으로 지능을 평가하는 데 사용되어 왔다고 생각합니다"라고 크리슈난은 말한다. "예를 들어, 모델이 변호사 시험 같은 시험을 치를 만큼 충분한 정보를 알고 있는가 하는 식이죠."

여기서 Vals는 차별화를 추구한다. 많은 벤치마킹 시스템이 공개된 시험을 제공하는데(이는 기업이 자사 모델을 그 시험에 맞춰 학습시켜 사실상 시험에서 부정행위를 할 수 있게 한다), Vals는 구체적인 시험 자료를 공개하지 않는다. 또한 Vals는 AI 모델의 일반 지식만 측정하는 것이 아니라, 법률·금융·코딩 등 특정 산업과 관련된 복잡한 과제를 완수하는 능력으로 모델을 평가한다.

"우리가 하는 일은 모델의 실제 영향을 살펴보는 것입니다"라고 크리슈난은 말했다. "모든 영역에서 모델이 인간과 동일한 품질의 결과물을 만들어내는 작업을 수행할 수 있는가를 보는 것이죠."

긍정적인 결과뿐 아니라 부정적인 결과도 확인해야 한다고 그는 말한다. "이 모델들이 세상에 무분별하게 퍼진다면 어떤 부정적 결과가 초래될지" 분석하는 것이 목표라고 한다.

Vals가 측정하는 역량의 범위는 계속 확대되고 있다. 전통적인 산업 외에도 이 스타트업은 더 독특한 영역으로 확장하고 있다. "재귀적 자기 개선(recursive self improvement)에 관한 벤치마크를 보유하고 있습니다. 정신 건강, 사이버 보안, 생물 보안, 그리고 모델이 제네바 협약을 어떻게 적용하는지 이해하기 위한 무력 분쟁법 분야에서도 작업하고 있습니다"라고 크리슈난은 말한다.

기업들은 Vals에 비용을 지불하고 자사 모델을 테스트하는데, 이는 다소 생소한 개념일 수 있다. 왜 기업이 자사 모델이 성능이 좋지 않다는 것을 알기 위해 돈을 지불할까? 하지만 효과적인 측정은 기업이 문제를 해결하고 시간이 지나며 개선하는 데 도움이 된다. 크리슈난은 이들의 수익 모델을 학생이 SAT 응시료를 칼리지 보드에 지불하는 것에 비유했다. 이러한 평가는 새로운 AI 모델 도입을 검토하는 기업들의 핵심 의사결정 요소가 되고 있다.

이 스타트업은 최근 현재 매출이 작년의 8배에 달한다고 밝혔다. 직원 수도 늘고 있다. 연초에 8명으로 시작한 Vals는 이미

원문 보기
원문 보기 (영어)
Benchmarking has become the industry norm for how AI companies validate their models' capabilities and, when the metrics swing in their favor, stand out from competitors and advertise their superiority. In other words, good benchmarks pretty much always mean good PR. Unfortunately, companies have also figured out how to outwit legacy benchmarking systems — many of which are older , and not built to measure the capabilities of modern models. Vals, a startup formed in 2024 , says that it is on a mission to fix this very imperfect system. In the span of less than two years, the company has established itself as a notable presence in the tech industry and, last year it managed to secure a seed round led by 8VC and Bloomberg Beta. Then, last month, after a period of rapid growth, it raised $40 million in a series A led by Andreessen Horowitz. Rayan Krishnan, the company's 25-year-old co-founder, previously interned at Palantir, and, as an undergraduate at Stanford, worked for Microsoft and the school's much lauded artificial intelligence lab. Krishnan says Vals was born from his own observations about how benchmarking was falling behind the advances of the industry it was designed to measure. "We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," Krishnan shares. With AI being integrated into every part of society, benchmarks should really exist to verify that models can do what companies advertise they can do, Krishnan said. Last week, the young founder showed me around his company's two-floor office on San Francisco's Folsom Street — an old brick building that, a century ago, served as the site of a large brewery . Instead of an industrial output of beer, the historical structure is now home to a number of different startups looking to ship the future of the tech industry. “Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," Krishnan tells me. "Like, do models know enough information to be able to take a bar exam type test?” Here, Vals seeks to differentiate itself. While many benchmarking systems offer tests that are publicly available (this can allow a company to train its model against those tests, thus arguably cheating on their exam ), Vals doesn't publicly disclose its specific test materials. Instead of measuring an AI model's general knowledge, Vals also evaluates models on their ability to complete complex tasks associated with specific industries like law, finance, and coding. “What we’re doing is actually looking at what are the real impacts of the models,” said Krishnan. “Can they do work that produces a product of the same quality as a human within every domain?” The idea is to check not just for positive outcomes but also for negative ones, he says. The hope is to analyze how, “if these models ran wild in the world, what the negative implications would be." The capabilities that Vals is measuring are growing. In additional to more traditional industries, the startup continues to push into more unique terrain. "We have a benchmark on recursive self improvement. We're doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention," Krishnan shares. Companies pay Vals to test their models, which can be an odd concept to wrap your head around. Why would a company pay to learn its model isn't performing well? But having an effective measurement helps companies troubleshoot and improve over time. Krishnan compares their revenue model to how a student might pay the College Board to take the SAT. In turn, these evaluations are becoming key decision-making factors for companies looking to acquire new AI models. The startup recently revealed that its revenue is currently eight times what it was last year. Its staff is also growing. Vals, which started the year with only eight people, has already tripled to a team of 25. Krishnan said that as the startup grows, the plan is to relocate to a significantly bigger office, as well as to bring on an additional 10 to 15 people. The company also recently launched a program centered around providing model evaluations to federal agencies. Krishnan sees his company's system of benchmarking as the future of how AI companies think about growing their businesses and establishing public trust. "AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they're going to make in AI," he said. Topics AI , AI , Andreessen Horowitz , Startups When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Lucas Ropek Senior Writer, TechCrunch Lucas is a senior writer at TechCrunch, where he covers artificial intelligence, consumer tech, and startups. He previously covered AI and cybersecurity at Gizmodo. You can contact Lucas by emailing lucas.ropek@techcrunch.com. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular A new kind of AI model from a ChatGPT inventor is thrilling developers Tim Fernholz OpenAI caught its models leaving notes to successors to hide bad behavior Rebecca Bellan Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal Rebecca Bellan Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant Salesforce and Nvidia's new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis