메뉴
BL
TechCrunch AI • 53일 전

디자인아레나(DesignArena), AI 모델 향상 위한 790만 달러 유치

IMP
8/10
핵심 요약

디자인아레나(DesignArena)를 개발한 스타트업 인텔리전스(Intelligence)는 AI 모델의 디자인 감각 및 품질을 평가하기 위해 사용자 피드백을 수집하는 플랫폼을 구축했습니다. 일반 사용자에게는 챗GPT(ChatGPT)처럼 프롬프트를 입력하고 A/B 테스트를 통해 최적의 결과물을 뽑아주는 서비스를 제공하며, AI 기업들은 사용자의 선택 데이터를 유료로 구매하여 모델을 고도화하는 방식입니다. 이미 연간 반복 수익(ARR) 6,000만 달러를 달성하며 수요를 입증했지만, 크라우드소싱 기반 평가 모델의 지속 가능성은 여전히 업계의 주요 과제로 남아있습니다.

번역된 본문

공동 창업자인 그레이스 리(Grace Li)의 말에 따르면, 그녀의 회사는 2025년 졸업 몇 주 전에 시작되었습니다. 소수의 대학 동기들과 함께 AI 게임 엔진을 작동시키려고 시도했습니다. AI 모델이 작동하는 게임을 만들어낼 수는 있었지만, 그 어떤 게임도 재미가 없었습니다. 이는 '어떻게 게임이 재미있는지 알 수 있을까?'라는 흥미로운 의문을 남겼습니다. 그들은 인간의 판단을 대체할 수 있는 것은 없다고 결론 내렸고, 곧 대규모로 진실된 인간의 피드백을 얻을 방법을 놓고 브레인스토밍을 시작했습니다. 그 결과 탄생한 것이 전 세계 530만 명이 사용하는 AI 도구인 디자인아레나(DesignArena)입니다.

결과적으로 보면, 확장 가능한 사용자 피드백을 찾고 있는 AI 기업들이 굉장히 많았고 그중 다수가 이를 위해 기꺼이 비용을 지불할 의사가 있었습니다. 리는 "이는 수많은 모델이 디자인 분야에서 발전하기 위해 놓쳤던 병목 현상이었습니다. 약 일주일 후, 우리는 최첨단 연구소(frontier lab)와 첫 번째 주요 계약을 성사시켰고, 이후는 역사 속으로 들어갔습니다."라고 말했습니다. 월요일, 디자인아레나를 운영하는 인텔리전스(Intelligence)라는 회사는 인덱스 벤처스(Index Ventures)가 투자를 주도하고 컨빅션(Conviction, 사라 구오 및 마이크 버널), A*, 발키리(Valkyrie) 등이 참여한 790만 달러(약 100억 원)의 시드 투자 유치를 발표했습니다.

기업 고객이 아닌 일반 사용자의 경우, 디자인아레나를 사용하는 것은 정교한 모델 라우터를 사용하는 것과 매우 유사합니다. 프롬프트를 입력할 수 있는 챗GPT(ChatGPT) 스타일의 창이 있으며, 웹사이트, 이미지, 그리고 10여 종의 다른 시각적 형식을 선택할 수 있는 드롭다운 메뉴가 따로 제공됩니다. 요청 사항, 형식 및 스타일을 입력하면, 몇 개의 결과물을 최고에서 최악의 순으로 랭킹을 매릴 때까지 일련의 'A vs. B' 선택지가 제시됩니다. 유용한 서비스이기는 하지만, 이 플랫폼의 진정한 가치는 기업용(B2B) 측면에서 발휘됩니다. 참여하는 AI 모델들은 이 플랫폼을 미디어 생성 모델을 위한 끝없고 즉각적인 피드백의 출처로 삼을 수 있기 때문입니다.

사용자들은 자신들이 순위를 매기는 모델이 무엇인지 무관심하는 경향이 있습니다. 리가 말하듯, 그들은 그저 얻을 수 있는 최고의 결과물을 원할 뿐이므로, 사용자가 매긴 순위는 대중이 진정으로 원하는 것이 무엇인지 알려주는 핵심적인 입력값이 됩니다. 리는 최첨단 연구소들에게 이는 비용을 지불할 만한 가치가 있는 서비스라고 덧붙였습니다. 또한 이 사이트는 현재 6,000만 달러의 연간 반복 수익(ARR)을 기록하며 있어, AI 산업에서 인간 주도 평가 데이터의 핵심 공급원으로서의 입지를 확고히 했습니다. 핵심적인 점은 사용자가 결과물을 얻으려면 반드시 로그인해야 한다는 것입니다. 따라서 인텔리전스(Intelligence)는 취향이 대륙별로, 그리고 시간이 지남에 따라 어떻게 변하는지 추적할 수도 있습니다. (리는 아시아의 웹 대시보드는 더 극대주의적(maximalist)인 디자인 스타일을 띠는 경향이 있다고 설명합니다.)

이러한 정성적 지표는 자동화된 벤치마크를 보완하는 중요한 역할을 합니다. 자동화 벤치마크는 훨씬 더 큰 규모로 작동할 수 있지만, 지난주 허깅 페이스(Hugging Face)의 침해 사고에서 극적으로 입증되었듯 조작되거나 다른 방식으로 악용될 가능성이 큽니다. 그렇다고 해서 크라우드소싱 기반의 인간 피드백이 무조건 승리하는 시장이 된다는 의미는 아닙니다. 올해 초, Yupp이라는 스타트업은 에이식즈(a16z) 크립토의 크리스 딕슨(Chris Dixon)으로부터 3,300만 달러를 투자받았음에도 불구하고 출시 1년 만에 문을 닫았습니다. 그들 역시 일부 최첨단 모델들을 고객으로 확보했고 130만 명 이상의 사용자를 보유했다고 밝혔지만, 지속 가능한 장기적인 비즈니스를 구축하지는 못했습니다.

그럼에도 불구하고, 인간 평가를 기반으로 하는 다른 스타트업들은 번창하고 있는 것으로 보입니다. 텍스트 기반 응답에 유사한 접근 방식을 취하고 있는 엘엠 아레나(LM Arena)는 유료 제품을 공식 출시한 지 불과 4개월 만인 1월에 시리즈 A에서 1억 5,000만 달러를 유치했습니다.

(주제: AI) 기사 내 링크를 통해 구매하시면 소정의 수수료를 받을 수 있습니다. 이는 본 매체의 편집 독립성에는 영향을 미치지 않습니다.

러셀 브랜덤 / AI 에디터 러셀 브랜덤은 2012년부터 플랫폼 정책 및 신흥 기술에 중점을 두고 기술 산업을 취재해 왔습니다. 그는 이전에 더 버지(The Verge)와 레스트 오브 월드(Rest of World)에서 일했으며, 와이어드(Wired), 더 올(The Awl), MIT 테크놀로지 리뷰(MIT's Technology Review)에 글을 기고했습니다. russell.brandom@techcrunch.com 또는 Signal(412-401-5489)로 연락할 수 있습니다.

10월 13일 ~ 15일 / 샌프란시스코 더 빠르게 스케일링하세요. 포트폴리오를 성장시키세요. 실용적인 전문 지식을 얻으세요. 목표가 무엇이든, 디스럽트(Disrupt) 행사가 여러분에게 힘을 실어줄 것입니다. 오늘 최대 $330를 절약하세요! 지금 등록하세요

가장 인기 있는 기사: 유튜버 행크 그린(Hank Green), 자신의 AI 사용이 '건강하지 않다'고 말하다 - 앤서니 하(Anthony Ha) 왓츠앱(WhatsApp), 대규모 발신자 메시지를 위한 새 폴더 테스트 중

원문 보기
원문 보기 (영어)
As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun? There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale. The result became DesignArena , an AI tool now used by 5.3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it. “It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.” On Monday, the company behind DesignArena — dubbed Intelligence — announced a $7.9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others. For non-enterprise users, using DesignArena is a lot like using a sophisticated model router. There's a Chat-GPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format and style, you'll be presented with a series of "A vs. B" choices until you've ranked the handful of outputs from best to worst. It's a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they're ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want. For frontier labs, that's a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry. Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time. (Li notes that web dashboards in Asia tend to have a more maximalist design style.) These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated, as the Hugging Face breach demonstrated in dramatic fashion last week. That's not to say that crowdsourced human feedback will be an automatic winning market. Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33M from a16z crypto’s Chris Dixon . It too nabbed some frontier models as customers and had, it said, over 1.3 million users, but still couldn't build a sustainable long-term business. Even so, other startups based on human evaluation seem to be thriving. LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January , just four months after formally launching its paid product. Topics AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 toda y! REGISTER NOW Most Popular YouTuber Hank Green says his AI usage is ‘not healthy’ Anthony Ha WhatsApp is testing a new folder for messages from large businesses Ivan Mehta Spotify adds a running mode to its app Ivan Mehta Claude Opus 5 became downright ruthless when tasked with running a vending machine Julie Bort DoorDash is building its own drone delivery business Kirsten Korosec Sam Altman is ready to decelerate Tim Fernholz Data centers may face temporary power cuts to prevent blackouts on largest US grid Tim De Chant