메뉴
BL
TechCrunch AI • 1일 전

일레븐랩스 CEO 인터뷰, 기업가치 220억 달러 평가

IMP
7/10
핵심 요약

AI 음성 합성 기업 일레븐랩스(ElevenLabs)가 연 반복 매출(ARR) 6억 달러를 기록 중이며, 투자자들로부터 220억 달러의 기업가치를 평가받고 있다고 전해졌다. 클라르나, 도이체 텔레콤, 시스코 등 대형 고객은 물론 여러 정부 고객을 확보했으며, 창립 4년 만에 음성 AI 시장을 선도하고 있다. CEO 마티 스타니셰프스키는 대화형 AI 튜링 테스트 통과가 다음 목표라고 밝혔다.

번역된 본문

일레븐랩스(ElevenLabs)는 AI의 음성 레이어, 즉 텍스트를 사람처럼 들리는 음성으로 변환하는 모델을 구축하는 회사다. 대부분의 사람들은 고객 서비스와 통화할 때, 종종 인식하지 못한 채 이 기술을 접한다. 예를 들어 클라르나(Klarna)는 미국 고객 3,500만 명을 대상으로 하는 1차 전화 지원을 일레븐랩스로 운영한다. 도이체 텔레콤, 시스코, 어도비, 그리고 점점 늘어나는 정부 기관들도 마찬가지다. 일레븐랩스는 오디오북, 더빙, 음악 등에 자사 플랫폼을 사용하는 크리에이터들에게도 판매한다.

이 회사는 이 분야에서 유일한 존재가 아니다. 실제로 점점 더 고객사와 경쟁하는 상황에 부딪히고 있다. 예컨대 대화형 AI 플랫폼인 데카곤(Decagon)은 일레븐랩스로 음성 제품을 학습시켰고, 지금은 일레븐랩스와 경쟁하고 있다. 하지만 투자자들은 크게 우려하지 않는 것으로 보인다. 연 반복 매출(ARR)이 6억 달러 수준이라고 밝힌 이 회사는 창립 4년 만에 투자자들로부터 220억 달러의 기업가치를 평가받고 있다고 전해진다.

이를 더 자세히 알아보기 위해 나는 토론토의 지역 창업 컨퍼런스인 'Nrth'(구 Elevate)에서 일레븐랩스의 공동 창립자이자 CEO인 마티 스타니셰프스키를 인터뷰했다. 우리는 짧은 시간 동안 다양한 주제를 다뤘다. 기업이 고객과 대화하는 상대가 AI임을 알려야 하는지(그는 알려야 한다고 생각한다), 회사의 매출 총이익률에 대해 이야기할 수 있는지 등이다. 놀랍지 않게도 스타니셰프스키는 구체적으로 논의할 수 없다고 말했지만, 시장 점유율 확대를 의미한다면 이익률이 더 압박받는 것을 개의치 않는다고 분명히 했다. 다음은 축약되고 가볍게 편집된 인터뷰 내용이다. 전체 대화는 여기서 확인할 수 있다.

작년 테크크런치 디즈럽트(TechCrunch Disrupt)에서 오디오 모델이 몇 년 안에 커모디티화(일반화)될 것이라고 말씀하셨는데, 지금은 그 예측을 어떻게 평가하시나요?

아직 해야 할 일이 많고, 모델 수준에서 달성할 수 있는 품질 차이는 여전히 큽니다. 장기적으로 보면, 아마 3~5년 후에는 그 차이가 더 작아질 것입니다. 우리가 하고 싶고, 가장 먼저 하고 싶은 것은 대화형 AI의 튜링 테스트를 통과하는 것입니다. 지능을 결합해야 하지만 감성 지능도 필요합니다. 상대방의 감정을 이해하고, 속도를 늦추거나 목소리를 높일 수 있어야 합니다. 이는 아직 실현되지 않았습니다.

현재 비즈니스에서 기업(엔터프라이즈) 비중은 얼마나 됩니까?

현재 ARR이 6억 달러입니다. 그중 55% 이상이 전통적인 기업 고객이며, 나머지 45% 중 상당 부분이 중소기업, 개발자, 빌더, 크리에이터입니다.

AI 업계의 다른 모두와 마찬가지로, 점점 고객사와 경쟁하고 계십니다. 데카곤은 귀사의 기술로 음성 제품을 학습시켰고, 지금은 자체 모델로 쿼리를 처리합니다.

경계가 점점 흐려지고 있습니다. 모델 기업, 플랫폼 기업, 애플리케이션 기업을 생각해 보면, 과거에는 어디서 시작하고 끝나는지 매우 명확하게 구분됐습니다. 오늘날 그 경계는 훨씬 흐릿합니다. 앤스로픽(Anthropic)의 경우, 모델 기업이었던 것이 확실히 플랫폼이 되었고, 점점 더 폭넓은 애플리케이션 집합이 되고 있습니다. 이런 흐름은 계속될 것이라고 생각합니다.

고객들은 일레븐랩스에서 여러 옵션 중에서 '추론 레이어(reasoning layer)'를 선택할 수 있습니다. 프론티어 연구소 모델 대비 오픈 웨이트(공개 가중치) 모델의 사용 추이는 어떻게 보이나요?

이진법적 선택은 아닙니다. 고객 경험 분야에서 전화가 단순히 정보 제공 목적이고 아무 작업도 실행하지 않는다면, 지식 베이스가 좋은 경험을 정의하기 때문에 많은 오픈소스 모델을 사용할 수 있습니다. 하지만 금융 서비스라면 인증이 필요하고, 거래 정보나 환불을 원할 수 있습니다. 오류의 여지가 없습니다. 여기서는 프론티어 모델이 여전히 선도할 것입니다.

그런 오픈 웨이트 모델 중 일부는 중국산입니다. 미국 정부가 고객이고, 유럽 정부들도 고객입니다. 그 대화는 어떻게 진행됩니까?

각기 다릅니다. 각 배포 건에서 우리가 배포하는 모델과 음성은 사례에 따라 달라집니다. 폴란드 정부나 브라질 정부와 협력한다면, 그들만의 요구사항이 있습니다. 오픈 웨이트 모델일 수도, 클로즈드 소스 모델일 수도, 그들 자신의 모델일 수도 있습니다.

원문 보기
원문 보기 (영어)
ElevenLabs builds the voice layer of AI, the models that turn text into speech that sounds human. Most people encounter it when they're talking with customer service — often without realizing it. Klarna runs first-line phone support for 35 million U.S. customers on it, for example. So do Deutsche Telekom, Cisco, Adobe, and a growing list of governments. ElevenLabs also sells to creators, who use its platform for audiobooks, dubbing, and music. It's not alone in what it does. In fact, it's increasingly bumping into customers, including Decagon, a conversational AI platform that trained its voice product on ElevenLabs and now competes with it. But its investors don't seem too concerned. The company, which says it's pacing at $600 million in annual recurring revenue, is now reportedly valued at $22 billion by its backers, despite being just four years old. To understand more, I interviewed ElevenLabs co-founder and CEO, Mati Staniszewski, at Nrth in Toronto, a local entrepreneurship conference formerly known as Elevate. We covered a range of topics in a short time, including whether businesses should tell customers when they're talking to an AI (he thinks they should), and whether he could discuss the company's gross margins. Unsurprisingly, Staniszewski said he couldn't discuss them in any detail, but he was clear that he doesn't mind them getting squeezed even further if it means expanding the company's market share. Our conversation follows, condensed and lightly edited. You can check out the full conversation here . You joined us at TechCrunch Disrupt last year, where you said audio models would be commoditized within a couple of years . How would you rate that prediction now? There is still a lot of work to be done, and the quality delta you can achieve just on the model level is still significant. If we think longer term, probably three, five years from now, those differences will be smaller. What we'd love to do, and be the first ones to do, is pass the Turing test for conversational AI. You need to combine intelligence, but you also need emotional intelligence. You need to understand the emotions of the other side, to be able to slow down or speak up. That hasn't yet been done. What percentage of the business is enterprise now? We are now $600 million in ARR. Fifty-five-percent plus is classic enterprise, and a big percentage of the [remaining] 45% are small and medium businesses, developers, builders, creators. Like everyone else in AI, you're increasingly competing with your customers. Decagon trained its voice product on you and now runs queries through its own models. The lines become more blurry. As we think about model companies, platform companies, application companies, in the past you'd have very clear splits where one starts and ends. Today that line is much more blurry. In Anthropic's case, what was a model company is definitely a platform and increasingly a wide set of applications. I think this will continue. Your customers can choose the "reasoning layer" from a menu of options at ElevenLabs. What are you seeing in terms of the use of frontier lab models versus open-weight? It's less of a binary choice. In customer experience, if you're calling in and it's just informational, you're not executing any actions — you can use a lot of the open-source models because your knowledge base defines what a good experience is. But if it's financial services, you want to be authenticated, you want information about a transaction, maybe a refund. There's no room for error. Here, frontier models will still lead. Some of those open-weight models are Chinese. The U.S. government is a customer. European governments are customers. What are those conversations like? Different. In each deployment, the models and the voices we deploy will depend on the case. If we work with the Polish government or the Brazilian government, they have their own set of requirements. It can be an open-weight model, a closed-source model, their own fine-tuned model. [In Poland,] it's a healthcare case. You have patients booking appointments across the public health system, and 18% never show up. The deployment is agents that call and remind you. They had a set of models optimized on their knowledge, and we integrate while keeping data residency. Should businesses disclose when someone is talking to an agent rather than a human? I think there should be disclosure at this time. Currently, people aren't used to it, and the common pattern is you don't want to feel cheated on that call. But in five years, when everybody has their own agent working on their behalf, you'll be calling in and expecting an agent. Then I think we'll shift as a society. There are good ways of doing it — if there's a 30-minute wait for a human, offer the customer a choice. In almost all cases, they choose the agent and then they're surprised by how good the experience is. What are your gross margins, given what you're paying for models and inference? I'm going to give a vague answer. Given we have that research element, we're able to fine-tune and constrain models in extremely smart ways. But if we can pass on any savings to the customer, we do that. The biggest thing is still proving the value and being there with the customer. So if we can invest and prove that value, we don't mind the margins going lower to actually benefit together as the value gets created in the next five years. You have millions of hours of customer service calls. Do you train on them? How much of your training data is synthetic? In certain companies, we created the models together. They wanted a specific model for their use case. Otherwise, the big part of the training hasn't been so much the volume of data, it was annotating the data. We have thousands of people internally on a contracting basis helping us annotate not only what was said, but when people were speaking, how they said things, what emotions were used. We had to bring voice coaches in to be able to detect accents accurately. It's been reported you're looking at 2028 for an IPO. Can you confirm that? We'd love to create a company that stands the test of time. We are preparing the foundation to be able to do it in the next years. But whether we do it will depend on the time and place. ‘Years' is very vague. [Laughs.] Backstage: Where do you land on whether the frontier labs should slow down? Everybody is aligned to work together on finding a way to pace. Whether they should be public about it, and how much of the media conversation or regulation it should involve, that's another topic. But yes, we should all take the right precautions as we deploy the technology. We don't train the text models and the intelligence side of models, which is the core key of the debate. Could ElevenLabs be exposed the way Hugging Face was? We're a step further, because we don't deploy self-replicating or recurrent parts of the intelligence of agents. Our technology doesn't allow you to let agents create more agents. Every customer goes through KYC. Cybersecurity risk is definitely a risk for the wider world, but we have a good set of precautions in place. Topics AI , ElevenLabs , TC When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Connie Loizos Editor in Chief & General Manager Loizos has been reporting on Silicon Valley since the late ’90s, when she joined the original Red Herring magazine. Previously the Silicon Valley Editor of TechCrunch, she was named Editor in Chief and General Manager of TechCrunch in September 2023. She’s also the founder of StrictlyVC, a daily e-newsletter and lecture series acquired by Yahoo in August 2023 and now operated as a sub brand of TechCrunch. You can contact or verify outreach from Connie by emailing connie@strictlyvc.com or connie@techcrunch.com , or via encrypted message at ConnieLoizos.53 on Signal. View Bio October 13 - 15 San Francisco Your next big connectio