메뉴
BL
TechCrunch AI 18시간 전

피쉬 오디오, AI 음성 모델 고도화 위해 5천만 달러 시드 유치

IMP
7/10
핵심 요약

AI 음성 생성 스타트업 Fish Audio(피쉬 오디오)가 크리에이터 및 기업을 위한 고도화된 AI 음성 모델 개발을 위해 5천만 달러(약 670억 원)의 시드 투자를 유치했습니다. 이 회사는 15,000개 이상의 자연어 제어 라이브러리를 바탕으로 연간 2,100만 달러의 매출을 올리고 있으며, 최근 음성 무단 도용 문제를 해결하기 위해 초고속 자동 삭제 시스템을 도입하며 크리에이터 권리 보호에 나섰다는 점에서 주목받습니다.

번역된 본문

AI 생성 음성 모델 시장은 거대합니다. 크리에이티드 사용 사례는 AI 음성 모델이 더 풍부한 표현력을 갖추기를 요구하는 반면, 고객 지원 및 영업 자동화를 원하는 기업들은 더 정밀하게 제어할 수 있는 모델을 필요로 합니다. 팔로알토에 본사를 둔 Fish Audio는 15,000개 이상의 자연어 제어 라이브러리를 통해 이 모든 사용 사례를 충족하고자 합니다. 작년에 출시된 이 스타트업은 오늘날 800만 명 이상의 사용자가 오픈소스 또는 호스팅 버전의 모델을 사용하고 있으며, 현재 2,100만 달러의 연간 반복 수익(ARR)을 창출하고 있습니다.

이러한 성장세를 이어가기 위해, 이 스타트업은 화요일 Coreline Ventures와 Capital Today가 주도한 시드 라운드에서 5천만 달러를 유치했다고 밝혔습니다. 이번 투자에는 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, 그리고 HF0도 참여했습니다.

Fish Audio는 전 엔비디아(NVIDIA) 연구원인 샤자 리아오(Shijia Liao)의 소규모 프로젝트로 시작되었습니다. 그는 시장에서 구할 수 있는 표현력 없는 합성 음성에 불만을 품고, 단일 GPU로 음성 생성 모델을 학습시킨 뒤 이를 오픈소스로 공개했습니다. 현재 GitHub의 Fish Speech 저장소는 31,000개 이상의 스타를 받았으며, 인디 개발자, 비디오 게임 디자이너 및 크리에이터들이 사용하고 있습니다.

이 회사는 지난 1년 동안 4개의 음성 생성 모델과 1개의 음성-텍스트 변환(STT) 모델 등 총 5개의 모델을 출시했습니다. 음성 생성 모델 중 3개는 오픈소스로 공개했지만, 최신 S2.1 Pro 모델은 유료 API를 통해서만 사용할 수 있습니다. Fish Audio는 크리에이터와 팀에 적합한 월간 유료 요금제를 제공하여 일정 시간의 음성 생성 및 음성 클로닝 기능을 사용할 수 있게 합니다. 또한 기업용 API 및 플랫폼 버전을 제공하며, HeyGen, Sanas, Plaud 등의 조직이 이미 이를 사용하고 있다고 밝혔습니다.

Fish Audio의 공동 창립자이자 CEO인 리사 카오(Rissa Cao)는 "모든 기업은 서로 다른 사용 사례와 선호도를 가지고 있습니다. 예를 들어, AI 아바타에 우리 음성을 사용하는 HeyGen과 같은 회사는 음성에서 현실감을 원하지만, 게임 스튜디오는 캐릭터를 위한 풍부한 표현력을 원합니다. 또한 LiveKit과 같은 음성 에이전트 기업은 통화에 적합할 만큼 표현력이 풍부하면서도 자연스럽고 저지연(low-latency)인 음성을 원합니다"라고 말했습니다.

이 스타트업이 음성 라이브러리를 구축한 한 가지 방법은 사용자에게 모델 학습을 위해 자신의 음성을 제출하도록 요청하고, 해당 음성이 사용될 경우 보상을 지급하는 것이었습니다. 하지만 이는 몇 달 전 문제를 일으켰습니다. 일부 크리에이터들이 자신의 동의 없이 음성이 Fish Audio에 업로드되었다고 주장했기 때문입니다. 이 스타트업은 이러한 우려를 해결하기 위해 DMCA 콘텐츠 삭제 프로세스를 마련해 두었지만, 삭제 자체에 오랜 시간이 걸렸습니다. 카오 CEO는 현재 회사가 삭제 프로세스를 자동화했다고 밝혔습니다. 그녀에 따르면, 크리에이터는 짧은 음성 샘플이나 계약서를 제출하여 업로드된 음성이 자신의 것임을 증명하기만 하면 3분 이내에 플랫폼에서 자신의 음성이 삭제됩니다.

그럼에도 불구하고, 이것이 누군가가 아티스트의 허락 없이 그들의 음성을 업로드하는 것을 완벽하게 막지는 못합니다. 아티스트가 사실을 알게 되어 삭제를 요청하기 전까지는 누군가의 음성이 계속해서 플랫폼에서 사용될 수 있습니다. Coreline Ventures의 파트너인 오스크 혼다(Oskue Honda)는 커뮤니티 주도 모델은 크리에이터가 플랫폼을 신뢰할 때만 작동한다고 말했습니다. "크리에이터가 플랫폼을 신뢰할 때만 커뮤니티 중심 접근 방식이 지속 가능한 경쟁 우위가 될 수 있습니다. 이는 동의, 투명성, 출처 표기가 사후 처리가 아닌 제품 자체에 내장되어야 함을 의미합니다. 저는 산업이 인증된 음성 소유권, 명확한 라이선스 조건, 쉬운 신고 및 삭제 프로세스를 지향해야 하며, 궁극적으로는 크리에이터의 음성이 라이선스되거나 상업적으로 사용될 때 재정적 혜택을 얻을 수 있는 수익 공유 모델로 나아가야 한다고 생각합니다."

카오 CEO는 스타트업이 크리에이터를 위한 요금제와 함께 오픈소스 프로젝트로만 제품을 제공하던 시절에는 효율적으로 운영되었고 자금이 필요하지 않았다고 말했습니다. 하지만 회사는 더 발전된 모델을 개발하고 기업의 요구를 수용하고자 했고, 이에 따라 투자자들의 관심이 높아졌습니다.

원문 보기
원문 보기 (영어)
The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable. Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million. To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators. The company has launched five models in the last year: four speech generation models and one speech-to-text model. It has open-sourced three of its speech generation models, but its latest S2.1 Pro model is available only through its paid API. Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation, plus voice cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen, Sanas and Plaud are already using it. "Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls," Cao said. One way the startup has built its library of voices is by simply asking users to submit their own voices for training its models, and compensating them if their voices are used. That resulted in some trouble a few months ago, however, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA content take-down process in place to address such concerns, but the take-downs themselves took a long time. Fish Audio's CEO and co-founder Rissa Cao told TechCrunch that the company has now automated the take-down process. Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup's platform in less than 3 minutes, she said. Still, that doesn't prevent anyone from uploading an artist's voice without their knowledge. And until the artist finds out, their voice will continue to be used on the platform until they file for it to be taken down. Oskue Honda, a partner at Coreline Ventures, said a community-driven model only works when creators trust the platform. "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially," he said. Cao said when the startup was only offering its product as an open-source project with plans for creators, it was running efficiently and didn't need money. But it wanted to develop more advanced models, and also wanted to accommodate enterprises as investor interest was ramping up, which led it to seek capital. Looking ahead, Fish Audio plans to release an audio understanding model this year. It's also building a speech-to-speech model. The speech generation market is crowded, with companies like ElevenLabs , WellSaid , Cartesia, Speechify, Async (previously Podcastle) , and Krisp competing for creators and enterprises' wallets. According to Rico Mallozzi, a partner at 359 Capital, fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs. "I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices," Mallozzi told TechCrunch over a call. Topics AI , Fundraising , open-source , text to speech , voice AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Ivan Mehta Ivan covers global consumer tech developments at TechCrunch. He is based out of India and has previously worked at publications including Huffington Post and The Next Web. You can contact or verify outreach from Ivan by emailing im@ivanmehta.com or via encrypted message at ivan.42 on Signal. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 toda y! REGISTER NOW Most Popular Librarians are hosting viral ‘Avoiding AI' workshops for people who are fed up with Big Tech Amanda Silberling SpaceX launches new V3 Starlink satellites but suffers another booster failure Sean O'Kane Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100M Marina Temkin US accuses American of allegedly wiping his phone using a ‘duress' password during border search Zack Whittaker Anduril reportedly in talks to raise funding at $100B valuation, more than 3x last year's mark Ram Iyer Tesla's robotaxis are moving in reverse Sean O'Kane How OpenAI’s human mistake led to the AI-powered hack on Hugging Face Lorenzo Franceschi-Bicchierai