메뉴
BL
TechCrunch AI • 22일 전

메타, 신규 AI 모델 사용 데이터 제공 시 최대 95% 할인

IMP
7/10
핵심 요약

메타가 새 AI 모델 '뮤즈 스파크(Muse Spark)'의 프롬프트와 출력 데이터를 공유하는 사용자에게 평균 약 95%의 요금 할인을 제공한다. 이는 사용자 데이터 확보에 어려움을 겪는 메타가 기업들에게 명시적인 보상을 제공해 학습 데이터를 확보하려는 전략으로, 프론티어 AI 기업 간 가격 경쟁과도 맞물려 주목된다.

번역된 본문

대부분의 AI 도구는 향후 버전 개선을 위해 모델 제공자와 사용 기록을 공유하는 것을 거부(옵트아웃)할 수 있다. 메타는 이 개념에 가격표를 달았다. 코딩 및 기타 에이전트 운영용으로 설계된 신모델 '뮤즈 스파크(Muse Spark)'에서, 메타는 자신의 프롬프트와 모델 출력을 공유해 미래 모델 개발에 '기여'하는 사용자에게 평균 약 95%에 달하는 명시적 할인을 제공한다. 표준 계약 기준 입력 토큰 100만 개당 1.25달러인 반면, 기여자 요금제에서는 단 10센트다. 출력 토큰은 표준 가격이 100만 개당 4.25달러이지만 기여자 모델에서는 20센트에 불과하다.

메타는 학습 데이터 확보에 어려움을 겪어왔다. 올해 초 자사 직원의 컴퓨터 사용을 추적하려는 시도는 내부에서 광범위한 비판을 받아 6월에 중단됐다. 메타는 새로운 요금제에 대한 테크크런치(TechCrunch)의 질의에 답변하지 않았다.

이런 사용자 데이터는 에이전트 도구의 성능 향상에 필수적이다. 오픈소스 하네스 'Pi'의 개발자인 마리오 체흐너(Mario Zechner)는 지난달 테크크런치에 "2025년 4월과 10월 사이 코딩 에이전트 능력이 크게 향상된 이유는 클로드 코드(Claude Code)가 기본적으로 모든 코딩 에이전트 세션을 저장해 강화학습 훈련에 사용했기 때문"이라고 말했다.

하지만 모델 개발사들이 점점 소프트웨어 엔지니어링 외부로 에이전트 도구 배치에 나서면서, 많은 전문 워크플로의 복잡성과 디지털 흔적 부족 때문에 이런 도구를 평가하고 개선하는 능력이 제약받고 있다.

프린스턴대 컴퓨터과학과 교수인 아르빈드 나라야난(Arvind Narayanan)은 대기업들이 자사 데이터가 모델 학습에 쓰이는 것을 원하지 않는다는 명확한 증거가 있다고 지적했다. 그는 소셜 미디어에서 "클로드 맥스(Claude Max)나 챗GPT 프로(ChatGPT Pro) 같은 구독형 소비자 플랜이 10~20배 이상 저렴함에도 불구하고, 대기업들은 토큰 과금 기반 엔터프라이즈 플랜을 고수한다(두 플랜의 주요 차이는 데이터 보존 여부와 엔터프라이즈 IT 거버넌스다)"라고 썼다.

어쩌면 이런 상황을 인식해서일까, 메타는 그 정보를 얻기 위해 기업들에게 명시적인 보상을 제공하고 있다. 메타의 가격 가이드는 기여자 티어가 "데이터 학습이 허용되는 프로토타이핑, 통합 테스트, 실험 확장의 진입 장벽을 낮춘다"고 설명한다. 나라야난은 이 프레임워크가 대기업들에게 어떤 데이터가 진정한 영업 비밀이고 어떤 데이터를 모델 제공자와 공유할 수 있는지 더 꼼꼼히 구분하도록 유도할 수 있다고 제안했다.

이런 구조는 프론티어 AI 기업 간 intensifying 가격 경쟁에도 영향을 미칠 수 있다. 어제 출시된 앤스로픽(Anthropic)의 최신 '페이블(Fable)'과 '미토스(Mythos)' 모델은 캐시된 토큰 처리 비용을 인하했고, 오픈AI(OpenAI)도 7월 말 최신 모델의 가격을 대폭 인하했다.

원문 보기
원문 보기 (영어)
Most AI tools allow you to opt out of sharing your usage with the model provider to improve future versions. Meta has taken that idea and put a price tag on it. For its new Muse Spark model, intended for operating coding and other agents, Meta is offering an explicit discount averaging out to about 95% for users who "contribute" to the development of future models by sharing their prompts and model outputs. While 1 million input tokens under a standard agreement costs $1.25, under the contributor pricing model they cost just 10 cents. For output tokens, the standard price is $4.25 per million, but that same million costs just 20 cents under the contributor model. Meta has had a rough time trying to obtain training data: An initiative to track the computer usage of its employees, launched earlier this year, attracted wide internal criticism and was paused in June. The company didn't respond to a question from TechCrunch about its new pricing model. This kind of user data is vital for making agentic tools work better. "The reason we saw a big jump in [coding agent] capabilities between April 2025 and October 2025 was that Claude Code, by default, would store all your coding agent sessions and use them for reinforcement learning training," Mario Zechner, the developer behind the open source harness Pi, told TechCrunch last month. But even as the imperative for model builders increasingly becomes deploying agentic tools for use outside of software engineering, their ability to evaluate and improve those tools is blocked by the complexity and lack of digital traces for many professional workflows. Arvind Narayanan, a Princeton computer science professor, noted that there is good evidence that large companies don't want their data to be used for model training. "They stick with token-billed Enterprise plans even though the subscription-based consumer plans like Claude Max and ChatGPT Pro are discounted by 10x-20x or even more! (The main difference between the plans is data retention + enterprise IT governance)," he wrote on social media. Perhaps in recognition of those dynamics, Meta is offering companies explicit compensation to obtain that information. Its pricing guide notes that the contributor tier "lowers the barrier to entry for prototyping, testing integrations, and scaling experiments where training on your data is acceptable." That, Narayanan suggested, could in turn incentivize large companies to be more diligent about which data is truly proprietary and which could be shared with model providers. The framework could also play into growing price competition between the frontier labs. Anthropic's newest Fable and Mythos models, released yesterday, came with lowered costs for processing cached tokens, while OpenAI's latest models got major price cuts at the end of July. Topics AI , Meta When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio October 13 - 15 San Francisco Don't miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era? REGISTER NOW Most Popular Norway considers ban on camera-enabled wearable ‘pervert glasses' Zack Whittaker Uber is laying off 10% of staff, or 3,300 people Ram Iyer AfterQuery reportedly becomes Y Combinator's fastest-ever unicorn, now valued at $3.2B Julie Bort GoPro to be acquired for $285M, will remain a public company Sean O'Kane Apple shares ‘shocking evidence' against former employee accused of stealing company data for OpenAI Amanda Silberling Microsoft tests fix for latest hours-long Outlook outage Sarah Perez MapQuest's app surges to No. 1 in Navigation after refusing to rename Lake Ontario Sarah Perez