메뉴
HN
Hacker News 11일 전

모질라 CTO가 전하는 오픈소스 AI의 현주소

IMP
9/10
핵심 요약

오픈소스(오픈 웨이트) AI 모델이 폐쇄형 최상위 모델과의 성능 격차를 거의 좁혔으며, 추론 비용은 36개월 만에 50배 이상 폭락했습니다. 이제 실무의 대부분의 작업 부하에서는 오픈소스 모델이 주력으로 사용되고 있으며, 기술적 가치는 단순한 모델 자체를 넘어 이들을 유기적으로 연결하는 에이전트 환경으로 이동하고 있습니다. AI 주도권이 특정 빅테크에 독점되지 않도록 상호운용성과 경쟁을 보장하는 오픈소스 생태계의 중요성이 그 어느 때보다 커졌습니다.

번역된 본문

모질라(Mozilla) CTO 라피 크리코리안(Raffi Krikorian)의 메시지

"뉴질랜드 최북단에서 마오리 방송인은 데이터를 자신들의 공동체에 묶어두는 라이선스 하에 '테 레오(te reo)'를 위한 음성 모델을 훈련합니다. 이 언어는 시장 규모가 너무 작아 기존 기업들의 관심을 끌지 못했죠. 세계 최대 회계법인 중 하나인 PwC는 재무 용어에 맞춰 오픈 모델을 미세조정(fine-tuned)했으며, 오늘날 자체 하드웨어에서 토큰당 비용을 지불할 필요 없이 수백 명의 고객을 위해 이를 운영하고 있습니다. 로잔의 연구원들은 적십자사와 함께 인도적 가이드라인에 맞춰 조정된 의료용 오픈 모델을 구축했으며, 자국과 탄자니아에서 임상 시험을 준비하고 있습니다. 동아프리카에서 농부들은 클라우드가 닿지 않는 들판에서 오프라인으로 작동하는 폰 자체의 모델을 사용해 카사바(cassava) 질병을 진단합니다. 스위스의 한 공공 컨소시엄은 공용 슈퍼컴퓨터에서 국가적 모델을 훈련하고 가중치, 데이터, 훈련 코드를 모두 공개했습니다. 이들 중 누구도 허락을 구하지 않았으며, 이런 방식은 기업의 유료 서비스를 빌리는 것으로는 불가능했을 것입니다. 그들은 이를 스스로 소유합니다. 이것이 오픈소스의 핵심 아이디어입니다.

우리는 이전에도 이런 상황을 겪었습니다. 모질라가 탄생한 이유도 한 기업이 웹의 현관문을 독점하려 했을 때, 오픈 커뮤니티가 일어나 그것을 막았기 때문입니다. 25년이 지난 지금, 누군가가 똑같은 패턴을 시도하고 있습니다. 우리는 지난번에 '오픈'에 베팅했고, 오픈이 승리했습니다. 함께한다면 우리는 다시 그렇게 할 수 있습니다. 우리의 믿음은 단순합니다. 앞으로 나아갈 길은 경쟁과 상호운용성(interoperability)에 있습니다. 우리는 수많은 모델이 존재하고, 이들을 표준적인 방식으로 연결할 수 있으며, 언제든 특정 벤더(vendor)의 속박에서 벗어날 자유가 있는 세상을 믿습니다. 오픈소스는 이에 대한 실적이 있습니다. 오픈소스는 파이를 키워 더 많은 사람들이 몫을 가질 수 있게 했습니다. 이어지는 내용을 하나의 지도로 읽어보세요: 오픈 AI가 승리하는 곳 (일부 수치는 우리조차 놀라게 했습니다), 그리고 취약점이 드러나는 곳. 약점을 숨기는 사례는 그저 광고일 뿐입니다."

라피의 전체 서한 읽기 → 여기에서 보고서 다운로드 ↓

오픈 웨이트(Open weights)가 성능 격차를 좁히는 동안 인공지능의 가격은 붕괴했습니다. 0% 최상위 폐쇄형(Closed) 모델과의 성능 격차 — 코딩에서는 동등한 수준, 추론에서는 뒤처짐 0× 36개월 동안 GPT-4 급 추론 비용의 하락: 100만 토큰당 $20 → $0.40

01 오픈소스 AI의 현재 상태

동등한 수준(Parity)에 도달했습니다. 이제 경쟁은 한 단계 더 높은 곳에서 벌어집니다. 오픈 웨이트(Open weights)는 더 이상 타협점이 아닙니다. 현재 실무 작업이 일어나는 곳입니다. 이제 실제 프로덕션 환경에서 처리되는 대부분의 토큰이 오픈 모델을 통과하며, OpenRouter에서 사용량이 가장 많은 상위 5개 모델 역시 모두 오픈소스입니다. 폐쇄형 모델은 여전히 추론 및 멀티모달리티(Multimodality) 같은 최전방(Frontier) 기술에서 앞서지만, 최전방 기술은 대부분의 실제 작업 부하에 필요하지 않습니다. 범용화된 입력값은 가격 결정력을 유지할 수 없습니다. 가치는 상위 계층인 에이전트 하네스(Agentic harness)로 이동하고 있습니다.

성능 격차: 8.04% → 0.5% → 3.3% 지난 24개월 동안 Chatbot Arena에서 나타난 오픈 vs 폐쇄형 모델 간의 격차입니다. 2024년 8월, 격차는 0.5%로 좁혀졌고, 2025년 2월에는 DeepSeek-R1이 미국의 최상위 모델과 잠시 동등한 수준을 기록했습니다. 2026년 3월에는 폐쇄형 추론 모델이 앞서나가면서 격차가 다시 3.3%로 벌어졌습니다. 하지만 이 3.3%라는 수치는 불규칙한 최전방 기술 평균치일 뿐입니다. 오픈 모델은 코딩, 지시 따르기, 일반 상식에서 동등하거나 근접한 수준이며, 격차는 주로 고차원적인 추론, 긴 문맥 검색, 에이전트 작업에 집중되어 있습니다. 따라서 중요한 것은 오픈 모델이 '충분히 좋은가?'가 아니라 '내 작업에 무엇이 필요한가?'입니다.

출처: Chatbot Arena, 2024년 1월 – 2026년 3월.

추론 비용은 36개월 만에 50배 하락했습니다. GPT-4 동급 모델의 100만 토큰당 가격 — 닷컴 버블 시대의 대역폭이나 PC 연산 비용 하락 곡선보다 더 빠른 속도입니다. (로그 스케일) 출처: Stanford HAI AI Index 2025 (18개월 동안 GPT-3.5 급 하락률 280배); Epoch AI (연간 9900배 감소); 2025년 11월 MIT 연구 (최전방 기술 기준 하드웨어를 고려하여 연간 510배 하락).

오픈 웨이트가 토큰 점유율을 가져갑니다 OpenRouter에서 오픈 웨이트 모델을 통해 라우팅되는 토큰의 비율은 미미한 수준에서 2025년 말 3분의 1, 2026년 중반에는 과반수를 차지할 만큼 성장했습니다. 출처: OpenRouter 100조 토큰 연구 (2024년 11월 – 2025년 11월) 및 실시간 리더보드; 중간 지점은 보간법 적용.

단, 요청 수(Request count)를 기준으로 하면 여전히 미국의 폐쇄형(Closed) 공급업체들이 선두를 달리고 있습니다. — 오픈 모델의 선두는 토큰 사용량 기준입니다.

원문 보기
원문 보기 (영어)
A Letter From Our CTO, Raffi Krikorian “ In New Zealand's far north, a Māori broadcaster trains speech models for te reo — a language too small for any market — under a license that keeps the data with its people. PwC, one of the largest accounting firms in the world, fine-tuned an open model on the language of finance and runs it today for hundreds of clients, on its own hardware, with no per-token meter running. Researchers in Lausanne built an open medical model with the Red Cross, tuned to its humanitarian guidelines, and are preparing clinical trials at home and in Tanzania. In East Africa, farmers diagnose cassava disease with a model that runs on the phone itself, offline, in fields the cloud has never reached. In Switzerland, a public consortium trained a national model on public supercomputers and released all of it: weights, data, training code. None of them asked permission, and none of them could have rented this. They own it — that is the whole idea. We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community rose up to make sure it never could. Twenty-five years later, someone is running the same play. We bet on open the first time. Open won. Together, we can do it again. Our belief is simple: the path forward is competition and interoperability. We believe in a world of many models, standard ways to plug them together, and the freedom to walk away from any vendor at any time. Open has a record here. It grew the pie and let more people own a slice of it. Read what follows as a map: where open AI is winning — some numbers surprised even us — and where it is exposed. A case that hides its weak points is an advertisement.” Read Raffi's full letter here → Download the report here ↓ Open weights closed the capability gap while the price of intelligence collapsed. 0% Capability gap to the top closed models — at parity on coding, behind on reasoning 0× Fall in GPT-4-class inference cost in 36 months: $20 → $0.40 per 1M tokens 01 The current state of open-source AI Parity reached. The contest is one layer up. Open weights are no longer a compromise. They are where the work happens: a majority of production tokens now route through them, and the five highest-volume models on OpenRouter are all open. Closed models still lead at the frontier, on reasoning and multimodality, but the frontier is not what most workloads need. Commodity inputs do not hold pricing power. Value moves up, to the agentic harness. The capability gap: 8.04% → 0.5% → 3.3% Open-vs-closed gap on Chatbot Arena over 24 months. By August 2024, the gap had collapsed to 0.5%, and in February 2025 DeepSeek-R1 briefly matched the top US model. By March 2026 it had reopened to 3.3% as closed reasoning models pulled ahead. But 3.3% is an average over a jagged frontier: open is at or near parity on coding, instruction-following and general knowledge, while the gap concentrates in reasoning, long-context retrieval and agentic tasks. The question is no longer whether open models are good enough. It's what you need for your workload. Hover the points. Source: Chatbot Arena, Jan 2024 – Mar 2026. Inference fell 50× in 36 months GPT-4-equivalent price per 1M tokens — faster than dotcom-era bandwidth or PC-compute price curves. Log scale. Sources: Stanford HAI AI Index 2025 (280× GPT-3.5-class drop over 18 months); Epoch AI (9–900× annual decay); Nov 2025 MIT study (5–10×/yr at the frontier, hardware-adjusted). Open weights win the tokens The share of tokens routed on OpenRouter through open-weight models grew from a negligible base to a third by late 2025 to a majority by mid-2026. Source: OpenRouter 100T-token study (Nov 2024–Nov 2025) and live leaderboard; intermediate points interpolated. By request count, closed US providers still lead — the open lead is a token-volume lead, concentrated in coding and agentic workloads. OpenRouter live leaderboard — trailing month, tokens routed The five highest-volume models are all open weights. Anthropic's closed Claude models are the next US-built entrants. Open weights Closed By mid-2026 the top nine models route roughly 18T weekly tokens for Chinese-built models against ~5.5T for US-built ones — more than 3:1 (FT analysis). Where developers route by cost, they route to open weights. Open ships easy. Open deploys hard. Data from the Mozilla / SlashData 2026 developer survey. Open models lead in adoption: 79% of developers adding AI functionality use them, against 71% for closed, and the two are largely complementary, with half of developers using both. But production is where teams stall: only 51% of open-model teams reach production versus 63% for closed. The gap is operational tooling and trust, not model capability. Open models lead in adoption, and mostly coexist with closed Share of developers adding AI functionality to their applications who currently use each model type, and how the two overlap. Open models 79% Closed models 71% How they combine 29% open only Use open-source models exclusively" style="width:29%;background:var(--green)">29% OS only 50% use both The two are largely complementary" style="width:50%;background:var(--seed-1)">50% Both 21% closed only Use closed-source models exclusively" style="width:21%;background:var(--grey-1)">21% CS only Source: Mozilla / SlashData 2026 developer survey. Open and closed aren't substitutes for most teams: 50% run both, 29% open only, 21% closed only. Where open adoption peaks, and where closed still edges it Open-model adoption by region. Greater China and East Asia lead at 89%; South America and Western Europe are the only two regions where closed adoption exceeds open. Same survey, by developer region. In South America and Western Europe, and only there, closed-model adoption runs ahead of open. Production rate by company size If the gap were about resources, scale would close it, and it doesn't. Closed climbs 54% → 73% with scale. Open barely moves: 53% → 57%. Closed models Open models Enterprises can buy their way through closed deployment. Open deployment waits on tooling nobody has finished. Source: Mozilla / SlashData 2026 developer survey. Why teams churn: challenges with open models Δ = churned − still using, in percentage points. The biggest gaps (performance, integration, maintenance) are operational, not capability. Hover the bars. Still using open Churned away Mozilla survey, n=1,410. “What are the main challenges you face when working with open or open-source AI models?” The same challenges, everywhere: what blocks open by region Share of current and churned open-model developers naming each challenge, by region. Warmer cells mean more developers blocked. The top rows are operational in every region: infrastructure cost, security and compliance, maintenance, deployment complexity. South Asia leans hardest on security and support; only North America and Greater China have more than 15% reporting no major challenges. Challenge W. Europe & Israel N. America Greater China South Asia East Asia ex GC S. America E. Europe & CIS Oceania All High infrastructure or compute costs 25% 26% 29% 28% 28% 28% 29% 18% 27% Security, privacy, or compliance concerns 20% 27% 18% 39% 29% 28% 25% 22% 26% Ongoing maintenance and updates 27% 26% 18% 26% 20% 31% 21% 25% 24% Complexity of deployment, hosting, or scaling 27% 24% 19% 24% 11% 30% 26% 25% 23% Lack of specialised support 17% 16% 21% 31% 24% 23% 23% 32% 22% Difficulty evaluating or comparing models 14% 17% 14% 23% 16% 26% 25% 18% 18% Difficulty fine-tuning or customising 22% 18% 18% 20% 11% 22% 18% 12% 18% Difficulty integrating into existing systems 19% 21% 14% 20% 7% 26% 19% 20% 18% Insufficient documentation or learning resources 18% 15% 15% 17% 15% 20% 24% 15% 17% Model performance is not good e