메뉴
BL
The Decoder 13일 전

전 오픈AI CTO 미라 무라티, 975B '잉클링' 모델 공개

IMP
8/10
핵심 요약

전 오픈AI CTO 미라 무라티가 설립한 Thinking Machines Lab이 텍스트, 이미지, 오디오를 처리하는 975B 규모의 멀티모달 오픈 웨이트 모델 '잉클링(Inkling)'을 공개했습니다. 이 모델은 미국 내 오픈소스 모델 중 최고 성능을 자랑하지만, 63%에 달하는 높은 환각 현상과 비교적 높은 비용으로 인해 정확성이 필수적인 실무 적용에는 한계가 있습니다.

번역된 본문

전 오픈AI CTO 미라 무라티의 스타트업, Thinking Machines Lab이 9,750억 개(975B)의 매개변수를 가진 오픈 웨이트 모델 '잉클링(Inkling)'을 출시했습니다. 이 모델은 효율성과 에이전트 기반 작업을 위해 설계되었지만, 전체 성능 면에서는 여전히 최고 수준의 중국 오픈소스 모델들에는 뒤처집니다.

Thinking Machines Lab이 첫 상용 언어 모델을 선보였습니다. 잉클링(Inkling)은 총 9,750억 개의 매개변수를 가진 혼합 전문가(Mixture-of-Experts, MoE) 트랜스포머 모델로, 실행 시 410억 개의 매개변수가 활성화됩니다. 이는 챗GPT(ChatGPT) 개발에 핵심적인 역할을 했던 전 오픈AI CTO 미라 무라티가 설립한 스타트업이 내놓은 첫 모델입니다.

파인튜닝을 비즈니스 모델로 채택하다 다른 많은 오픈소스 AI 모델과 달리, 잉클링은 텍스트, 이미지, 오디오를 기본적으로 처리하며 최대 100만 토큰의 컨텍스트 윈도우를 지원합니다. 모델 가중치(weights)는 허깅페이스(Hugging Face)에서 무료로 다운로드할 수 있습니다. 또한 Thinking Machines는 AI 모델을 특정 작업에 맞게 조정할 수 있는 자체 플랫폼인 '텅커(Tinker)'를 통해 접근 권한을 제공합니다.

회사는 잉클링을 맞춤화할 수 있는 유연한 기본 모델로 포지셔닝하고 있습니다. 공지에서도 "잉클링은 오늘날 사용 가능한 가장 강력한 모델 전체는 아니다"라고 솔직하게 밝혔습니다. 대신 Thinking Machines는 멀티모달 지원, 효율적인 처리, 그리고 파인튜닝 옵션의 결합이 이 모델의 차별점이 될 것으로 기대하고 있습니다.

Thinking Machines에 따르면, 잉클링은 45조 개의 토큰 규모인 공개 및 합성 텍스트, 이미지, 오디오 녹음, 비디오 데이터를 사전 학습(Pre-train)했습니다. 이 학습 데이터 세트에는 "지적재산권 보호를 받을 수 있는" 공개 데이터도 포함되어 있습니다. 회사는 합성 데이터를 생성하는 방법 중 하나로 중국의 AI 모델인 'Kimi K2.5'를 활용했습니다. Kimi K2.5는 코드 에디터 커서(Cursor)의 코딩 모델의 기반으로도 사용된 바 있습니다. 자세한 기술적 세부 사항은 모델 카드(model card)에서 확인할 수 있습니다.

미국 오픈 모델을 이끌지만 중국의 최고 모델에는 뒤처져 AI 벤치마크 플랫폼인 Artificial Analysis에 따르면, 잉클링은 Artificial Analysis 지능 지수(Intelligence Index)에서 41점으로 데뷔했습니다. 이로써 미국 연구소가 배포한 오픈 웨이트 모델 중 선두를 차지했습니다. 이는 이전 선두주자였던 Nemotron 3 Ultra(38점)보다 3점 높은 수치이며, Gemma 4 31B(29점)와 gpt-oss-120b(24점)를 크게 앞서는 기록입니다.

지식 노동 작업을 시뮬레이션하는 에이전트 기반 벤치마크인 GDPval-AA v2에서 잉클링은 1,238의 Elo 레이팅을 기록했습니다. 이는 1,190점의 Kimi K2.6과 1,189점의 DeepSeek v4 Flash max를 뛰어넘는 점수입니다. 또한 은행 업무 벤치마크인 Tau-3에서도 24%의 점수로, Kimi K2.6(21%)과 DeepSeek v4 Flash max(23%)를 앞섰습니다.

하지만 사실 관계의 정확도 측면에서는 다소 부진한 성과를 보입니다. Artificial Analysis의 'AA 옴니사이언스(Omniscience)' 벤치마크에서 잉클링은 단 2점(+2)을 받았습니다. 이는 선두권 오픈 웨이트 모델들에는 미치지 못하는 수준이지만, -1점을 받은 Nemotron 3 Ultra와 같은 다른 미국 모델들보다는 여전히 앞서는 점수입니다. 특히 잉클링의 정확도는 40%이며, 환각(Hallucination) 비율은 63%에 달합니다. 이러한 결과는 고도로 정확한 정보가 필요한 애플리케이션에서 이 모델의 사용을 제한할 가능성이 높습니다.

비용 측면에서 잉클링은 64K 컨텍스트 윈도우를 기준으로 백만 입력 토큰당 1.87달러, 백만 출력 토큰당 4.68달러의 비용이 듭니다. 이는 GLM-5.2 및 DeepSeek과 같은 중국의 오픈소스 모델들보다 약간 더 비싼 가격대입니다.

원문 보기
원문 보기 (영어)
Ex-OpenAI CTO Murati's Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails China Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 16, 2026 Artificial Analysis Key Points Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, a multimodal open-weights model with 975 billion parameters that natively processes text, images, and audio. According to the analysis platform Artificial Analysis, Inkling is currently the most powerful U.S. open-weights model, outperforming competitors like Kimi K2.6 and DeepSeek v4 Flash max on agentic tasks while also demonstrating high token efficiency. Despite its strong benchmark performance, the model shows notable weaknesses in factual accuracy, with a hallucination rate of 63 percent, and comes at a higher cost than comparable Chinese models. Ask about this article… Search Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, an open-weights model with 975 billion parameters. It's built for efficiency and agent-based tasks, but it still trails the best open-source Chinese models in overall performance. Thinking Machines Lab has shipped its first production-ready language model. Inkling is a Mixture-of-Experts Transformer with 975 billion total parameters, 41 billion of which are active at any given time. It's the first model from the startup founded by Mira Murati, the former OpenAI CTO who played a key role in developing ChatGPT. Fine-tuning as a business model Unlike many other open-source AI models, Inkling natively handles text, images, and audio and supports a context window of up to one million tokens. The weights are freely available on Hugging Face . Thinking Machines also offers access through Tinker, its platform for adapting AI models to specific tasks. Ad The company is positioning Inkling as a flexible base model for customization. "Inkling is not the strongest overall model available today," the announcement states. Thinking Machines expects the mix of multimodal support, efficient processing, and fine-tuning options to set the model apart. Ad DEC_D_Incontent-1 Thinking Machines says it pre-trained Inkling on 45 trillion tokens of public and synthetic text, images, audio recordings, and videos. The training set also includes public data that "may be subject to intellectual property protection." The company used the Chinese AI model Kimi K2.5, among other methods, to generate synthetic data. Kimi K2.5 also served as the basis for Cursor's coding model . More technical details are available in the model card . Inkling leads U.S. open models but trails China's best According to AI benchmarking platform Artificial Analysis , Inkling debuts with a score of 41 on the Artificial Analysis Intelligence Index. That makes it the leading open-weights model from a U.S. lab. It ranks three points above the previous leader, Nemotron 3 Ultra at 38, and well ahead of Gemma 4 31B at 29 and gpt-oss-120b at 24. Ad On GDPval-AA v2, an agent-based benchmark that simulates knowledge-work tasks, Inkling reaches an Elo rating of 1,238. It beats Kimi K2.6 at 1,190 and DeepSeek v4 Flash max at 1,189. Inkling also scores 24 percent on the Tau-3 banking benchmark, ahead of Kimi K2.6 at 21 percent and DeepSeek v4 Flash max at 23 percent. Inkling performs rather poorly on factual accuracy. Artificial Analysis gives the model a score of just +2 on its AA Omniscience benchmark. That puts it below the leading open-weights models, though still above other U.S. models such as Nemotron 3 Ultra at -1. Inkling's accuracy is 40 percent, while its hallucination rate is 63 percent. Those results are likely to limit its use in applications that need highly accurate information. Ad DEC_D_Incontent-2 With a 64K context window, Inkling costs $1.87 per million input tokens and $4.68 per million output tokens. That's slightly more than open-source Chinese models such as GLM-5.2 and DeepSeek v4, which offer similar or better performance on text and code tasks. For context windows up to 256,000 tokens, pricing rises to $3.74 for input, $0.748 for cached input, and $9.36 for output. Ad But Inkling uses fewer output tokens than comparable open-weights models. According to Artificial Analysis, it averages 25,000 output tokens per Intelligence Index task. GLM-5.2 max uses 43,000, Kimi K2.6 uses about 38,000, and DeepSeek v4 Pro max uses about 37,000 tokens on the same tasks. Thinking Machines says Inkling offers continuously adjustable "thinking effort." Users can choose their preferred balance between cost and performance, reducing token use while maintaining the same result quality. Inkling-Small beats the larger model on some benchmarks Thinking Machines is also previewing Inkling-Small , a more compact model with 276 billion total parameters and 12 billion active parameters. The smaller model delivers similar or better results than Inkling on several benchmarks. Inkling-Small scores 88.3 percent on GPQA Diamond, compared with 87.2 percent for Inkling. On the HLE benchmark with tools, it scores 46.6 percent, slightly ahead of Inkling at 46.0 percent. Thinking Machines credits changes to the pre-training data and training process for the results. The company plans to publish the full weights once testing is complete. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Thinking Machines