메뉴
BL
The Decoder • 56일 전

미라 무라티 AI 랩, '잉클링 스몰' 출시... 효율성 승부

IMP
8/10
핵심 요약

전 오픈AI CTO 미라 무라티의 AI 연구소 씽킹 머신스(Thinking Machines)가 두 번째 모델인 '잉클링 스몰(Inkling Small)'을 공개했습니다. 이 모델은 기존 대형 모델 대비 파라미터를 1/3 수준으로 줄였음에도 코딩 및 추론 벤치마크에서 더 뛰어난 성능을 기록하며 압도적인 토큰 효율성을 입증했습니다. 아파치 2.0(Apache 2.0) 라이선스로 가중치가 공개되어 기업들이 자체 데이터로 파인튜닝(fine-tuning)하기에 매우 적합한 모델로 평가받습니다.

번역된 본문

씽킹 머신스(Thinking Machines), 두 번째 모델 '잉클링 스몰(Inkling Small)'로 크기보다 효율성에 승부를 걸다. 마티아스 바스티안(Matthias Bastian) - 2026년 7월 31일.

전 오픈AI CTO 미라 무라티(Mira Murati)가 설립한 AI 연구소 씽킹 머신스가 '잉클링 스몰(Inkling Small)'을 출시했습니다. 벤치마크 평가 기관 Artificial Analysis에 따르면, 이 오픈 웨이트(open-weights) 추론 모델은 지능 지수(Intelligence Index)에서 40점을 기록했습니다. 이는 기존 모델인 잉클링(41점)보다 1점 낮지만, 파라미터(parameters)는 3분의 1 이하(총 2,760억 개, 활성 120억 개)로 줄인 수치입니다. Artificial Analysis는 동일하거나 더 작은 규모의 오픈소스 모델 중 이보다 높은 점수를 기록한 모델이 없다고 밝혔습니다.

잉클링 스몰은 '휴머니티스 라스트 엑스(Humanity's Last Exam)'(32% vs 30%)와 'GPQA Diamond'(89% vs 87%)를 비롯한 여러 코딩 및 추론 테스트에서 기존의 더 큰 형제 모델을 능가하는 성능을 보였습니다. 에이전트 기반 작업(agent-based tasks)과 사실적 지식에서는 다소 뒤처지지만, 토큰 효율성은 훨씬 뛰어납니다. Deepseek V4 Flash(4만 5천 토큰)나 GPT-5.4 mini(7만 8천 토큰)와 비교했을 때, 작업당 평균 2만 4천 개의 출력 토큰만을 사용합니다.

이 모델은 텍스트, 이미지 및 음성 입력을 처리할 수 있으며, 25만 6천(256K) 토큰의 컨텍스트 윈도우(context window)를 지원하고 아파치 2.0(Apache 2.0) 라이선스 하에 배포됩니다. 모델 가중치는 허깅페이스(Hugging Face)에 공개되어 있으며, 사용자는 '팅커 플레이그라운드(Tinker Playground)'를 통해 브라우저에서 직접 파인튜닝(fine-tune)할 수 있습니다. 씽킹 머신스는 자사의 모델을 사용자가 자체 데이터로 파인튜닝할 수 있는 강력한 기반 모델로 포지셔닝하고 있으며, 일각에서는 이것이 AI의 다음 전선이 될 것으로 보고 있습니다.

원문 보기
원문 보기 (영어)
Thinking Machines bets on efficiency over size with its second model, Inkling Small Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 31, 2026 Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. According to Artificial Analysis , the open-weights reasoning model scores 40 on the Intelligence Index, one point below Inkling (41), with less than a third of the parameters (276 billion total, 12 billion active). AA says no open model of equal or smaller size scores higher. Inkling Small beats its bigger sibling on several coding and reasoning tests, including Humanity's Last Exam (32% vs. 30%) and GPQA Diamond (89% vs. 87%). It falls behind on agent-based tasks and factual knowledge but is far more token-efficient, averaging 24K output tokens per task compared to 45K for Deepseek V4 Flash and 78K for GPT-5.4 mini. The model handles text, image, and speech inputs, has a 256K-token context window, and ships under Apache 2.0. Weights are on Hugging Face , and users can fine-tune it in the browser via Tinker Playground . Thinking Machines positions its models as a foundation for fine-tuning with users' own data . Some see this as the next frontier in AI . Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Thinking Machines | Artificial Analysis Ask about this article… Search