전 오픈AI CTO 미라 무라티가 설립한 씽킹 머신 랩(Thinking Machines Lab)이 첫 오픈 웨이트(Open-weight) 기반 AI 모델인 '잉클링(Inkling)'을 공개했습니다. 이 모델은 단순히 성능을 경쟁하기보다 기업이 자체적으로 맞춤화하여 사용하는 데 초점을 맞추고 있으며, 범용 AI 모델보다 맞춤형 모델이 더 높은 가치를 창출한다는 업계의 흐름을 강력하게 지지합니다. 기업들은 이 모델을 시작점 삼아 자사의 데이터와 전문성을 결합해 비용과 효율성을 모두 잡는 최적화된 AI를 구축할 수 있습니다.
번역된 본문
전 오픈AI CTO 미라 무라티(Mira Murati)가 설립한 AI 스타트업 씽킹 머신 랩(Thinking Machines Lab)이 수요일 오전 첫 자체 AI 모델인 '잉클링(Inkling)'을 출시했습니다. 오픈AI, 앤스로픽(Anthropic), 구글의 주력 모델들과 달리 이 모델은 오픈 웨이트(Open-weight) 방식을 채택하여 외부 개발자와 기업들이 모델을 다운로드하고 직접 수정할 수 있습니다.
회사의 공개 자료에 따르면, 잉클링은 총 9,750억 개의 매개변수(parameters)를 가진 전문가 혼합(Mixture-of-Experts) 시스템입니다. 하지만 특정 작업을 수행할 때는 그중 일부인 약 410억 개의 매개변수만 사용하여 초대형 모델을 더 빠르고 저렴하게 실행할 수 있는 일반적인 설계를 따릅니다. 또한 45조 개의 토큰 규모의 텍스트, 이미지, 오디오, 비디오 데이터로 학습되었으며, 이 세 가지 모드를 기본적으로 통합해 처리 및 추론할 수 있습니다. 이는 회사가 약 1년 반 동안 대중의 눈을 피해 AI 인프라를 구축해 온 이후 첫 공개 성과입니다. 이러한 작업의 일부는 이미 5월에 '상호작용 모델(interaction models)'에 대한 연구 미리 보기를 통해 공개된 바 있습니다. 이는 일반적인 챗봇처럼 멈추고 기다리는 대신, 듣고 말하며(심지어 말을 끊고) 사용자와 소통하도록 설계된 AI였습니다.
이는 씽킹 머신의 핵심 전략을 시험하는 자리이기도 합니다. 즉, 조직이 직접 맞춤 설정할 수 있는 AI가 현재 최대 규모 연구소들이 판매하는 '획일적인(one-size-fits-all)' 모델보다 더 나은 성능을 발휘할 것이라는 청사진입니다. 잉클링은 매우 흥미로운 모델로, 무작정 추측하기보다는 불확실성을 명확히 표시하는 등 보정된 답변을 제공하도록 설계되었습니다. 또한 속도를 위해 타협해야 할 때 사용자가 '사고 노력(thinking effort)' 단계를 직접 조절할 수 있게 해줍니다. 회사에 따르면 특정 벤치마크에서 잉클링은 엔비디아의 네모트론 3 울트라(Nemotron 3 Ultra)와 동일한 코딩 성능을 달성하는 데 단 3분의 1의 토큰만 사용했습니다.
흥미로운 점은 씽킹 머신이 잉클링을 동급 최강의 모델이라고 주장하지 않는다는 것입니다. 회사의 브리핑 자료는 잉클링이 "현재 사용 가능한 폐쇄형 또는 개방형 모델 중 가장 강력한 모델은 아니다"라고 명시하고 있습니다. 대신 이 모델이 명백히 추구하는 것은 '균형 잡힌(전반적으로 뛰어난) 성능'입니다.
물론 이는 누가봐도 명백한 '기업용(엔터프라이즈) 제품'이라는 점을 넘어, 대체 누구를 타겟으로 한 제품인지라는 큰 의문을 낳습니다. 현재 씽킹 머신은 잉클링을 완성된 형태의 결과물이라기보다는, 조직이 회사의 모델 맞춤화 플랫폼인 '틴커(Tinker)'를 통해 직접 파인튜닝(fine-tune)할 수 있는 '시작점'으로 마케팅하고 있습니다. (오픈AI, 앤스로픽, 구글은 챗GPT, 클로드, 제미나이를 각각 처음부터 범용 챗봇으로 경쟁하기 위해 구축한 뒤 그 위에 에이전트 및 자율 기능을 추가하는 매우 다른 접근 방식을 취했습니다.)
씽킹 머신이 지난주 게시한 글은 분명 이번 출시를 위한 배경 설명을 위함이었습니다. 회사는 그 글에서 한 회사가 중앙 집중식으로 학습시킨 뒤 굳어버린 AI보다, 각 조직이 스스로 형태를 잡아가는 AI가 훨씬 더 성능이 뛰어나다고 주장했습니다. 왜냐하면 실제 현업의 엄청난 전문 지식은 그것을 가진 사람마다 매우 구체적이기 때문입니다. 더 넓은 의미에서, 중앙 집중형 연구소들은 모두에게 동일한 제품을 반복적으로 개선해가며 판매하는 반면, 자체 모델을 소유하고 맞춤화하려는 기업들은 훨씬 더 많은 가치를 창출할 수 있다는 것입니다.
이러한 주장은 점점 더 힘을 얻고 있습니다. 일요일에 게시된 블로그 포스트에서 오픈AI와 앤스로픽에 수십억 달러를 투자한 마이크로소프트의 사티아 나델라(Satya Nadella) CEO는 독점형(폐쇄형) AI 모델을 사용하는 기업들은 사실상 비용을 두 번 지불한다고 경고했습니다. 즉, 한 번은 구독 비용으로 지불하고, 또 한 번은 수많은 프롬프트와 수정을 통해 전달되는 비즈니스 지식(이는 향후 모델 버전에 흡수될 수 있음)을 넘겨주는 방식으로 지불한다는 것입니다.
페이스(Hugging Face)의 클레멘스 델랑게(Clem Delangue) CEO 역시 지난주 테크크런치(TechCrunch)와의 대화에서 비슷한 전망을 내놓았습니다. 그는 최첨단 모델(Frontier models)은 실험 및 고부가가치 작업용으로 점차 제한적으로 사용될 것이며, 대부분의 실제 프로덕션 AI 작업은 민간 또는 오픈소스 대안으로 이동할 것이라고 말했습니다. 이는 씽킹 머신이 구축하고자 하는 정확한 시장 분할입니다. 이러한 주장의 가장 확실한 증거는 최근 세계 최대의 헤지펀드인 브리지워터 어소시에이트(Bridgewater Associates)와의 프로젝트(참고로 브리지워터는 씽킹 머신의 투자자가 아닙니다)에서 나왔습니다. 양사의 연구원들은 기존의 오픈소스 모델을...
Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, released its first proprietary AI model Wednesday morning, called Inkling — and unlike the flagship models from OpenAI, Anthropic, or Google, it's open-weight, meaning outside developers and companies can download it and modify it directly. Inkling is a mixture-of-experts system with 975 billion total parameters, though it only draws on a fraction of that — about 41 billion — for any given task, a common design that keeps very large models faster and cheaper to run. It was trained on 45 trillion tokens of text, image, audio, and video, and reasons natively across all three, according to the company's own release materials. It's the company's first public proof point after a year and a half spent building AI infrastructure largely out of public view. Some of that work surfaced already, in a May research preview of "interaction models" — AI designed to listen and speak (and even interrupt) instead of stop and wait as with typical chatbots. It's also a test of the central bet behind Thinking Machines, which is that AI that organizations can adapt for themselves will outperform the one-size-fits-all models the biggest labs currently sell. It's an interesting model, one that's designed to give calibrated answers, including flagging uncertainty rather than guessing, and which lets users dial "thinking effort" up or down when they want to trade for speed. On one benchmark, the company says, Inkling uses a third as many tokens as Nvidia's Nemotron 3 Ultra in order to hit the same coding performance. It's worth noting that Thinking Machines doesn't claim Inkling is best-in-class. Its briefing materials state explicitly that Inkling is "not the strongest model available today, closed or open." What it's evidently going for instead is well-rounded performance. Of course, that raises a big question, which is who this product is targeting, beyond the obvious — this is definitely an enterprise product. Thinking Machines is, for now, marketing it less as a finished work than as a starting point, something for organizations to fine-tune themselves through Tinker, the company's model-customization platform. (OpenAI, Anthropic, and Google have all taken a very different approach with ChatGPT, Claude, and Gemini, respectively, which were all built to compete as general-purpose chatbots first, with agentic, autonomous features layered on top.) A post published by Thinking Machines last week was clearly meant as the backdrop for this release. AI that's trained centrally by one company and then set in stone, the company argued in that post, underperforms AI that organizations shape themselves because so much expertise is specific to the people who hold it. The broader idea is that centralized labs are selling everyone the same product, repeatedly refined by the lab that built it, while enterprises willing to own and customize their own models can wring far more value from them. It's an argument that's gaining steam. In a blog post published Sunday, Microsoft CEO Satya Nadella — whose company has invested billions in both OpenAI and Anthropic — warned that enterprises using proprietary AI models effectively pay twice: once in subscription costs, and again by handing over business knowledge embedded in their thousands of prompts and corrections, which can be absorbed into future model versions. Hugging Face CEO Clem Delangue made a similar prediction in conversation with TechCrunch last week. Frontier models, he said, will increasingly be reserved for experimentation and high-value tasks, while most production AI work shifts to private or open-source alternatives — the exact split Thinking Machines is building around. The clearest evidence for that argument came recently from a project with Bridgewater Associates, the world's largest hedge fund (which is not, for what it's worth, a Thinking Machines investor). Researchers from both companies took an existing open-source model and trained it further on Bridgewater's own financial expertise. The result scored 84.7% on financial reasoning tests, beating top proprietary AI models, while costing roughly a fourteenth as much to run, though those results, published jointly in late June, come from the two companies' own evaluation, not an independent one. Thinking Machines has also emphasized how quickly it got here: OpenAI took roughly five years, and Anthropic roughly three, to bring tech to market and show revenue; Thinking Machines says it did the same in about nine months. Some will wonder whether Inkling was trained on outputs from competitors' models, a practice known as distillation that has drawn scrutiny industry-wide. The short answer, per the company's own materials, is partly. Thinking Machines pretrained Inkling from scratch, but it says it used other open-weight models — including Moonshot AI's Kimi K2.5 — to help generate some of its early post-training data before large-scale reinforcement learning took over. The next model, the company insists, will use fully self-contained post-training instead. On the cost side, Thinking Machines has been more guarded. It struck a strategic partnership with Nvidia in March to deploy a gigawatt of Vera Rubin computing capacity, and says Inkling itself was trained entirely on Nvidia's GB300 NVL72 systems. But the company hasn't said how it plans to balance that against revenue that, by most accounts, hasn't been a primary focus so far. (A reported $50 billion fundraising round was said to be coming together last November, which multiple outlets reported had stalled by January; the company has declined to talk about its funding picture since, though Nvidia said it made a "significant investment" in Thinking Machines when the companies announced that March partnership.) A related question is whether Thinking Machines' spending will ever reach the scale of OpenAI's or Anthropic's, or whether its efficiency-driven approach means the economics look different. Put another way, the company's bet may be less that it will eventually spend like its larger rivals than that it won't need to at all — because once weights are public, nothing obligates anyone who downloads them to pay Thinking Machines to run them, unlike the metered access OpenAI and Anthropic sell. It's Tinker, not the model itself, where the company's revenue has to come from, via training, fine-tuning, and, now, a cut of the hosting ecosystem built around it. Headcount, at least, looks more settled. Thinking Machines now employs roughly 200 people, up from levels reported after a wave of departures earlier this year, including two co-founders who left for OpenAI in January. Thinking Machines, for its part, doesn't seem interested in playing up individual moves the way much of the industry does. According to a source inside the company, its culture, by design, favors continuity over reliance on any one personality. It makes sense: it's less of a setback when people change teams if they were never put on a pedestal to begin with. It's also a remarkable thing for a company to insist on, given how much of its own story is still associated with the name of its now-famous co-founder, whether she planned it or not. Topics AI , Mira Murati , TC , thinking machine labs When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Connie Loizos Editor in Chief & General Manager Loizos has been reporting on Silicon Valley since the late ’90s, when she joined the original Red Herring magazine. Previously the Silicon Valley Editor of TechCrunch, she was named Editor in Chief and General Manager of TechCrunch in September 2023. She’s also the founder of StrictlyVC, a daily e-newsletter and lecture series acquired by Yahoo in August 2023 and now operated as a sub brand of TechCrunch. You can contact or verify outreach from Connie by emailing connie@strictlyvc.com or connie@techcrunch.co