메뉴
HN
Hacker News 13일 전

오픈 웨이트 모델 '잉클링(Inkling)' 공개

IMP
8/10
핵심 요약

새로운 오픈 웨이트(Open-Weights) 기반 파운데이션 모델인 '잉클링(Inkling)'이 공개되었습니다. 이 모델은 총 975B(활성 41B) 매개변수와 최대 1M 토큰의 컨텍스트를 지원하며, 특히 텍스트, 이미지, 오디오를 아우르는 멀티모달推理 및 에이전트 코딩에 강점을 보입니다. 개발자들은 Tinker 플랫폼을 통해 이 모델을 손쉽게 파인튜닝(Fine-tuning)하여 다양한 실제 작업 환경에 맞춤형으로 활용할 수 있다는 점에서 의미가 큽니다.

번역된 본문

원문 제목: 잉클링: 우리의 오픈 웨이트 모델

[Tinker 모델 카드] [Hugging Face]에서 사용해 보세요.

우리의 사명은 인간의 의지와 판단을 확장하는 AI를 구축하는 것입니다. 우리는 누구나 모델을 맞춤 설정할 수 있는 플랫폼을 개발하고, 대화형 협업을 위해 구축된 AI 시스템을 미리 공개하며, 새로운 연구를 발표해 왔습니다. 오늘 우리는 사람들이 자신만의 모델로 만들 수 있도록 처음부터 직접 학습시키고 전체 가중치(weights)를 공개하는 모델을 출시함으로써 이 사명을 한 단계 더 발전시킵니다.

잉클링(Inkling)이라는 이름의 이 모델은 총 975B(9천7백5십억)의 매개변수와 41B(410억)의 활성 매개변수를 가진 전문가 혼합(Mixture-of-Experts, MoE) 트랜스포머입니다. 최대 1M(백만) 토큰의 컨텍스트 창을 지원합니다. 45조 개의 텍스트, 이미지, 오디오, 비디오 토큰으로 사전 학습(Pretrain)되었습니다. 이 모델은 다양한 크기의 모델 패밀리 중 첫 번째 모델입니다. 이와 함께, 유사한 방식으로 학습되어 더 낮은 비용과 대기 시간으로 강력한 성능을 달성하는 12B 활성 매개변수의 더 가벼운 모델인 '잉클링-스몰(Inkling-Small)'의 프리뷰도 공유합니다.

잉클링은 텍스트, 이미지 및 오디오에 대해 기본적으로 추론(Reasoning)하며, 효율적이고 제어 가능한 '사고(Thinking)' 노력을 통해 비용과 성능의 균형을 맞춥니다. 우리는 이 모델이 다양한 도메인에 걸쳐 강력하고, 적응할 수 있을 만큼 유연한 폭넓고 균형 잡힌 파운데이션 모델이 되도록 학습시켰습니다.

잉클링은 현재 사용 가능한 모델(오픈소스 및 폐쇄형 포함) 전체를 통틀어 가장 강력한 모델은 아닙니다. 대신, 멀티모달 기능, 효율적인 추론, Tinker를 통한 파인튜닝(Fine-tuning) 가용성 등 여러 특성의 결합으로 인해 맞춤형 설정(Customization)을 위한 훌륭한 오픈 웨이트(Open-weights) 기반 모델이 됩니다.

잉클링은 시작에 불과합니다. 우리가 계속해서 구축해 나갈 모델 패밀리의 첫 번째 출시입니다. 우리는 더 많은 사용 사례에 맞춤 설정을 접근하기 쉽게 만들고 싶기 때문에, 오늘부터 Tinker에서 잉클링을 파인튜닝할 수 있습니다. 파인튜닝을 위한 올바른 기반 모델을 선택하는 것은 측정 가능한 벤치마크와 모델을 직접 사용해 보면서 느껴지는 고유한 감각을 결합하는 정성적인 판단입니다. 우리는 후자를 가능하게 하기 위해 Tinker 콘솔에 '잉클링 플레이그라운드(Inkling Playground)'를 추가했습니다. 이는 개발자가 잉클링과 대화할 수 있는 인터페이스입니다. 맞춤 설정이 실제로 어떤 의미인지 보여주기 위해, 우리는 잉클링에게 스스로를 파인튜닝하도록 요청했습니다. Tinker를 사용하여 이 모델은 자체적인 파인튜닝 작업을 작성하고, 실행하고, 결과를 평가했습니다. 공유 가능한 버전이 출시되면 확인하실 수 있습니다.

--> 주요 기능

실제 응용 프로그램에는 파인튜닝을 통해 결합하고 개선할 수 있는 광범위한 기능을 갖춘 모델이 필요합니다. 우리는 잉클링이 할 수 있는 일과 신뢰성 및 안전성과 같은 중요한 특성에서 어느 정도 수준인지 보여줍니다.

일반목적 모델 (Generalist model) 잉클링은 폭넓은 활용을 목표로 설계되었습니다. 우리는 특정 도메인에 대해서만 좁게 최적화하는 대신, 에이전트(Agent), 추론, 코딩, 명령어 수행, 사실 확인, 비전(Vision) 및 오디오 작업에 걸쳐 이 모델을 학습시켰습니다. 이러한 폭넓음은 맞춤형 설정 및 실제 사용에 있어 중요합니다. 즉, 사용자들은 단순히 벤치마크에서 뛰어난 성적을 내는 것을 넘어, 매우 다양한 워크플로우에 적응할 수 있는 모델을 필요로 합니다.

에이전트 코딩 및 도구 활용 (Agentic coding and tool use) 강력한 파인튜닝 기반 모델은 에이전트의 도구 활용을 통해 다양한 작업을 유연하게 해결할 수 있어야 합니다. 잉클링은 대부분의 에이전트 벤치마크에서 오픈 웨이트 모델들 중에서도 상위권의 성능을 보여줍니다. 우리는 잉클링이 다양한 코딩 및 에이전트 환경 내부에서 실행되도록 학습시켰으며, 훈련 중에 도구 세트와 스키마를 무작위화하여 특정 환경에 대한 민감도를 줄였습니다. 다음 섹션에서 설명하는 잉클링의 제어 가능한 '사고(Thinking)' 노력은 환경 내에서 설정할 수 있습니다. 아래에는 잉클링의 에이전트 코딩 및 도구 활용 방식과 그 결과물을 보여주는 몇 가지 데모가 있습니다.

  • 임베디드 브라우저를 활용한 원샷 웹 앱 (One-shot web app with embedded browser use) 잉클링은 단 한 번의 실행으로 기능적인 웹 앱을 구축한 다음, 자연어 지시를 통해 웹 앱 인터페이스를 작동할 수 있는 임베디드 AI 어시스턴트를 구동합니다.

  • 디자인 아레나 (Design Arena) 잉클링은 디자인 아레나의 에이전트 웹 개발 리더보드에서 평가되었으며, 여기서 맹검 심사를 한 인간 평가자들이 생성된 웹 앱을 일대일로 비교합니다. 이 모델은 가장 강력한 오픈 웨이트 모델들 중 하나로 랭크되었습니다.

  • 통일된 스타일의 결과물 (Cohesively styled artifacts) 잉클링은 정확한 명령 수행, 정확한 정보, 그리고 전체에 걸친 일관된 스타일 및 디자인을 갖춘 여러 페이지의 결과물을 생성합니다.

  • 긴 다듬기 과정을 통한 멀티플레이어 게임 생성 (Multiplayer game created through long refinement loop) 잉클링은 [자체적으로 생성한 코드와 결과물을 지속적으로 다듬어 멀티플레이어 게임을 완성했습니다...]

원문 보기
원문 보기 (영어)
Try on Tinker Model card Hugging Face Our mission is to build AI that extends human will and judgment. We have developed a platform that lets anyone customize models, previewed an AI system built for interactive collaboration, and published novel research . Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own. Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency. Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort. We trained it to be a broad, balanced foundation model: strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Inkling is just the start: our first release in a model family we will continue to build on. We want to make customization accessible for more use cases, so Inkling is available for fine-tuning on Tinker today. Picking the right base model to fine-tune is a qualitative judgment that combines measurable benchmarks with the unique feel of a model that comes from playing with it. To enable the latter we’re adding the Inkling Playground in the Tinker console: a developer-facing interface for chatting with Inkling. To show what customization means in practice, we asked Inkling to fine-tune itself. Using Tinker, the model wrote its own fine-tuning job, ran it, and evaluated the result: once a shareable cut exists. --> Capabilities Real-world applications require models with a wide range of capabilities that can be combined and improved with fine-tuning. We showcase what Inkling can do and how it measures up on important qualities such as trustworthiness and safety. Generalist model Inkling is designed to be broad. We trained it across agentic, reasoning, coding, instruction-following, factuality, vision, and audio tasks, rather than narrowly optimizing for one domain. That breadth matters for customization and real-world use: different users need models that can adapt to very different workflows, not just excel on benchmarks. Agentic coding and tool use A strong base for fine-tuning needs to flexibly solve a wide variety of tasks with agentic tool use. Inkling scores well among open-weights models on most agentic benchmarks. We trained Inkling to run inside a variety of coding and agent harnesses, and we randomized the tool set and schema during training to reduce sensitivity to any particular one. Inkling’s controllable thinking effort, described in the next section, can be set from within the harness. Below are a few demos showcasing Inkling’s agentic coding and tool use and the artifacts it creates. One-shot web app with embedded browser use Inkling builds a functional web app in a single shot, then powers an embedded AI assistant that can operate the web app interface through natural language instructions. Design Arena Inkling was evaluated on Design Arena’s Agentic Web Dev leaderboard, where blinded human evaluators compare generated web apps head to head. It ranks among the strongest open-weight models. Cohesively styled artifacts Inkling creates multi-page artifacts with precise instruction following, accurate information, and cohesive styling and design throughout. Multiplayer game created through long refinement loop Inkling refined an online snake game through 40 iterations of feedback from GPT Codex serving as a reviewer. The ability to sustain a long process of refinement and improve from feedback is crucial to creating the best collaborative work. Controllable thinking effort Test-time scaling and problem-solving are the core capability of every model, but that capacity is hard to capture with a single number. Developers fine-tuning models for a specialized task care as much about efficiency as about the max-effort performance on a public benchmark. Cost and latency are often binding constraints in real-world applications, and low latency in particular is crucial for enabling collaboration and improvement through iteration. Inkling supports controllable thinking effort, allowing you to balance performance with token efficiency. The chart above shows the effort/performance curve of Inkling as well as other open-weights models on a range of benchmarks: Terminal Bench 2.1 for agentic coding, HLE for advanced reasoning, and IFBench for instruction following. Inkling spends one third as many tokens to achieve the same performance as Nemotron 3 Ultra on Terminal Bench. Cost and latency matter for a model that you run millions of times and as part of longer workflows; looking at the full cost curve allows developers to choose the best model for each use case. Multimodality A major goal of Inkling’s design is to serve as the background reasoning model in the interaction models system we recently introduced. Interaction models enable the user to collaborate naturally, using voice and vision in real-time. This requires a model natively trained for broad multimodal capabilities. Open weights Closed weights Inkling effort=0.99 Qwen3-Omni Nemotron-3 Nano-Omni Kimi K2.5 Kimi K2.6 Qwen3.5 Omni-Plus Gemini 3.1 Pro (high) Audio Audio MC 56.6% 24.3% 23.2% – – 37.6% 66.8% MMAU 77.2% 77.5% 76.7% – – 81.1% 82.5% VoiceBench 91.4% 88.8% 89.4% – – 92.4% 94.3% Vision MMMU Pro (Standard 10) 73.5% 60.0% 53.0% 75.0% 79.0% 71.0% 82.0% Charxiv RQ 78.1% 61.1% 63.6% 77.5% 80.4% 72.5% 80.2% Charxiv RQ with python 82.0% – – 78.7% 86.7% – 89.9% Audio and vision benchmarks against specialist omni models (open- and closed-weight), reported at effort=0.99. The multimodal components are trained from scratch on general-domain data. We opted for an encoder-free architecture for audio and vision inputs, consistent with the interaction model design. Audio signals are input as discrete dMel spectrograms Bai et al., 2024 . , while images are encoded as patches of 40×40 pixels using a four-layer hMLP Touvron et al., 2022 . . Both are transformed via a light-weight embedding layer and processed jointly with text tokens. Inkling transcribes speech, follows spoken instructions, answers questions about recordings, and reasons over longer-form audio. These capabilities place it among the strongest open-weight audio models on VoiceBench, MMAU, and AudioMC. For vision, Inkling accepts images as input and can describe visual content, answer questions, and perform in-depth reasoning based on the provided visual information. It demonstrates strong performance on charts, diagrams, and mathematical visual reasoning tasks. During inference, Inkling can also leverage a Python tool to support image understanding through operations such as zooming and cropping, while seamlessly integrating visual reasoning with code-based reasoning. As our first release, Inkling establishes a robust multimodal foundation for future work. We expect its multimodal capabilities to continue improving as we expand the model and training pipeline in subsequent iterations. Epistemics We trained Inkling for calibration, instruction following, and resistance to censorship, which we refer to collectively as the model’s epistemics . Getting the facts right requires more than memorizing a large corpus of knowledge. A useful model must be well-calibrated, expressing the right amount of confidence in its answers — including on questions which a