메뉴
HN
Hacker News • 22일 전

K2 호라이즌: 여섯 개 오픈 모델로 구성된 연결 플릿 공개

IMP
8/10
핵심 요약

IFM이 0.9B부터 375B-A23B까지 여섯 개 모델로 구성된 'K2 Horizon'을 Apache 2.0 라이선스로 공개했습니다. 0.9B, 3.7B, 7B 모델은 각 크기급에서 새로운 SOTA를 달성했으며, 사전학습부터 에이전트 후행학습까지 전체 학습 과정(체크포인트, 데이터, 코드, 로그, 가중치)을 완전 개방한 것이 핵심입니다. 엣지 기기부터 기업 배포까지 전 구간을 커버하는 최초의 완전 오픈 에이전트 모델 패밀리라는 점에서 중요합니다.

번역된 본문

오늘 IFM은 여섯 개 모델로 구성된 연결된 플릿 'K2 Horizon'을 공개합니다: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, 0.9B입니다. 추론, 수학, 코딩, 에이전트 작업, 범용 능력 전반에서 K2 Horizon은 모든 크기급에서 최상위 성능을 제공하며, 0.9B, 3.7B, 7B 모델은 각각의 규모에서 새로운 SOTA(최고 수준)를 기록했습니다.

K2 Horizon은 우리의 지금까지 가장 포괄적인 오픈 릴리스이기도 합니다. 모든 모델에 대해 사전학습(pretraining)부터 추론 및 에이전트 후행학습(post-training)까지 전체 학습 라이프사이클을 개방합니다. 중간 체크포인트, 학습 데이터 또는 상세한 데이터 구성 레시피, 오픈 아키텍처, 믹스처 구성, 학습 코드, 설정, 세밀한 로그, 평가 결과, 최종 가중치를 모두 공개합니다. 모델과 코드는 Apache 2.0 라이선스로 배포되며, 데이터셋은 ODC-BY 등 해당 라이선스를 따릅니다. 재배포가 불가능한 경우에는 데이터가 어떻게 구성·혼합되었는지를 공개합니다.

규모를 아우르는 새로운 성능 프론티어. 0.9B, 3.7B, 7B 모델은 널리 쓰이는 평가에서 각 크기급에서 세계 최고 성능을 달성했습니다. 새로운 MoVA(Mixture-of-Value-Attention) 메커니즘을 탑재한 36B-A4B 모델은 활성 파라미터당 탁월한 능력을 발휘하며, 훨씬 큰 일부 모델을 능가합니다. 32B와 375B-A23B 모델도 각각의 비교군에서 최상위권에 속합니다. 여섯 모델은 엣지 기기부터 기업 환경까지 다양한 배포 환경에서 경쟁력 있는 성능을 제공합니다.

에이전트를 위한 최초의 완전 오픈 모델 플릿. K2 Horizon은 에이전트 후행학습까지 전체 개발 과정을 공개한 최초의 오픈 모델 패밀리입니다. 모든 단계의 체크포인트, 데이터(또는 데이터 레시피), 코드, 설정, 학습 로그를 공개함으로써 추론, 도구 사용, 계획, 에이전트 능력이 어떻게 발현되는지 연구하고, 이를 만드는 방법을 재현하며, 새로운 도구·환경·도메인에 적용하는 것을 가능하게 합니다.

엣지부터 기업까지 아우르는 여섯 모델. 0.9B 모델은 시계, 안경 같은 극도로 제약된 환경을 위해 설계되었고, 3.7B와 7B 모델은 스마트폰 등 온디바이스 애플리케이션에 고급 기능을 제공합니다. 밀집(dense) 32B 모델과 희소(sparse) 36B-A4B 모델은 로컬 워크스테이션과 효율적 서빙에 강력한 선택지를 제공하며, 375B-A23B 모델은 요구가 많은 기업 배포에 플릿 최고의 능력을 제공합니다. 여섯 모델 모두 양자화(quantization)를 지원합니다.

하나의 연결된 플릿. 여섯 모델은 핵심 아키텍처, 어휘, 학습 방법론, 인터페이스, 평가 인프라, 배포 도구를 공유합니다(0.9B 모델은 더 작은 어휘 사용). 이러한 일관성 덕분에 크기 간 이동, 동적 라우팅, 규모에 따른 능력·효율 연구가 쉬워집니다.

규모 전반의 세계 최고 성능. 0.9B, 3.7B, 7B 모델은 수학, 추론, 범용 능력, 코딩, 에이전트 작업에서 각 크기급 SOTA를 달성했습니다. 36B-A4B 모델은 활성 파라미터 수로 예상되는 수준을 넘어서는 성능을 보이며, 어텐션 값을 계산할 때의 독특한 MoE(Mixture-of-Expert) 설계의 효율성을 입증합니다. 32B와 375B-A23B 모델은 각 비교군 최상위권입니다.

소형 모델이 특히 주목할 만합니다. K2 Horizon 0.9B는 AIME 2026에서 48점 이상을 기록하며 강력한 추론, 도구 사용, 에이전트 능력을 갖췄습니다. 3.7B와 7B 모델은 이러한 능력을 더 까다로운 소프트웨어 엔지니어링 및 다단계 환경으로 확장하여 SWE-bench와 BrowseComp에서 우수한 성적을 보였습니다. 광범위한 탐색과 반복적 복구를 요구하는 TerminalBench 같은 복잡한 작업은 여전히 가장 작은 모델에게 어렵지만, K2 Horizon은 모든 규모에서 가능성의 경계를 확장합니다.

원문 보기
원문 보기 (영어)
Today IFM is releasing K2 Horizon, a connected fleet of six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B. Across reasoning, mathematics, coding, agentic tasks, and general capabilities, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B, and 7B models setting new state of the art at their respective scales. K2 Horizon is also our most comprehensive open release to date. For every model, we are opening the training lifecycle from pretraining through reasoning and agentic post-training. We are releasing intermediate checkpoints, training data or detailed data-construction recipes, open architecture, mixture compositions, training code, configurations, fine-grained logs, evaluation results, and final weights. The models and code are released under the Apache 2.0 license. Datasets are released under their applicable licenses, such as ODC-BY; We disclose how the data was constructed and mixed when redistribution is not possible. Together, K2 Horizon represents the most comprehensive open model release to date: A new performance frontier across scales. The 0.9B, 3.7B, and 7B models achieve world-leading performance in their size classes across widely used evaluations. The 36B-A4B model, equipped with our new Mixture-of-Value-Attention (MoVA) mechanism, delivers exceptional capability per active parameter, outperforming some much larger models. The 32B and 375B-A23B models rank among the top models in their respective classes. Together, the six models provide competitive performance across deployment environments ranging from edge devices to the enterprise. The first fully open model fleet for agents. K2 Horizon is the first open model family to expose the complete development process through agentic post-training. By releasing checkpoints, data (or data recipe), code, configurations, and training logs across every stage, K2 Horizon makes it possible to study how reasoning, tool use, planning, and agentic capabilities emerge; reproduce the methods that create them; and adapt those methods to new tools, environments, and domains. Six models spanning edge to enterprise. The 0.9B model is designed for highly constrained environments such as watches and glasses, while the 3.7B and 7B models bring advanced capabilities to phones and other on-device applications. The dense 32B model and sparse 36B-A4B model provide powerful options for local workstations and efficient serving. The 375B-A23B model brings the fleet’s strongest capabilities to demanding enterprise deployments. All six models include quantization support. One connected fleet. The six models share core architecture, vocabulary, training methodology, interfaces, evaluation infrastructure, and deployment tooling, with a smaller vocabulary for the 0.9B model. This consistency also makes it easier to move between sizes, route work dynamically, and study capability and efficiency across scale. World-leading performance across the scales The 0.9B, 3.7B, and 7B models achieve state-of-the-art results in their respective classes across mathematics, reasoning, general capability, coding, and agentic tasks. The 36B-A4B model performs beyond the level normally expected from its active parameter count, demonstrating the efficiency of our unique Mixture-of-Expert design when computing attention values. The 32B and 375B-A23B models place among the top models in their respective comparison classes. The small models are especially notable. K2 Horizon 0.9B achieves an AIME 2026 score above 48, along with strong reasoning, tool-use, and agentic capabilities. K2 Horizon 3.7B and 7B extend these capabilities to more demanding software-engineering and multi-step environments, demonstrated on strong performance in SWE-bench and BrowseComp. Although complex tasks that require extensive exploration and repeated recovery, such as those in TerminalBench, remain difficult for the smallest models, K2 Horizon moves the boundary of what is possible at every scale. Why the Horizon Fleet matters A transparent model that falls far behind the capability frontier has limited value as a foundation, even for research. At the same time, a powerful model released only as final weights allows people to run it, but provides little insight into how its capabilities were created. K2 Horizon brings these two together. The fleet provides highly competitive models and releases the recipes used to train them. Researchers can study advanced capabilities in models strong enough to exhibit them, while developers can reproduce, adapt, and extend the methods rather than treating the final checkpoint as an opaque starting point. Since introducing the fully open principle in our 2023 LLM360 paper , we have released open models every year while extending that commitment to larger scales, stronger capabilities, and now the complete lifecycle through agentic post-training. A Deep Dive into The K2 Horizon Fleet K2 Horizon 375B-A23B: the enterprise powerhouse K2 Horizon 375B-A23B is the fleet’s largest and most capable model. Its sparse MoE architecture provides 375 billion parameters of total capacity while activating approximately 23 billion parameters for each token, allowing it to draw on the capacity of a much larger model without using every parameter for every token. The model ranks among the top models below 400 billion parameters across general, reasoning, coding, and agentic evaluations. It is designed for demanding workloads where model quality matters most, including complex reasoning, software engineering, research, and long-horizon agentic tasks. Like every model in the Horizon fleet, 375B-A23B is released not as a single endpoint but as a development tree. Its intermediate checkpoints and post-training branches expose how the base model develops into reasoning, instruction-following, and specialized agentic variants. K2 Horizon 32B and 36B-A4B: strong performance for local deployment Horizon 32B is the fleet’s most powerful dense model, providing a strong balance of capability, adaptability, and local deployability. It ranks among the top dense models below 40 billion parameters. Horizon 36B-A4B reaches nearly the performance of the dense 32B model while activating only approximately 4 billion parameters per token. Its efficiency comes from MoVA, our new sparse attention architecture, together with MoE feed-forward layers. These two models serve as an important reference point for studying how dense and sparse architectures behave under similar training conditions. These models occupy the fleet’s local performance sweet spot. They are powerful enough for demanding reasoning, coding, and agentic applications while remaining practical for local workstations and efficient serving systems. K2 Horizon 7B, 3.7B, and 0.9B: frontier capability at small scale k2 Horizon 7B and 3.7B deliver strong reasoning, mathematics, coding, tool-use, and agentic performance while remaining suitable for local and on-device deployment. On several evaluations, their results approach or exceed those of models many times larger from the previous generation. K2 Horizon 0.9B carries many of the same capabilities into highly constrained environments. It can perform mathematical reasoning, use tools, and complete simple agentic tasks while remaining compact enough for applications on watches, glasses, and other edge devices under quantization. The appropriate task changes with scale: the 0.9B model is best suited to focused interactions and lightweight tool use, while the 3.7B and 7B models can handle more demanding coding and multi-step workflows. Together, they demonstrate how much capability can now be retained in models small enough to run almost anywhere. Designing K2 Horizon One family from the beginning Horizon was designed as a connected family rather than a collection of unrelated models. The six models share core architectural decisions, training methodology, interfaces, evaluation infrastructure, and deployment tooling. This