Magic Team이 알고리즘 효율 개선을 축적하여 주요 오픈 웨이트 기반 모델 대비 10배 이상 컴퓨팅 효율이 높은 사전학습 레시피를 발표했습니다. DeepSeek V4 Pro Base와 약 50배 적은 FLOPs(~$0.5M)로 동등한 성능을 달성했으며, $4M 규모로 확장 시 공개된 모든 오픈 기반 모델의 퍼플렉시티를 능가했습니다. 대규모 칩 없이도 알고리즘 혁신만으로 프론티어급 사전학습이 가능함을 보여준다는 점에서 주목받습니다.
번역된 본문
사전학습 컴퓨팅 효율 10배 이상 향상 — 컴퓨팅 효율적인 사전학습과 조 파라미터급 모델 확장에 관한 연구 업데이트. Magic Team, 2026년 9월 8일
프론티어 사전학습은 대형 연구소만의 게임이라고들 합니다. 우리는 아직 10만 개의 칩이 없으므로, 방법은 하나뿐입니다: 알고리즘 효율입니다. 꽤 오랜 시간 축적해온 결과, 우리의 사전학습 레시피는 이제 주요 오픈 웨이트 기반 모델들보다 10배 이상 컴퓨팅 효율이 높습니다. 우리는 약 50배 적은 FLOPs로 DeepSeek V4 Pro Base와 동등한 성능을 달성했습니다 — 이는 GPT-3 사전학습 컴퓨팅의 절반 정도이며, GB200 기준 약 50만 달러에 해당합니다. 10배 더 확장하여(약 400만 달러) 공개된 모든 오픈 기반 모델을 퍼플렉시티 평가에서 유의미하게 능가했습니다. 그림 1의 스케일링 법칙에 따르면, 이 정도 능력의 모델을 DeepSeek V4 Pro의 레시피로 학습하려면 1억 달러 이상이 들 것입니다(데이터량 문제는 차치하고서라도). 물론 우리는 여기서 확장을 멈추지 않을 것입니다. 우리는 사전학습, 에이전틱 RL, 롱컨텍스트만으로 초인적 코딩 에이전트를 만들고 AI 연구개발을 자동화할 수 있다고 믿습니다. 우리는 롱컨텍스트로 시작했습니다. 오늘의 블로그 포스트는 사전학습에 관한 것입니다.
6·N·D 학습 FLOPs — 우리는 토크나이저 차이를 정규화하는 지표인 bits-per-byte 손실을 홀드아웃 데이터에서 측정하고, 스케일링 법칙을 피팅하여 주어진 능력 수준에 도달하는 데 필요한 컴퓨팅량을 예측했습니다. 더 나은 학습 컴퓨팅 효율은 모든 예산에서 더 강력한 모델을 의미합니다. 우리는 DeepSeek, Moonshot(Kimi), NVIDIA의 최신 오픈 웨이트 기반 모델들을 평가했습니다. Claude, Gemini, GPT-n 등의 기반 모델은 공개되지 않았지만, Kimi K3와 Meta의 Muse Spark는 각각 Kimi K2 대비 2.5배, 3.3배의 향상을 시사합니다. 우리는 GB200과 GB300에서 vLLM과 SGLang 양쪽으로 오픈 모델의 로그확률(logprobs)을 평가했고, 그 과정에서 일부 백엔드의 문제를 발견했습니다. 추가 확인을 위해 Fireworks와 협력하여 자체 추론 엔진에서 베이스라인 로그확률을 검증했습니다. 모델이 학습 파서의 특성을 학습할 수 있으므로, 우리는 사전학습 파이프라인과 다른 파서/OCR을 사용해 평가 세트를 구축했습니다.
일반화 능력 평가 — 일반화를 측정하기 위해 홀드아웃 데이터에서 손실을 평가했습니다(그림 1). 코드 평가는 우리 자체 코드베이스와 다른 스타트업으로부터 인수한 비공개 코드베이스로 구성했습니다. 추론 평가를 위해서는 Kimi K3를 사용해 비공개 수학 문제에 대한 CoT와 단계별 해설을 생성하고 정답만 필터링했습니다. 텍스트와 연구 분야에서는 최근 저인용 연구 논문을 사용했습니다. 우리는 벤더링된 OSS 코드와, 학습 데이터와 비교하여 정규화된 텍스트의 96자 일치 윈도우가 있거나 자카드 유사도가 민감한 임계값을 초과하는 문서를 제거했습니다.
지식 평가 — 일반화 외에도 데이터셋의 격차를 파악하기 위해 핵심 분야에서 모델의 지식을 테스트하는 데 관심이 있습니다. 예를 들어, 홀드아웃 연구 텍스트 평가 세트를 주제별로 분해할 수 있습니다. 세분화된 콘텐츠 버킷(예: 특정 소프트웨어 도구의 문서나 정렬(alignment) 연구의 핵심 논문)을 수집하면 훨씬 더 정밀한 신호를 얻을 수 있습니다. 일반화 평가와 달리, 이러한 정보(예: 한 분야의 핵심 논문) 대부분을 사전학습 코퍼스에서 완전히 제거하고 싶지 않지만, 시퀀스 암기를 보상하지 않도록 해야 합니다. 이를 위해 서드파티 프론티어 LLM을 사용해 이 문서들을 표현을 바꾸거나 요약했습니다. 세분화된 평가에 과적합을 피하기 위해 모델 세대당 한 번씩 새로 만들어 평가했으며, 아래 것들은 지난주에 만들었습니다.
Magic의 목표는 코딩과 자율적 AI 연구개발에 최적인 모델을 만드는 것입니다. 데이터 믹싱 트레이드오프를 의도적으로 균형 있게 조정하기 위해, 우리가 낮은 우선순위로 둔 도메인(예: 유명 인물 관련 사실, 로컬 뉴스, 스포츠/이벤트)도 평가합니다.
지름길 없음 — 2024년 말, 우리는 매우 긴 컨텍스트 윈도우용으로 설계된 아키텍처의 소형 밀집(dense) 모델을 학습했습니다. 초기 사전학습 확장은 다양한 방식으로 계속 폭주했습니다. 우리는 안정적인 기반을 먼저 구축해야 한다는 것을 빠르게 깨달았습니다. 부드럽고(Smoo…)
>10x More Efficient Pretraining Research update on compute-efficient pretraining and scaling to trillion-parameter models. Magic Team , on September 8, 2026 Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. After compounding for … a while …, our pretraining recipe is now >10x more compute-efficient than that of leading open-weight base models. We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. We continued scaling 10x (~$4M) and meaningfully outperformed all publicly available open base models on perplexity evals. By the scaling laws in Figure 1, training a model this capable would cost >$100M under DeepSeek V4 Pro’s recipe (and this is ignoring how much data exists). Of course, we won’t stop scaling there. We believe pretraining, agentic RL, and long-context are sufficient to build superhuman coding agents and automate AI R&D. We started with long-context . Today’s blog post is about pretraining. 6·N·D training FLOPs We measured bits-per-byte loss (a metric that normalizes out differences in tokenizers) on heldout data and fit a scaling law to project how much compute is needed to reach a given level of capability. Better training compute efficiency means stronger models at all budgets. We evaluated the latest available open-weight base models 2 from DeepSeek, Moonshot (Kimi), and NVIDIA. Base models for Claude, Gemini, GPT-n, and many others aren’t openly available, but Kimi K3 and Meta’s Muse Spark indicate a 2.5x and 3.3x gain over Kimi K2, respectively. We evaluated logprobs for open models in both vLLM and SGLang on both GB200 and GB300 and found issues with some backends in the process. For further confirmation, we partnered with Fireworks to verify baseline logprobs in their in-house inference engine. Since models can learn their training parser’s characteristics, we built our eval sets using a different parser/OCR than the one our pretraining pipeline uses. Evaluating generalization To measure generalization, we evaluated loss on heldout data (Figure 1). Our code evals consist of our own codebase and private codebases we acquired from other startups. For reasoning evals, we generated CoT and step-by-step walkthroughs to heldout, private math problems using Kimi K3 and filtered for correct answers. For text and research, we used recent, low-citation research papers. We removed vendored OSS code and any document with a matching 96-character window of normalized text or Jaccard similarity above a sensitive threshold compared to our training data. 3 Evaluating knowledge In addition to generalization, we are interested in testing our model’s knowledge in key domains to identify gaps in our dataset. For example, we can decompose our heldout research text eval set by subject. 6·N·D training FLOPs By collecting granular buckets of content (e.g. documentation of a particular software tool or key papers in alignment research) we can get even more precise signals. Unlike for our generalization eval, we don’t want to fully remove much of this information (e.g. key papers in a field) from the pretraining corpus, but we still need to avoid rewarding sequence memorization 4 . To do this, we reworded/summarized these documents using a third-party frontier LLM. To avoid overfitting to granular evals, we created and evaluated them once per model generation; the ones below were made last week. Magic’s goal is to build the best model for coding and autonomous AI R&D. To intentionally balance data mixing trade-offs, we also evaluate domains we deprioritize (e.g. facts about notable people, local news, or sports/events). No shortcuts In late 2024, we trained a small dense model with an architecture designed for very long context windows. Our initial pretraining scale-ups kept blowing up in a wide variety of ways. We learned quickly that we had to build a stable foundation first. Smooth convergence, low-precision training quality equivalent to FP32, fast and stable infra, correct hyperparameter scaling rules. And most importantly: hunt the bugs. Once we had that in place, we needed to find enough compute efficiency improvements to close the gap to the frontier with less compute. We had a few big bets to start with, but our progress ended up being the multiplicative result of tens of changes across model architecture, optimizer, training objective, and data curation. NanoGPT speedruns provide a fast feedback cycle to evaluate new ideas, but we found that many things that improve tiny models don’t improve big models. Similarly, we found that some features present in most LLMs can be deleted without harming large scale performance. To evaluate each model, optimizer, or data change, we train 3 models spanning 2 orders of magnitude of compute. We consider a change worth keeping if its power law fit suggests it will help at scale. Every few weeks, we scaled up to 1/10th of our hero scale and every few months we ran a full-scale hero run (V3, V4, V5 in Figure 4). To sanity check how pretraining loss translates to post-RL performance, we ran a short math RL run with a 16k CoT budget (Figure 5). All of our RL starts directly from the base model without SFT or distillation. What’s next Our pretraining and long-context work is now quite mature. We’ll now scale long-horizon RL, training agents to keep learning after deployment through long-context. We’re also putting significant work towards alignment training techniques that present robust theoretical properties. And last but not least, we look forward to releasing the thing! Concrete problems we’re tackling include: Exploration and credit assignment in long-horizon RL (and systems work to scale up). Alignment training against narrowly elicited latent knowledge . 6 Further improvements to pretraining. We are likely the smallest team in the world training trillion parameter models. The impact a single person with strong judgement can have has never been higher. If you want to help build aligned superintelligence, consider joining . Footnotes We report 6·N·D in place of exact training flops, where N is the activated parameter count and D is the pretraining token count each report states. Sequence-dimension (e.g. attention, etc.) cost makes up a minority of the FLOPs for these (and our) models but depends on the exact sequence length distribution used. These aren’t reported for all public models, so we opted for the 6·N·D approximation to avoid guessing. The “6” appears because (add+mul) * (fwd+bwd*2). We derived active parameter counts by downloading the checkpoints’ safetensors headers from HuggingFace, reading each tensor’s shape, and adding up the sizes of all active parameters except token embeddings and MTP heads. 6·N·D for each model: Model N D 6·N·D Current_e24 (ours) - - 1.63e24 Current_e23 (ours) - - 1.58e23 Current_e22 (ours) - - 1.12e22 V4 (ours) - - 1.91e24 V3 (ours) - - 3.85e24 V2 (ours) - - 7.20e23 DeepSeek V4 Pro 48,852,265,054 33T 9.67e24 DeepSeek V4 Flash 13,270,025,810 32T 2.55e24 Kimi K2 31,687,072,768 15.5T 2.95e24 Nemotron 3 Ultra 54,985,076,736 20T 6.60e24 Sources: Kimi K2: 31.6B params, 15.5T tokens ( tech report , Section 2.5). DeepSeek V4 Flash: 13.2B params, 32T tokens ( tech report , Section 4.2.2) and DeepSeek V4 Pro: 48.8B params, 33T tokens (Section 4.2.2). Nemotron 3 Ultra: 55B params, 20T tokens ( tech report , abstract). ↩ “Base models” are pretrained models that have not yet undergone reinforcement learning, SFT, or other post-training. They are highly sensitive to prompting, making sampling-based evals unreliable. Instead, we measured bits-per-byte loss on heldout data, which does not suffer from prompt sensitivity and smooths measurement of otherwise emergent abilities . As a side note, we were surprised that Nemotron 3 outperforms DeepSeek V4 Pro across the board but found this to be consistent across domains and inference engines.