메뉴
BL
The Decoder • 49일 전

바이트댄스, 10조 매개변수 중국 최대 AI 모델 개발 중

IMP
8/10
핵심 요약

틱톡의 모회사인 바이트댄스는 최대 10조 개의 매개변수(Parameter)를 갖춘 거대 AI 모델을 현재 사전 학습(Pretraining) 중입니다. 이는 기존 중국 최대 모델이었던 문샷(Moonshot)의 카미(Kimi)보다 3배 이상 큰 규모로, 글로벌 최고 수준의 AI 모델을 목표로 한다는 점에서 중요합니다. 창업자 장이밍의 지시에 따라 1년 넘게 타사 모델의 결과물을 활용하는 증류(Distillation) 방식을 배제하고 독자적인 역량 강화에 집중하고 있습니다.

번역된 본문

중국 최대 규모의 AI 모델이 바이트댄스에서 개발 중이다.

영국 파이낸셜 타임스(FT) 보도에 따르면, 바이트댄스는 최대 10조 개의 매개변수(parameters)를 보유한 AI 모델을 학습하고 있다. 이는 현재 중국에서 가장 큰 규모인 문샷(Moonshot)의 '카미 K3(Kimi K3)'보다 3배 더 큰 규모다. 이规模는 틱톡의 모회사인 바이트댄스를 업계 추정치 약 8조 매개변수 수준인 앤스로픽(Anthropic)의 최고 시스템 '미토스 5(Mythos 5)'와 맞먹는 위치에 올려놓을 것이다. 앤스로픽 측은 자체적인 정확한 수치를 공개한 바 없다.

세 명의 내부자들이 FT에 밝힌 바에 따르면, 해당 모델은 현재 사전 학습(pretraining) 단계에 있으며, 이 과정은 일반적으로 3~6개월 정도 소요된다. 매개변수는 모델이 저장할 수 있는 정보량을 결정하지만, 실제 성능은 데이터 품질과 학습 방법에도 크게 좌우된다.

소식원 중 한 명에 따르면, 바이트댄스는 1년이 넘는 기간 동안 증류(distillation), 즉 다른 기업 모델의 출력물을 바탕으로 학습하는 방식을 피해왔다. 창업자인 장이밍(张一鸣)은 2,000명 규모의 시드(Seed) 팀을 향해 장기적으로 세계 최고 수준의 모델 역량을 목표로 삼을 것을 내부적으로 지시한 바 있다.

일론 머스크에 따르면, xAI 역시 콜로서스 2(Colossus 2) 클러스터에서 6조 및 10조 매개변수를 갖춘 그록(Grok) 변형 모델을 학습하고 있는 것으로 알려졌다.

광고

과장 없는 AI 뉴스 – 전문가가 직접 엄선합니다. THE DECODER를 구독하고 광고 없는 환경에서 읽기, 주간 AI 뉴스레터, 연 6회 제공되는 독점 프론티어 보고서인 "AI 레이더", 전체 아카이브 열람 및 댓글 섹션 작성 혜택을 누려보세요. 지금 구독하세요.

출처: FT

원문 보기
원문 보기 (영어)
China's Largest AI Model Is Being Developed at Bytedance Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 7, 2026 Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That's three times the size of Moonshot's Kimi K3, currently the largest Chinese model. It would put the TikTok parent company in the same ballpark as Anthropic's top system Mythos 5 , which industry estimates place at around eight trillion parameters. Anthropic hasn't disclosed its own numbers. Three insiders told the FT that the model is in pretraining, a phase that typically takes three to six months. Parameters determine how much a model can store, but performance also depends on data quality and training methods. One of the sources says Bytedance has avoided distillation, meaning training on outputs from other companies' models, for over a year. Founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster , according to Elon Musk. Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: FT Ask about this article… Search