메뉴
HN
Hacker News 36일 전

VibeThinker-3B, 새로운 학습법으로 대형 모델 능가

IMP
9/10
핵심 요약

파라미터 30억(3B) 개의 소형 언어 모델인 VibeThinker-3B가 혁신적인 사후 학습 파이프라인을 통해 최신 대형 모델들을 능가하는 추론 성능을 달성했습니다. 단계별 지도 미세조정(SFT), 다중 도메인 강화학습, 오프라인 자가 증류를 결합해 수학 및 코딩 벤치마크에서 DeepSeek V3.2 등 대형 모델들과 동등하거나 더 뛰어난 결과를 보여주며, 작은 모델도 최적화된 학습법을 통해 최고 수준의 성능을 낼 수 있음을 증명했습니다.

번역된 본문

--> 컴퓨터 과학 > 인공지능 arXiv:2606.16140 (cs) [2026년 6월 15일 제출]

제목: VibeThinker-3B: 소형 언어 모델에서 검증 가능한 추론의 최전선 탐구 저자: Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou, Junlin Zhang

초록: 이 기술 보고서는 엄격한 소형 모델 환경 내에서 검증 가능한 추론 기능을 어디까지 끌어올릴 수 있는지 조사하기 위해 개발된 30억(3B) 파라미터의 컴팩트한 덴스 모델(dense model)인 VibeThinker-3B를 소개합니다. Spectrum-to-Signal 사후 학습(post-training) 패러다임을 기반으로, 우리는 커리큘럼 기반 지도 미세조정(SFT), 다중 도메인 강화학습(RL), 오프라인 자가 증류(self-distillation)를 포함하는 최적화된 파이프라인을 통해 모델을 체계적으로 향상시켰습니다.

실험 평가에 따르면 VibeThinker-3B는 까다로운 검증 가능 작업에서 최고 수준의 성능을 달성했습니다. 구체적으로, AIME26에서 94.3점(주장 수준 테스트 시간 스케일링을 적용하면 97.1점으로 향상됨)을 획득했고, LiveCodeBench v6에서 80.2%의 Pass@1을 기록했으며, 최근 공개된 보이지 않는 LeetCode 대회에서 96.1%의 수용률을 보여 강력한 분포 외(OOD) 일반화 능력을 입증했습니다. 이는 이 모델을 DeepSeek V3.2, GLM-5, Gemini 3 Pro와 같이 파라미터 규모가 수십 배 더 큰 플래그십 모델들과 일치하거나 능가하는 1티어 추론 시스템의 성능대역에 위치시킵니다. 또한, IFEval에서 93.4점을 기록한 것은 이러한 극한의 추론 능력 강화가 엄격한 명령어 통제 가능성을 훼손하지 않음을 확인해 줍니다.

이전의 1.5B 연구를 확장한 이러한 결과는 '파라미터 압축-적용범위 가설(Parametric Compression-Coverage Hypothesis)'을 도출하게 했습니다. 이 가설은 검증 가능한 추론 기능은 컴팩트한 '추론 코어(reasoning core)'로 압축될 수 있는 반면, 개방형 도메인 지식과 범용 역량은 사실, 개념 및 롱테일(Long-tail) 시나리오에 대한 폭넓은 파라미터 적용범위를 필요로 한다고 설명합니다. 이러한 관점은 컴팩트 모델이 단순히 배포 효율성을 위한 대체재가 아니라, 파라미터 밀집 역량 환경에서 최고 수준의 성능을 달성하기 위한 상호 보완적인 경로임을 시사합니다.

주제: 인공지능(cs.AI); 계산 및 언어(cs.CL) 인용: arXiv:2606.16140 [cs.AI] (또는 해당 버전의 경우 arXiv:2606.16140v1 [cs.AI]) https://doi.org/10.48550/arXiv.2606.16140

DataCite를 통해 발급된 arXiv DOI 제출 내역 보낸 사람: Sen Xu [이메일 보기] [v1] 2026년 6월 15일 월요일 02:57:19 UTC (552 KB) 전문 링크: 논문 접근: Sen Xu 및 8명의 다른 저자가 쓴 'VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models'라는 제목의 논문 PDF 보기 PDF HTML 보기(실험적) TeX 소스 라이선스 보기 현재 탐색 컨텍스트: cs.AI < 이전 | 다음 > 새 글 | 최근 글 | 2026-06 다음으로 탐색 변경: cs cs.CL

참고문헌 및 인용 NASA ADS Google 스칼라 Semantic Scholar 내보내기 BibTeX 인용 로딩 중... BibTeX 형식 인용 및 기타 로딩 중... 제공된 데이터: 책갈피 서지 도구 서지 및 인용 도구 서지 탐색기 토글 서지 탐색기(탐색기란?) Connected Papers 토글 Connected Papers(Connected Papers란?) Litmaps 토글 Litmaps(Litmaps란?) scite.ai 토글 scite 스마트 인용(스마트 인용이란?) 코드, 데이터, 미디어 이 기사와 관련된 코드, 데이터 및 미디어 alphaXiv 토글 alphaXiv(alphaXiv란?) 코드 링크 토글 논문용 CatalyzeX 코드 파인더(CatalyzeX란?) DagsHub 토글 DagsHub(DagsHub란?) GotitPub 토글 Gotit.pub(GotitPub이란?) Huggingface 토글 허깅 페이스(Hugging Face란?) ScienceCast 토글 ScienceCast(ScienceCast란?) 데모 데모 복제 토글 복제(복제란?) 스페이스 토글 허깅 페이스 스페이스(스페이스란?) 스페이스 토글 TXYZ.AI(TXYZ.AI란?) 관련 논문 추천 도구 및 검색 도구 인플루언스 플라워 링크 Influe

원문 보기
원문 보기 (영어)
--> Computer Science > Artificial Intelligence arXiv:2606.16140 (cs) [Submitted on 15 Jun 2026] Title: VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Authors: Sen Xu , Shixi Liu , Wei Wang , Jixin Min , Yingwei Dai , Zhibin Yin , Yirong Chen , Xin Zhou , Junlin Zhang View a PDF of the paper titled VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models, by Sen Xu and 8 other authors View PDF HTML (experimental) Abstract: This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectrum-to-Signal post-training paradigm, we systematically enhance the model through an optimized pipeline that includes curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation. Experimental evaluations demonstrate that VibeThinker-3B achieves frontier-level performance on highly demanding verifiable tasks. Specifically, it attains a score of 94.3 on AIME26 (improving to 97.1 with claim-level test-time scaling), an 80.2 Pass@1 on LiveCodeBench v6, and exhibits strong out-of-distribution generalization with a 96.1\% acceptance rate on recent unseen LeetCode contests. This effectively places it in the performance band of first-tier reasoning systems, matching or exceeding flagship models that are orders of magnitude larger, such as DeepSeek V3.2, GLM-5, and Gemini 3 Pro. Furthermore, a score of 93.4 on IFEval confirms that this extreme reasoning enhancement does not compromise strict instruction controllability. Extending our previous 1.5B work, these findings motivate the Parametric Compression-Coverage Hypothesis, which views verifiable reasoning as compressible into compact reasoning cores, while open-domain knowledge and general-purpose competence require broad parameter coverage over facts, concepts, and long-tail scenarios. This perspective suggests that compact models are not merely deployment-efficient substitutes, but a complementary path toward frontier-level performance in parameter-dense capability regimes. Subjects: Artificial Intelligence (cs.AI) ; Computation and Language (cs.CL) Cite as: arXiv:2606.16140 [cs.AI] (or arXiv:2606.16140v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2606.16140 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Sen Xu [ view email ] [v1] Mon, 15 Jun 2026 02:57:19 UTC (552 KB) Full-text links: Access Paper: View a PDF of the paper titled VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models, by Sen Xu and 8 other authors View PDF HTML (experimental) TeX Source view license Current browse context: cs.AI < prev | next > new | recent | 2026-06 Change to browse by: cs cs.CL References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation &times; loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) scite.ai Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle Gotit.pub ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle TXYZ.AI ( What is TXYZ.AI? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs . Which authors of this paper are endorsers? | Disable MathJax ( What is MathJax? )