메뉴
HN
Hacker News • 36일 전

DiffusionGemma 기술 보고서

IMP
8/10
핵심 요약

Gemma 팀이 기존 자기회귀(AR) 모델을 디퓨전 방식으로 파인튜닝한 실험적 오픈 웨이트 언어 모델 'DiffusionGemma'를 공개했습니다. 이 모델은 256 토큰 블록을 병렬로 반복 정제하여 H100 GPU 한 장에서 초당 약 1,500 토큰을 생성하며, 최첨단 추측 디코딩을 적용한 AR 모델보다도 훨씬 빠릅니다. 원본 AR 모델 학습 토큰 예산의 10% 미만으로 학습되었으며, AR 생성 능력도 유지해 하이브리드 디퓨전-AR 디코딩의 가능성을 보여줍니다.

번역된 본문

--> 컴퓨터 과학 > 계산 및 언어

arXiv:2608.00146 (cs) [2026년 7월 31일 제출]

제목: DiffusionGemma 기술 보고서

저자: DiffusionGemma 팀: Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor 외

초록:

우리는 이산 디퓨전(discrete diffusion)을 사용해 매우 빠른 속도로 텍스트를 생성하는 실험적 오픈 웨이트 언어 모델인 DiffusionGemma를 소개합니다. 한 번에 토큰 하나씩 디코딩하는 대신, DiffusionGemma는 256개 토큰 블록을 병렬로 반복적으로 정제함으로써 기존 자기회귀(AR) 대규모 언어 모델의 순차적 디코딩 병목을 회피합니다.

처음부터 학습하는 대신, 우리는 활성 파라미터 38억 개, 총 파라미터 252억 개의 전문가 혼합(MoE) 모델인 Gemma 4를 파인튜닝하여 DiffusionGemma를 얻었습니다. 컴퓨팅 효율적인 2단계 학습 파이프라인은 출발 AR 모델의 전체 학습 토큰 예산의 10% 미만을 사용합니다. 첫 번째 단계는 지도 파인튜닝(SFT)으로 양방향 디노이징(denoising)을 학습시키고, 두 번째 단계는 강화학습과 샘플러 증류(sampler distillation)를 결합하여 생성 품질과 추론 효율을 함께 개선합니다.

DiffusionGemma는 생성 속도와 모델 능력 간의 트레이드오프에서 새로운 파레토 프론티어를 수립합니다. 전체 평가 스위트를 평균하여 순방향 패스당 약 20개의 토큰을 생성하며, 단일 NVIDIA H100 GPU에서 초당 약 1,500개의 출력 토큰을 달성합니다. 이는 최첨단 추측 디코딩(speculative decoding)을 적용한 AR 모델보다도 상당히 빠른 속도입니다.

또한 DiffusionGemma는 원본 모델의 사고 모드(thinking mode), 멀티모달 입력, 긴 컨텍스트 지원을 유지합니다. 디퓨전 파인튜닝에도 불구하고 미미한 성능 저하만으로 AR 생성이 여전히 가능하며, 이는 하이브리드 디퓨전-AR 디코딩으로 가는 길을 시사합니다.

주제: 계산 및 언어 (cs.CL); 인공지능 (cs.AI)

인용: arXiv:2608.00146 [cs.CL] https://doi.org/10.48550/arXiv.2608.00146

제출 이력: 작성자: Jean Tarbouriech [v1] 2026년 7월 31일 (금) 16:11:46 UTC (6,116 KB)

전문 링크: PDF 보기, HTML 보기(실험적), TeX 소스, 라이선스 보기

현재 탐색 컨텍스트: cs.CL

참고문헌 및 인용: NASA ADS, Google Scholar, Semantic Scholar, BibTeX 내보내기

서지 도구: Bibliographic Explorer, Connected Papers, Litmaps, scite.ai 스마트 인용

코드, 데이터, 미디어: alphaXiv, CatalyzeX 코드 파인더, DagsHub

원문 보기
원문 보기 (영어)
--> Computer Science > Computation and Language arXiv:2608.00146 (cs) [Submitted on 31 Jul 2026] Title: DiffusionGemma Technical Report Authors: DiffusionGemma Team : Adrien Ali Taïga , James Assiene , Daniele Calandriello , Rahma Chaabouni , João Gante , Tamara von Glehn , Nate Keating , Chris Knutsen , Martin Kukla , Tianlin Liu , Ivan Lobov , Ofir Nabati , João Gabriel Oliveira , Nicolas Perez-Nieves , Nastasia Prutianova , Bobak Shahriari , Jean Tarbouriech , Pavel Tyletski , Çağlar Ünlü , Cindy Wu , Glenn Cameron , Jerome Connor , Sertan Girgin , Maarten Grootendorst , Alon Levkovitch , Eliya Nachmani , Omar Sanseviero , Piotr Stanczyk , Quentin Berthet , Andrew Campbell , Clément Crepy , Valentin De Bortoli , Arnaud Doucet , Romuald Elie , Alexandre Galashov , Klaus Greff , Alexis Jacq , David Ruhe , Yu-Han Wu , Sebastian Flennerhag , Brendan O'Donoghue , George Scrivener , Shantanu Thakoor View a PDF of the paper titled DiffusionGemma Technical Report, by DiffusionGemma Team: Adrien Ali Ta\&#34;iga and 42 other authors View PDF HTML (experimental) Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding. Subjects: Computation and Language (cs.CL) ; Artificial Intelligence (cs.AI) Cite as: arXiv:2608.00146 [cs.CL] (or arXiv:2608.00146v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2608.00146 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Jean Tarbouriech [ view email ] [v1] Fri, 31 Jul 2026 16:11:46 UTC (6,116 KB) Full-text links: Access Paper: View a PDF of the paper titled DiffusionGemma Technical Report, by DiffusionGemma Team: Adrien Ali Ta\&#34;iga and 42 other authors View PDF HTML (experimental) TeX Source view license Current browse context: cs.CL < prev | next > new | recent | 2026-08 Change to browse by: cs cs.AI References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation &times; loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) scite.ai Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle Gotit.pub ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle TXYZ.AI ( What is TXYZ.AI? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs . Which authors of this paper are endorsers? | Disable MathJax ( What is MathJax? )