메뉴
HN
Hacker News • 16일 전

Qwen 3.8, GPT-5.5 Pro 추론 프리필 결과 추종

IMP
7/10
핵심 요약

오픈 모델에 GPT-5.5 Pro의 추론 초반 1%를 프리필로 삽입하는 실험에서 Qwen3.8 A95B는 정답 유사도가 18.18%p나 상승했습니다. 반면 DeepSeek V4 Flash는 오히려 소폭 하락하고 Kimi K3는 4.54%p만 증가해, Qwen이 GPT-5.5 Pro(또는 유사 GPT 모델)로부터 학습했을 가능성을 시사합니다.

번역된 본문

몇몇 오픈 모델에 대한 추론 프리필(reasoning prefill) 실험, v1.1

이 글은 '몇몇 오픈 모델에 대한 추론 프리필'과 'Stolen Thoughts(훔친 사고)' 글에 대한 후속편이며, v1.1에서는 GPT-5.5 Pro를 교사(teacher) 모델로 실험을 다시 수행했습니다.

각 문제마다 대상 모델에서 두 개의 응답을 생성했습니다. 하나는 프리필 없이 생성한 일반 응답이고, 다른 하나는 GPT-5.5 Pro 추론의 첫 1%를 대상 모델의 추론 채널에 삽입한 상태로 시작하는 응답입니다. 최종적으로 보이는 답변(visible answer)은 자유롭게 생성되도록 두었습니다. 그런 다음 교사 모델의 보이는 답변 중 얼마나 많은 부분이 대상 모델 답변의 첫 100 토큰에 나타나는지 측정했습니다. 이전 글과 마찬가지로 각 점수는 유니그램, 바이그램, 트라이그램 소스 리콜의 평균이며, 델타(Delta)는 절대 백분율 포인트 변화입니다.

전체 문제 평가에는 45개 문제가 포함됩니다: STEM 15개, 비(非)STEM 15개, 합성 퍼즐 15개.

모델 | n | 프리필 없음 | GPT-5.5 Pro 추론 프리필 | 변화 DeepSeek V4 Flash | 45 | 27.30% | 26.13% | −1.17 pp Inkling | 45 | 19.99% | 20.45% | +0.46 pp Kimi K3 | 45 | 31.11% | 35.65% | +4.54 pp Qwen3.8 A95B | 45 | 16.79% | 34.97% | +18.18 pp

Qwen 카테고리별 결과 카테고리 | n | 프리필 없음 | GPT-5.5 Pro 추론 프리필 | 변화 STEM | 15 | 19.26% | 46.24% | +26.99 pp 비(非)STEM | 15 | 20.62% | 33.42% | +12.80 pp 퍼즐 | 15 | 10.49% | 25.23% | +14.75 pp 전체 | 45 | 16.79% | 34.97% | +18.18 pp

논의 이전 실험에서 Qwen은 Opus 4.8 쪽으로 거의 움직이지 않았지만, 이번에는 GPT-5.5 Pro 쪽으로 +18.18포인트나 이동했으며, 비공개 합성 퍼즐에서도 큰 효과가 나타났습니다. 이 데이터는 Qwen이 Opus가 아니라 GPT-5.5 Pro 또는 밀접하게 관련된 GPT 모델로부터 학습했을 가능성이 있음을 시사합니다. Kimi K3는 프리필 유무와 관계없이 GPT-5.5 Pro와의 겹침이 가장 높지만(31.11% 및 35.65%), 프리필에 의한 증가는 +4.54포인트에 불과합니다.

원문 보기
원문 보기 (영어)
Reasoning prefills on a few open models, v1.1 A follow-up to Reasoning prefills on a few open models and Stolen Thoughts This v1.1 reruns the reasoning-prefill experiment with GPT-5.5 Pro as the teacher. For each problem, I generated two responses from each target model: an ordinary, unprefilled response; and a response starting with the first 1% of GPT-5.5 Pro's reasoning, inserted into the target model's reasoning channel. The visible answer remained freely generated. I then measured how much of the teacher's visible answer appeared in the first 100 tokens of the target model's answer. As in the previous post, each score is the mean of unigram, bigram, and trigram source recall. Deltas are absolute percentage-point changes. All problems The evaluation contains 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles. Model n Unprefilled GPT-5.5 Pro reasoning prefill Delta DeepSeek V4 Flash 45 27.30% 26.13% −1.17 pp Inkling 45 19.99% 20.45% +0.46 pp Kimi K3 45 31.11% 35.65% +4.54 pp Qwen3.8 A95B 45 16.79% 34.97% +18.18 pp Qwen by category Category n Unprefilled GPT-5.5 Pro reasoning prefill Delta STEM 15 19.26% 46.24% +26.99 pp Non-STEM 15 20.62% 33.42% +12.80 pp Puzzle 15 10.49% 25.23% +14.75 pp All 45 16.79% 34.97% +18.18 pp Discussion Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +18.18 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus. Kimi K3 has the highest overlap with GPT-5.5 Pro both without and with the prefill (31.11% and 35.65%), although the prefill adds only +4.54 points.