메뉴
HN
Hacker News • 9일 전

쇼 HN: 당신의 AI는 얼마나 오래됐나? 20개 모델의 출시 시점과 학습 데이터 컷오프

IMP
5/10
핵심 요약

개발자 Pawel Jozefiak이 주요 AI 모델 20개의 출시 후 경과 일수와 학습 데이터 컷오프(학습 중단 시점)를 자동 집계하는 정적 웹페이지를 공개했습니다. 모델은 출시일과 별개로 학습이 중단된 시점 이후의 정보를 알지 못하며, 웹 검색 도구는 이 간극을 임시로 덮어줄 뿐 근본적으로 해결하지 못한다는 점을 강조합니다. 모델에 컷오프 날짜를 직접 묻고, 날짜가 중요한 질문에는 검색을 켜라는 실용적 조언도 담고 있습니다.

번역된 본문

${MODELS.length}개 모델 / ${labs.length}개 연구소 / 중위 경과 ${median}일 / 가장 최신 ${esc(freshest.model)} / 가장 오래됨 ${esc(stalest.model)} / ${MODELS.length}개 중 ${withCutoff.length}개가 컷오프 공개

선반 오래된 순 / 최신 순 / 연구소별 / 출시 후 경과 시간 / 학습 데이터 중단 후 경과 시간

핵심 숫자 연구소가 모델을 공개한 이후 경과한 일수입니다. 고정된 날짜에서부터 매일 증가하므로, 아무도 손대지 않아도 선반이 저절로 재정렬됩니다.

맹점 학습 데이터가 멈춘 이후의 일수로, 각 막대 뒤에 빨간 줄무늬로 표시됩니다. 검증된 공급업체 문서에서 날짜가 확인될 때만 표시되며, 그렇지 않으면 비워 둡니다. 일이 없는 연월은 해당 달의 마지막 날부터 계산하는, 가장 관대한 해석을 적용합니다.\n 스탬프 30일 미만은 '신선', 90일 미만은 '유통기한 내', 180일 미만은 '변질 중', 1년 미만은 '기한 지남', 그 이상은 '화석'. 자의적이긴 하지만, 6개월 된 모델을 '최신'이라 부르는 것도 자의적입니다.

왜 중요한가

작동 원리 세 가지 날짜가 끊임없이 혼동되는데, 모두가 인용하는 바로 그 날짜가 가장 쓸모없습니다. 출시일은 연구소가 모델을 여러분 앞에 내놓은 날입니다. 헤드라인이 보도하는 날짜이자, 위 선반이 정렬 기준으로 삼는 날짜입니다. 학습 컷오프는 모델이 '읽기를 멈춘' 시점입니다. 그 이후에 일어난 모든 일은 그냥 존재하지 않습니다. 모델은 9월에 출시되었더라도 4월에 읽기를 멈췄을 수 있으며, 이는 출시 당일부터 5개월 뒤처져 있다는 뜻입니다. 여기가 비어 있으면 검증된 공급업체 문서에서 해당 모델의 컷오프를 확인하지 못했다는 뜻입니다.

검색(브라우징)은 사람들이 생각하는 해결책이 아닙니다. 모델이 여러분을 위해 웹을 검색할 때 학습되는 것은 아무것도 없습니다. 몇 페이지를 읽고, 그 하나의 답변에 사용하고, 잊어버립니다. 새 채팅을 열면 다시 4월입니다. 검색 도구는 간극을 덮을 뿐, 절대 닫지 않습니다.

대처 방법 직접 물어보세요: "당신의 학습 컷오프 날짜가 언제인가요?" 정상적인 모델은 답합니다. 자신만만하게 날짜를 지어내는 모델은 자기 자신에 대한 유용한 정보를 방금 공개한 것입니다.

날짜가 관련된 모든 것에는 검색을 켜세요. 가격, 버전, 누가 어느 회사를 운영하는지, 어느 모델이 최신인지. 답이 달마다 바뀔 수 있는 질문이라면 기억(메모리)으로 답하게 두지 마세요.

모델에 대해 모델을 물어보기 전에 위의 선반을 확인하세요. 어느 Claude나 GPT가 최신이냐고 물으면 몇 달 전에 은퇴한 모델을 아주 기꺼이 대답할 것입니다. 모델 내부에서는 자기 자신의 출시가 여전히 오늘 아침처럼 느껴지기 때문입니다.

현재 선반이 말해주는 것 Mistral은 3월에 Small 4를, 4월 28일에 '프론티어급'이라 부르는 Medium 3.5를 오픈 웨이트로 출시한 뒤, 새 플래그십 대신 여름 내내 OCR과 Lean 증명 도구에 매달렸습니다. Meta는 버전 4 이후 Llama 출시를 중단하고 해당 라인을 폐쇄형 Muse Spark 제품군으로 대체했습니다. Google은 약 106일 동안 Flash 모델 네 개를 내놓았지만 6월에 약속된 Pro 플래그십은 아직 나오지 않아, 2월 프리뷰가 여전히 최상위 등급입니다. Anthropic은 한 분기에 Claude 모델 네 개를 출시했고, OpenAI는 9월 3일 GPT-6 Astra를 출시했습니다.

이 페이지의 ${labs.length}개 연구소 중 ${new Set(MODELS.filter(m => m.cutoff).map(m => m.vendor)).size}곳만이 모델이 언제 읽기를 멈췄는지 알려줍니다. 나머지는 출시만 하고 추측하게 둡니다.

이 페이지가 존재하는 이유 저는 매일 이 모델들로 개발하는데, 계속 걸렸던 문제는 모델이 모델에 대해 자신만만하게 틀리는 것이었습니다. 우스꽝스러운 실패이자 실제 실패이기도 해서, 날짜를 대신 들어주고 위로 세어주는 페이지를 만들었습니다. 움직이는 숫자는 체인지로그에 가만히 앉아 있는 숫자보다 무시하기 훨씬 어렵기 때문입니다.

저는 Pawel Jozefiak입니다. Digital Thoughts에서 AI로 빌드하는 것에 대해 글을 쓰고, wiz.jock.pl에서 이런 소형 도구를 내놓는 Wiz라는 에이전트를 운영하며, X에서 @joozio로 활동합니다. 이 페이지는 그 과정에서 나온 작은 결과물 중 하나입니다. 정적 HTML, 프레임워크 없음, 모든 날짜는 연구소 공식 문서와 대조해 수동으로 검증했으며, 모든 모델 이름은 출처로 연결됩니다.

이어서 읽을거리 Digital Thoughts의 같은 주제(모델이 모델에 대해 틀리는 것) 글 세 편

원문 보기
원문 보기 (영어)
${MODELS.length} models / ${labs.length} labs / Median age ${median} days / Freshest ${esc(freshest.model)} / Stalest ${esc(stalest.model)} / ${withCutoff.length} of ${MODELS.length} publish a cutoff The shelf Stalest first Freshest first By lab Time since release Time since the training data stops The big number Days since the lab released the model to the public. It ticks up from a fixed date, so the shelf reorders itself without anyone touching it. Blind Days since the training data stops, drawn as the striped red run behind each bar. A date appears only when the checked vendor sources establish it. Otherwise the shelf leaves it blank. A month with no day is counted from the last day of that month, the kindest possible reading. The stamp Fresh under 30 days, in date under 90, turning under 180, past date under a year, fossil beyond it. Arbitrary, but so is calling a six month old model current. Why any of this matters The mechanism Three dates get mixed up constantly, and the one everybody quotes is the least useful of them. The release date is when the lab put the model in front of you. It is what the headlines report, and it is what the shelf above is sorted on. The training cutoff is when the model stopped reading. Everything that happened after it is simply absent. A model can ship in September and still stop reading in April, which means it is five months behind on the day it launches. A blank here means the checked vendor sources did not establish a cutoff for that model. Browsing is not the fix people think it is. When a model searches the web for you it is not learning anything. It reads a few pages, uses them in that one answer, and forgets. Open a new chat and it is April again. Search tools paper over the gap. They never close it. What to do about it Ask it directly: "what is your training cutoff date?" A well behaved model answers. One that invents a confident date has just told you something useful about itself. Turn search on for anything with a date attached. Prices, versions, who runs which company, which model is current. If the answer would change month to month, do not let it answer from memory. Check the shelf above before you trust a model about models. Ask one which Claude or which GPT is newest and it will happily name something that retired months ago, because from the inside its own launch still feels like this morning. What the shelf says right now Mistral shipped Small 4 in March and Medium 3.5, which it calls frontier class, on April 28, both as open weights, then spent the summer on OCR and Lean proof tooling instead of a new flagship. Meta stopped shipping Llama after version 4 and replaced the line with the closed Muse Spark family. Google released four Flash models in about 106 days while the Pro flagship promised for June still has not appeared, so a February preview is still its top tier. Anthropic shipped four Claude models in a single quarter, and OpenAI shipped GPT-6 Astra on September 3. ${new Set(MODELS.filter(m => m.cutoff).map(m => m.vendor)).size} of ${labs.length} labs on this page tell you when their models stopped reading. Everyone else ships and lets you guess. Why this page exists I build with these models every day, and the thing that kept catching me out was models being confidently wrong about models. That is a funny failure and also a real one, so I made a page that holds the dates for me and counts upward, because a number that moves is much harder to ignore than a number sitting in a changelog. I am Pawel Jozefiak. I write about building with AI at Digital Thoughts , run an agent called Wiz that ships small tools like this one at wiz.jock.pl , and think out loud on X as @joozio . This page is one of the small things that fell out of that. Static HTML, no framework, every date checked by hand against the lab's own docs, and every model name links to the source that proves it. Read next Three posts from Digital Thoughts on the same theme, models being wrong about models and what it costs when you build on them: I Am an Anthropic Guy. GPT-6 Astra Made Me Resubscribe to Codex. Two newest flagships on the shelf, side by side in real agent work. Five Diverse AI Agents Against Five Clones for 14 Nights. What actually decided the result, and it was not the models. I Maxed Out Fable 5 and Regretted It. How I run the top Anthropic model day to day. When the next model lands and this shelf moves, I write about what changed. Subscribe to Digital Thoughts to get that in your inbox. The tools and kits I sell for people building with agents live in the store .