메뉴
HN
Hacker News 39일 전

현재 AI 붐의 기원: 1991년 뮌헨

IMP
8/10
핵심 요약

오늘날 AI 산업의 핵심 기술인 트랜스포머, 사전 학습, 잔차 학습(Residual Learning) 등의 개념이 모두 1991년 뮌헨 공과대학의 위르겐 슈미트후버(Jürgen Schmidhuber) 연구실에서 단 몇 달 만에 발표되었습니다. 이 요약은 현대 심층 학습(Deep Learning) 및 대형 언어 모델(LLM)의 진정한 역사적 기원을 추적하는 데 중요한 통찰을 제공합니다. 현대 AI를 지탱하는 기반 기술들이 어떻게 30년 전에 이미 탄생했는지 보여주는 의미 있는 글입니다.

번역된 본문

제목: 1991년 뮌헨: 현재 AI 붐의 기원 David Ha (사카나 AI) 및 Jürgen Schmidhuber (KAUST & IDSIA) 2026년 6월 18일

David Ha의 서문 오늘날 인공지능(AI) 붐의 엄청난 규모를 보면, 이 수조 원 규모의 산업 기반이 30년 전인 뮌헨에서 마련되었다는 사실을 잊기 쉽습니다. 오늘날 전 세계 최고의 기술 기업들은 ChatGPT와 같은 대형 언어 모델(LLM)의 확장에 수천억 원을 투자하고 있습니다. 하지만 기계 학습 커뮤니티의 일부 역사 매니아나 올드 스퀘어들을 제외하고는, 이러한 현대 시스템의 핵심 빌딩 블록 거의 모두가 1991년 불과 몇 달이라는 짧은 기간 동안 모두 발표되었다는 사실을 아는 사람은 거의 없습니다.

놀랍게도 이 기술들은 모두 Jürgen Schmidhuber가 이끄는 뮌헨 공과대학(TU Munich)의 단 하나의 연구실에서 나왔습니다. 그해가 끝나기 전에, 그의 팀은 본질적으로 심층 학습(Deep Learning)의 현대 시대를 설계해 두었습니다. 그들은 최초의 트랜스포머(Transformer) 변형 기술(ChatGPT의 'T'에 해당)을 발표했고, 비지도 사전 학습(unsupervised pre-training, ChatGPT의 'P') 개념을 도입했으며, 신경망 증류(Neural Network Distillation) 기술을 개척했습니다.

또한 그들은 각각 20세기와 21세기의 가장 많이 인용된 AI 논문인 LSTM과 ResNet의 핵심인 '심층 잔차 학습(deep residual learning)'을 도입했습니다. 이 네 가지 기술이 오늘날 가장 진보된 LLM을 구동하는 원동력입니다. 게다가 그들은 '생성형 AI'의 기반이 되는 생성적 적대 신경망(Generative Adversarial Networks)의 초기 토대도 마련했습니다. 내가 구글 브레인에서 일하던 시절부터 현재 사카나 AI에서 추진하고 있는 재귀적 자기 개선(RSI) 연구에 이르기까지, Jürgen의 기여는 수년간 제 사고에 깊은 영향을 미쳤습니다. 특히 그의 연구실이 1990년대에 소개한 개념을 바탕으로 2018년 세계 모델(World Models)을 대중화하는 데 도움을 준 것을 자랑스럽게 생각합니다. 이 아이디어들이 어떻게 시간의 시험을 견뎌내고, 규모를 확장하여 전 세계 AI 커뮤니티에서 완전히 수용되고 있는지 보는 것은 정말 놀라운 일입니다!

심층 학습의 진정한 역사에 관심이 있는 분들을 위해, Jürgen은 1991년 뮌헨에 이 씨앗들이 어떻게 심어졌는지에 대한 상세한 타임라인을 아래에 정리해 두었습니다. David Ha, 2026년 6월

Jürgen Schmidhuber의 1991년 타임라인 및 주해 참고 문헌 계산 비용이 오늘날보다 수백만 배 더 비쌌던 시절, 고향 도시에서 내 팀이 1991년에 수행한 연구와, 그곳에서 그 이후에 함께 일했던 모든 훌륭한 사람들을 자랑스럽게 생각합니다. 뮌헨 공과대학의 1991년 3월 26일부터 8월 31일로 거슬러 올라가는 다음 핵심 AI 간행물들을 확인해 보세요.

★ 1991년 3월 26일: 최초의 형태의 트랜스포머(ChatGPT의 T 참조) — 현재 비정규화 선형 트랜스포머(정규화된 이차 트랜스포머의 전신)라고 불립니다. 이 모델은 효율성 측면에서 여전히 중요합니다. 입력 크기에 따라 계산 비용이 제곱이 아닌 선형적으로 증가하기 때문입니다. ★ 1991년 4월 30일: 심층 신경망(NN)을 위한 사전 학습 — ChatGPT의 P에 해당합니다. 이를 통해 매우 깊은 심층 학습이 가능해졌습니다. ★ 1991년 4월 30일: 신경망 증류 — 2025년 유명한 DeepSeek '스푸트니크' 및 기타 대형 언어 모델(LLM)의 핵심이 되는 기술입니다. ★ 1991년 6월 15일: 매우 깊은 신경망을 위한 잔차 연결을 사용한 심층 잔차 학습. 20세기 가장 많이 인용된 AI 기술이자 2010년대 최초의 LLM(ELMO, ULMFiT)의 기반이 된 장단기 메모리(LSTM)의 핵심 요소입니다. 21세기 최다 인용 과학 논문 역시 심층 잔차 학습에 초점을 맞추고 있으며, 우리의 LSTM에서 영감을 받은 이전의 순방향 신경망보다 10배 더 깊은 심층 잔차 하이웨이 네트워크(Highway Net)의 변형을 다루고 있습니다. 심층 잔차 학습은 현재 사실상 모든 LLM에서 사용되고 있습니다. ★ 1991년 8월 31일: 인공적인 호기심과 창의성을 통해 훈련된 신경망 세계 모델을 위한 생성적 적대 신경망에 대한 최초의 동료 심사 논문.

원문 보기
원문 보기 (영어)
David Ha , Sakana AI Jürgen Schmidhuber , KAUST & IDSIA 18 June 2026 @hardmaru @SchmidhuberAI AI Blog Munich 1991: the Roots of the Current AI Boom Preface by David Ha When we look at the massive scale of today’s Artificial Intelligence boom, it is easy to forget that the foundations of this trillion-dollar industry were laid down over 30 years ago in Munich. Today, the world's top tech companies are investing hundreds of billions into scaling up Large Language Models (LLMs) such as ChatGPT. Yet, outside of a few history buffs or old-school folks in the Machine Learning community, people might not realize that virtually every core building block of these modern systems was published in a span of just a few months back in 1991 . Incredibly, they all emerged from a single lab at the Technical University Munich led by Jürgen Schmidhuber . Before that year ended, his team had essentially mapped out the modern era of deep learning. They published the very first Transformer variant (see ChatGPT's "T"), introduced the concept of unsupervised pre-training (ChatGPT's "P"), and pioneered neural network distillation . They also introduced deep residual learning , the centerpiece of both LSTMs and ResNets , the most cited AI papers of the 20th and the 21st century, respectively. These four techniques power today's most advanced LLMs. Furthermore, they laid the early groundwork for generative adversarial networks , foundational for "Generative AI." Jürgen’s contributions have deeply shaped my own thinking over the years, from my time at Google Brain to our recursive self-improvement (RSI) research we're currently pushing at Sakana AI . I am especially proud to have helped popularize World Models back in 2018, building directly on concepts his lab introduced in the 1990s. It is amazing to see how well some of these ideas have stood the test of time, scaling up to be fully embraced by the global AI community! For those interested in the real history of deep learning , Jürgen has put together a detailed timeline below of exactly how these seeds were planted in Munich in 1991. David Ha , June 2026 Jürgen Schmidhuber's 1991 Timeline, with Annotated References I am proud of the work my team did in 1991 in my home city when compute was millions of times more expensive than today [RAW] , and of all the great people I worked with there and afterwards. Check out TU Munich's following key AI publications dated 3/26/1991—8/31/1991. ★ 26 March 1991: the first kind of Transformer (see the T in ChatGPT)—now called the unnormalized linear Transformer [ULTRA] [FWP0-6] [WHO10] [DLH] : the predecessor of the normalized quadratic Transformer [TR1] . ULTRA is still important, also because of its efficiency: its computational costs scale linearly in input size, rather than quadratically . ★ 30 April 1991: Pre-Training for deep neural networks (NNs) —the P in ChatGPT [UN0] [UN1] [UN2] [UN] [DLH] . This enabled very deep learning [WHO5] . ★ 30 April 1991: Neural network distillation —central to the famous 2025 DeepSeek "Sputnik" and other Large Language Models (LLMs) [UN0] [UN1] [UN2] [WHO9] [DLH] . ★ 15 June 1991: deep residual learning with residual connections for very deep NNs [WHO11] (see Sepp Hochreiter's diploma thesis [VAN1] ): the core ingredient of Long Short-Term Memory [LSTM1] , the most cited AI of the 20th century, basis of the first LLMs in the 2010s (ELMO, ULMFiT). The most-cited scientific article of the 21st century [MOST25-26] is also about deep residual learning , focusing on a variant of our LSTM-inspired deep residual Highway Net [HW1-25b] that was 10 times deeper than previous feedforward NNs [WHO11] [DLH] . Deep residual learning is now being used in virtually all LLMs. ★ 31 August 1991: first peer-reviewed publication [GAN91] on generative & adversarial networks [GAN90-25] for neural world models [WM26,WM26b] trained through artificial curiosity & creativity —now controversially used for deepfakes and other applications of Generative AI [WHO8] [DLH] . As of January 2026, the two most frequently cited papers of all time (with the most citations within 3 years—manuals excluded) are directly based on the work of 1991 [MOST26] [MOST] [MIR] . In 1991, however, it was already totally obvious that LLM-like NNs alone are not enough to achieve Artificial General Intelligence (AGI). No AGI without mastery of the real world [DLH] ! That's why we started working on additional techniques required to achieve AGI, e.g., planning with adaptive world models [PLAN1-6] [WM26,WM26b] created by artificial scientists [AC] (since 1990 at TU Munich), meta learning & recursive self-improvement (since 1987) [META1] [META] , and others [DLH] [AIB] . Around the same time , Munich also was the origin of the first self-driving cars in traffic [AUT] (by Ernst Dickmanns's team), going up to 175 km/h. The city was truly the epicenter of AI. In the past 3 decades, however, most of commercial AI has shifted to the Pacific Rim, far away from Munich. How could that happen? Can anything be done about it? See [95-25] for answers! See [WHO3-11] for the broader historical context [DLH] of the work published in 1991 [MIR] . I am still hoping that I may live to see our great field of Machine Learning realize my 1970s teenager vision of building something much smarter than myself, such that I can retire. Jürgen Schmidhuber , June 2026 Acknowledgments Thanks to several expert reviewers for useful comments. (Let us know if you can spot any remaining error.) The contents of this article may be used for educational and non-commercial purposes, including articles for Wikipedia and similar sites. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License . Annotated References [95-25] J. Schmidhuber ( AI Blog , 2025). 1995-2025: The Decline of Germany & Japan vs US & China. Can All-Purpose Robots Fuel a Comeback? In 1995, in terms of nominal gross domestic product (GDP), a combined Germany and Japan were almost 1:1 economically with a combined USA and China, according to IMF. Only 3 decades later, this ratio is now down to 1:5! Self-replicating AI-driven all-purpose robots may be the answer. Based on a 2024 F.A.Z. guest article . [AC] J. Schmidhuber ( AI Blog , 2021, updated 2025). 3 decades of artificial curiosity & creativity . Schmidhuber's artificial scientists not only answer given questions but also invent new questions. They achieve curiosity through: (1990) the principle of generative adversarial networks, (1991) neural nets that maximise learning progress, (1995) neural nets that maximise information gain (optimally since 2011), (1997) adversarial design of surprising computational experiments, (2006) maximizing compression progress like scientists/artists/comedians do, (2011) PowerPlay... Since 2012: applications to real robots. [AIB] J. Schmidhuber's AI Blog. With lessons on the history of AI & computing, e.g.: Who invented deep learning? Who invented backpropagation? Who invented convolutional neural networks? Who invented artificial neural networks? Who invented generative adversarial networks? Who invented Transformer neural networks? Who invented deep residual learning? Who invented neural knowledge distillation? Who invented the computer? Who invented the transistor? Who invented the integrated circuit? ... [ATT] J. Schmidhuber ( AI Blog , 2020, updated 2025). 30-year anniversary of end-to-end differentiable sequential neural attention. Plus goal-conditional reinforcement learning. Schmidhuber had both hard attention for foveas (1990) and soft attention in form of Transformers with linearized self-attention (1991-93). [FWP] Today, both types are very popular. [AUT] J. Schmidhuber ( AI Blog , 2005). Highlights of robot car history . Around 1986, Ernst Dickmanns and his group at Univ. Bundeswehr Munich built the world's first real autonomous robot car