메뉴
HN
Hacker News 43일 전

기계학습 연구와 선(禪)의 기술

IMP
8/10
핵심 요약

훌륭한 AI 연구자가 되기 위해서는 논문을 읽고 직접 모델을 구축하는 과정을 병행하는 규칙적인 수행이 필요합니다. 최신 유행을 쫓기보다는 교차 엔트로피나 SVD 같은 기본 개념을 깊이 이해하고, 단순한 벤치마크 점수 향상이 아닌 새로운 가능성을 시험할 수 있는 문제에 집중해야 합니다. 또한 기존의 작은 규모의 경험에 얽매이지 않고 '초심(Shoshin)'으로 돌아가 확장성(Scaling) 중심의 현대 AI 트렌드를 유연하게 받아들이는 태도가 중요합니다.

번역된 본문

AI 연구와 선(禪): 연구자의 자질(temperament)이 재능(talent)보다 중요하다 - Jack Morris (2026년 6월 15일)

당신은 AI 연구를 하고 싶은가? 아무도 그 방법을 제대로 가르쳐주지 않는다는 것은 사실입니다. 적어도 직접적으로는 말입니다. 하지만 시작하는 방법은 꽤 간단합니다. (i) 읽고 배우기와 (ii) 직접 만들어보기의 조합입니다. 둘 중 하나만 해서는 안 됩니다. 이 둘을 결합할 때 비로소 연구자가 되는 것입니다.

훌륭한 연구자가 되는 과정은 참선을 배우는 것과 다르지 않습니다.

I. 시작하는 방법은 (a) 읽고 배우기, 그리고 (b) 직접 만들어보는 것의 조합으로 매우 단순합니다. 하나만 해서는 안 됩니다. 이 조합을 통해 연구자로 성장하게 됩니다.

Token for Token 블로그를 읽어주셔서 감사합니다! 새로운 글을 받아보고 제 작업을 지원하려면 무료로 구독하세요.

이런 옛 선(禪)의 격언이 있습니다. "깨달음을 얻는 날에도 앉고, 깨달음을 얻지 못하는 날에도 앉는다." 연구하는 것은 기본적으로 이와 같습니다. 과학적 통찰은 무작위로 찾아오는 것처럼 보입니다. 대부분의 날에는 찾아오지 않죠. 성공을 위한 중요한 특성은 그저 시간과 노력을 투자하는 것입니다. 음악, 스포츠, 영업 등 다른 어떤 분야와 마찬가지로 세계 최고가 되고 싶다면 엄청난 규율과 훈련이 필요합니다.

노암 샤저(Noam Shazeer)는 SwiGLU 논문에서 성공적인 연구 아이디어의 본질적인 무작위성에 대해 재치 있게 이렇게 언급했습니다. "우리는 왜 이 아키텍처들이 잘 작동하는지에 대한 어떠한 설명도 제공하지 않습니다. 다른 모든 것들과 마찬가지로, 그 성공을 신의 자비에 돌립니다."

이와 관련된 한 가지 조언은, 논문을 너무 많이 읽는 것도 문제가 될 수 있다는 것입니다. 어떤 문제를 해결하고 싶다면, 검증된 성공의 길은 해결책을 시도해 보고, 실행해 보고, 병목 현태에 부딪히고, 이를 해결하려 노력하며, 스스로 아이디어가 바닥났을 때 비로소 문헌을 찾아보는 것입니다.

II. 좋습니다, 그럼 저는 무엇을 연구해야 할까요? 이제 막 시작하는 분이라면 제 솔직한 대답은 이렇습니다. 정확한 주제가 크게 중요하지는 않다고 생각합니다. 그럼에도 불구하고, 6개월 미만으로 유행한 주제를 선택하는 것은 경고하고 싶습니다. AI는 빠르게 변하지만, 근본적인 아이디어는 40년 동안 변하지 않았습니다. 이것을 평생의 직업으로 삼고 싶다면, '2026년의 개념'들(하네스, 에이전트, 컨텍스트 엔지니어링 등)에 대해 너무 깊이 고민하지 마세요. 이런 것들은 또 바뀔 테니까요.

대신, 기본기로 돌아갈 때 더 많은 것을 배울 수 있습니다. 교차 엔트로피(cross-entropy)가 무엇인지 배우세요. 작은 분포에 대해 직접 손으로 계산해 보세요. 머릿속에서 시각화할 수 있을 정도로 특이값 분해(SVD)를 깊이 이해하세요. 코딩을 위한 강화학습(RL)에만 너무 집착하지 말고, 대신 정책 경사도(Policy Gradient) 뒤에 있는 아이디어, 왜 유용한지, 그리고 왜 수십 년 동안 인기가 있었는지를 배우세요.

한 가지 더 메타적인 조언이 있습니다. 만약 여러분의 연구 프로젝트가 달성할 수 있는 최고의 결과가 기존 벤치마크에서 더 높은 점수를 받는 것이라면, 여러분은 충분히 깊이 들어가지 못한 것입니다. 종종 기존 데이터셋은 새롭고 흥미로운 기능을 테스트하지 못합니다. 제이슨 웨이(Jason Wei)도 비슷한 점을 지적했습니다. "AI 연구에서 (10년 전에는 실제로 존재하지 않았던) 과소평가되지만 때로는 성패를 가르는 기술은, 여러분이 작업 중인 새로운 방법을 실제로 제대로 테스트해 볼 수 있는(exercise) 데이터셋을 찾는 능력입니다."

구체적인 제안을 해달라면, 저는 할 수 없습니다. 그것은 여러분에게서 나와야 합니다. 깊이 들어가고, 기본에 집중하고, 벤치마크를 쫓지 마세요. 물속에 계속 머물러 있으면, 아이디어는 저절로 떠오를 것입니다.

III. 초심자의 마음에는 많은 가능성이 있지만, 전문가의 마음에는 가능성이 적다 - 스즈키

요즘 실리콘밸리에서 자주 반복되는 말 중 하나는, 현대에 있어서 AI 연구의 경험이 오히려 좋은 연구 직관을 방해할 수도 있다는 것입니다. 저는 이 현상을 가까이서 지켜봤습니다. 스케일링(Scaling) 시대 이전의 많은 연구자들이 여전히 소규모에서는 작동하지만 대규모로 테스트하면 분명히 실패할 방법을 설계하는 데 관심을 두고 있습니다.

OpenAI에 대해 정말 인상 깊은 점 중 하나는 (적어도 기술적인 측면에서) 회사를 운영하는 대부분의 사람들이 35세 미만이라는 것입니다. 챗GPT(ChatGPT) 뒤에 있는 중요한 의사결정권자들 중 상당수는 30세 미만입니다. 여기서 우리가 얻을 수 있는 교훈 중 하나는, AI는 아직 초창기이고 급변하는 분야이기 때문에 (글 하단 누락)...

원문 보기
원문 보기 (영어)
Zen and the Art of AI Research temperament >> talent Jack Morris Jun 15, 2026 35 4 3 Share So you want to do AI research? It’s true that no one really teaches you how. Not directly, anyway. But it turns out that the way to get started is pretty simple: some combination of (i) reading and (ii) building stuff. You can’t do one without the other. You become a researcher through the combination. It turns out the process of becoming a great researcher is not unlike learning to meditate: I. The way to get started is pretty simple, through some combination of (a) reading and learning, and (b) building stuff. You can’t only do one. You’ll become a researcher through this combination. Thanks for reading Token for Token! Subscribe for free to receive new posts and support my work. Subscribe There’s an old Zen saying that goes something like this – on days we find insight, we sit. on days we do not find insight, we sit. Doing research is basically like this. Scientific insights can come seemingly at random. Most days they will not come. An important trait for success is just putting in the time & effort. Like any other pursuit (music, sports, sales, etc.), if you want to become world-class, it will take a tremendous amount of discipline. Noam Shazeer makes a nice hat-tip to the inherent randomness of successful research ideas in the SwiGLU paper: “We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence.” A related comment is that it’s possible to read too many papers . If you want to solve a problem, the tried-and-true path to success is to attempt a solution, try it, reach a bottleneck, try to solve it, and only reach for literature when you’ve run out of ideas yourself. II. Fine, but what should I work on? If you’re just starting out, here’s my honest answer: I don’t think the exact topic matters much. That said, I would warn you against choosing things that have been popular for less than six months. AI moves fast, but the fundamental ideas haven’t changed in forty years. If you want to make a career out of this, I wouldn’t advise you to think too hard about the concepts of 2026: harnesses, agents, context engineering, etc. These will change. Instead, you’ll learn more by going back to the basics: learn what cross-entropy is. Compute it by hand for a small distribution. Deeply understand SVD, to the point where you can start to visualize it in your head. Don’t think too much about RL for coding specifically, instead learn the ideas behind policy gradients, why they’re useful, and why they’ve been popular for decades. One more meta-comment: if the best possible outcome of your research project is a higher score on an existing benchmark, you are not going deep enough. Often, existing datasets won’t test new interesting capabilities. Jason Wei makes a similar point : An underrated but occasionally make-or-break skill in AI research (that didn’t really exist ten years ago) is the ability to find a dataset that actually exercises a new method you are working on. As for a concrete suggestion, I can’t make one; that has to come to you. Go deep, focus on the basics, and don’t chase benchmarks. Stay in the water and the ideas will come. III. in the beginner’s mind there are many possibilities; in the expert’s mind there are few – Suzuki Something often-repeated in Silicon Valley these days is how experience in AI research might actually be counterproductive to good research intuition in the modern day. I’ve observed parts of this up-close; many researchers from the pre-scaling-era remain interested in designing methods that work at a small scale but will obviously fail when tested at scale. One really impressive thing about OpenAI is that most of the people running the company (on the technical side, at least) are under 35. Many of the important decisionmakers behind chatGPT are under 30. One thing we can take away from this is that since AI is such a nascent field (chatGPT is less than four years old!) no one has a huge advantage , because no one has been working on it for very long. In short, holding on to ideas for too long can actually be counterproductive. Stay open-minded and refuse to let ego cloud your judgement. IV. Inspiration strikes when you least expect it. Here are two examples from history: The discovery of the structure of the benzene ring famously came in a dream: the structure had never been seen before, but was imagined as a snake biting its own tail. Ozempic basically comes from lizards . The GLP-1 hormone it mimics was first found in the venom of the Gila monster, a desert lizard that eats just a few times a year. Somehow we figured out how to make this work for humans too. One important takeaway is that to do good research, you must do things other than research . Most of my personal “aha moments” happened away from the keyboard, especially when going on walks. Darwin, Tesla, Feynman, Aristotle. Many great thinkers of history proclaimed the outsized benefits of stretching your legs and going for a little stroll. Even if you don’t do research, you should probably go on more walks. V. Even when inspiration strikes, nature may not be benevolent: even with a perfect implementation, our idea might just not be true in some fundamental sense. Or perhaps it was, or seems to be. When the results come in, how should we react? Another principle we can borrow from Zen is (experimental) equanimity. When analyzing an experiment, we can channel the following mentality: Did it go well? Great! Did it go poorly? Also great! Both outcomes teach you the same amount of information. In fact, it’s often possible to learn more from a string of negative results than a single positive result. “Wow, it’s still not working – incredible!” Now that’s a healthy attitude for research. The converse of this is that you shouldn’t get that excited about good results. In fact, most good results come because of a bug; it’s not that the results themselves were good, it’s that you measured incorrectly, and convinced yourself. Everyone wants their ideas to work – and this is a good thing! – but one thing all experienced researchers share is extreme skepticism, especially in the face of outcomes that seem too-good-to-be-true. Unfortunately, they almost always are. VI. A flower does not think of competing with the flower beside it. It just blooms. Research is extremely outcome-driven. Especially in academia, it’s easy to look at others’ successes on paper and turn to emotions. People succeed for different reasons. Some people get lucky. The academic reviewing process, in particular, is neither consistent nor fair. When new research comes out in your area that you admire, ask yourself the following question: Am I operating at the proper level of depth to have made this insight myself? Now there are two possible outcomes. If the answer is yes – great. Your process is sound, but you didn’t make this finding; you were busy, you were doing something else, but you could’ve. And if the answer is no – then take this as motivation to go deeper. VII. before enlightenment, chop wood, carry water. after enlightenment, chop wood, carry water. Many successful projects typically involve hundreds of hours of gruntwork behind the scenes. Andrej Karpathy labeled a nontrivial portion of ImageNet by hand . The creators of SWEBench , who were ahead of their time in many ways, spent hundreds of hours painstakingly filtering GitHub data to get a small, tractable set of GitHub issues useful for evaluation. If you look at the career of great researchers, they likely spent lots of time working in obscurity before finding success. Get used to this. The more ambitious and forward-thinking an idea, the more work it may be to thoroughly implement and evaluate. This difficulty is a feature, not a bug. VIII. Collin Raffel, an amazing researcher whom I deeply respect, once mentioned that he thinks many ideas fail not because they’re bad ideas,