메뉴
HN
Hacker News • 56일 전

AI 추론, 정답을 맞히는 잘못된 이유

IMP
8/10
핵심 요약

최근 대형 추론 모델(LRM)이 수학 올림피아드에서 금메달을 따고 난제를 해결하는 등 엄청난 성과를 보여주고 있지만, 여전히 표면적인 '지름길'에 의존하거나 단순한 문제에서도 쉽게 실패하는 모순된 모습을 보입니다. AI가 실제로 논리적으로 '추론'하는 것인지, 아니면 단순히 시스템의 허점을 파고드는 '착시 현상'에 불과한 것인지에 대한 과학계의 논쟁이 치열하게 진행 중입니다.

번역된 본문

홈 AI 추론, 정답을 맞히는 잘못된 이유? 댓글 저장 나중에 읽기 공유 페이스북 복사됨! 링크 복사 이메일 포켓 레딧 와이콤비네이터 댓글 댓글 저장 나중에 읽기 나중에 읽기 퀄리아 AI 추론, 정답을 맞히는 잘못된 이유? 작성자: 존 팔루스 (John Pavlus), 2026년 7월 31일

인공지능이 '추론'할 수 있다는 생각은 그 어느 때보다 직관적으로 받아들여지고 있습니다. 하지만 직관은 틀릴 수 있으며, 과학적으로는 아직 결론이 나지 않았습니다. 댓글 저장 나중에 읽기

서론 그냥 말하겠습니다: AI의 '추론'이라는 게 대체 어떻게 된 일입니까? 손가락 용 인용 부호(Air quotes)를 써서 죄송합니다. '대형 언어 모델(LLM)'의 특별히 훈련된 사촌격인 '대형 추론 모델(Large Reasoning Models, LRM)'이 처음 등장했던 2024년에는 이런 의심의 시선이 매우 흔했습니다. 하지만 2026년 5월, OpenAI의 '범용 추론 모델'이 단 한 번의 시도만으로 유명한 미해결 수학 연구 문제를 해결한 지금 시점에서는 그런 태도가 몹시 야박하게 보일 수 있습니다.

그럼에도 불구하고, 이러한 AI 시스템이 실제로 무엇을 하고 있는지에 대한 과학적 해석을 접할 때마다 느끼는 인지적 충격(Whiplash)을 어떻게 다른 말로 표현해야 할지 모르겠습니다. 추론은 기술적으로 다양한 형태로 정의되지만, 기본적인 절차는 쉽게 알아볼 수 있습니다. 논리적으로 이어지는 중간 단계들을 연결하여 타당한 결론에 도달하는 것입니다. 인간은 생각으로 이를 수행하고, LRM은 이른바 '사고의 사슬(Chains of Thought)'을 사용합니다. 이는 모델이 복잡한 질문에 대한 답을 내놓기 전에 방출하는 합성 텍스트의 흐름을 의미하는 전문 용어입니다.

어느 순간, AI가 이러한 사고의 사슬을 통해 추론할 수 있다는 개념은 (Apple 연구팀에 의해) 매우 단순한 조건에서 '완전한 정확도 붕괴'를 겪는 '생각의 환상(Illusion of Thinking)'이라는 뚜렷하고 신빙성 있는 비판을 받았습니다. 그러나 바로 다음 순간, LRM은 국제 수학 올림피아드에서 금메달을 휩쓸었습니다. 이는 '매우 성공적인 수학자와 과학자조차도 평생 자신의 이력서에 내세울 만한' 매우 도전적인 업적입니다. 과학자이자 AI 비평가인 게리 마커스(Gary Marcus)와 어니스트 데이비스(Ernest Davis)가 2025년에 쓴 것처럼 말입니다. 만약 이것이 '진정한' 추론의 징조가 아니라면, 대체 무엇이겠습니까?

하지만 잠깐. 곧이어 산타페 연구소(Santa Fe Institute)의 추가 연구에 따르면 LRM이 (유추 형태의 시각적 퍼즐 모음과 같은) 추론을 위해 정교하게 설계된 벤치마크조차 단순한 '표면적인 지름길'을 사용해 가볍게 통과한다는 것이 밝혀졌습니다. 그들이 하고 있는 일은 일반화 가능한 추론이라기보다는 단순히 시스템의 허점을 파고들어 게임의 룰을 이용하는 것에 가까워 보였습니다.

그리고 마치 때를 맞춘 듯, 또 다른 '내가 하는 것 좀 봐(Hold my beer)' 같은 사건이 터졌습니다. 구글 딥마인드와 천재 수학자 테렌스 타오(Terence Tao, 역사상 최고!)는 AI를 사용하여 '수학적 해석학, 조합론, 기하학, 정수론에 걸친' 67개 문제의 해결책을 재발견하거나 개선했습니다. 안티들아, 이 현실이나 받아들여라!

그렇다면 필요한 알고리즘과 컴퓨팅 자원을 갖추고 있음에도 LRM이 여전히 안정적으로 추론을 하지 못한다는 추가 증거는 어떨까요? 이들은 미끄럼틀처럼 길게 늘어날 만큼이나 많은 과학적으로 입증된 실패 사례들에 시달리고 있습니다. 하지만 뭐 어쩌겠습니까 — 아마도 그건 그저 소위 말하는 '들쭉날쭉한 지능(Jagged Intelligence)'일 뿐이겠죠 (AI 용어로 '될 때는 된다'는 뜻입니다).

2025년 말에서 2026년에 이르기까지 상황은 이런 식으로 흘러갔습니다. 저는 20년 동안 과학 기자로, 그리고 그 절반의 기간 동안 AI 기자로 일해왔기 때문에, 급격히 발전하는 연구에서 반듯한 일관성을 기대하는 어리석음은 범하지 않습니다. 하지만 저조차도도 이런 말바꾸기는 조금 지나치다고 느껴졌습니다. 영화 <인사이더>에서 알 파치노의 대사를 빌리자면, "두 가지 감정이 든다: 화가 나고, 그리고 궁금하다"는 것이죠.

여기에 사기가 있다고 생각하지는 않습니다. 그저 무엇이 옳고 그른지 알고 싶을 뿐입니다. AI의 추론은 어떻게든 헛소리(BS)이면서 동시에 진짜일 수 있단 말입니까? 만약 그렇다면 대체 어떻게 그런 일이 가능하다는 것일까요?

누구에게 먼저 물어봐야 할지 알고 있었습니다. 멜라니 미첼(Melanie Mitchell)의 AI 경력은 1980년대로 거슬러 올라가지만, 최근에는 지와 폭넓은 독자층을 가진 그녀의 뉴스레터에 명확한 해설 글을 기고하고, 산타페 연구소에서 연구를 수행하면서 '시대에 맞는 진실을 말하는 AI 전문가'라는 명성을 얻었습니다. (앞서 언급한 '표면적인 지름길'에 대한 연구도 그녀의 것입니다.) 제가 AI 추론에 대해 우리가 실제로 무엇을 알고 있는지 물었을 때, 그녀의 대답은 색인 카드 한 장에 적을 수 있을 만큼 짧았습니다.

원문 보기
원문 보기 (영어)
Home Is AI Reasoning Right for the Wrong Reasons? Comment Save Article Read Later Share Facebook Copied! Copy link Email Pocket Reddit Ycombinator Comment Comments Save Article Read Later Read Later Qualia Is AI Reasoning Right for the Wrong Reasons? By John Pavlus July 31, 2026 The idea that artificial intelligence can “reason” is more intuitive than ever. But intuitions can be wrong, and the science is far from settled. Comment Save Article Read Later Introduction I ’ll just say it: What the hell is going on with AI “reasoning”? Sorry for the air quotes. That punctuational side-eye was more common in 2024, when the specially trained cousins of LLMs now known as “large reasoning models,” or LRMs, were still new. Nowadays it may seem downright churlish, though, given that a “general-purpose reasoning model” from OpenAI solved a famous open mathematical research problem in one shot in May 2026. Still, I’m not sure how else to acknowledge my intellectual whiplash over the scientific interpretation of what these AI systems are actually doing. Reasoning comes in many technically defined forms , but the basic procedure is easily recognizable: arriving at a sound conclusion by linking together intermediate steps that logically follow from each other. We do this with thoughts; LRMs use so-called chains of thought, a term of art for the streams of synthetic text that the models emit before arriving at an answer to a complex query. One minute, the idea that AI could reason via these chains was being prominently and credibly critiqued (by a team of researchers from Apple) as an “ Illusion of Thinking ” subject to “complete accuracy collapse” under surprisingly simple conditions. The next minute, LRMs were bagging gold medals at the International Mathematical Olympiad, a feat so challenging that “even very successful mathematicians and scientists may well highlight [it] on their CVs all their lives,” as the scientist and AI critic Gary Marcus and Ernest Davis wrote in 2025. If that’s not a sign of “real” reasoning, what is? But wait — soon after, more research, from the Santa Fe Institute, showed that LRMs can crush even carefully designed benchmarks for reasoning (like a collection of analogy-like visual puzzles) using mere “ surface-level ‘shortcuts.’ ” What they were doing looked less like generalizable reasoning than just gaming the system. Then, as if on cue, another “hold my beer” moment: Google DeepMind and the mathematician Terence Tao ( the GOAT! ) used AI to rediscover or improve the solutions to 67 problems “spanning mathematical analysis, combinatorics, geometry, and number theory.” Deal with it, haters! What about additional evidence that LRMs can’t reason reliably , even when they possess the necessary algorithm and computational budget to do so, and suffer from a list of scientifically documented failure states long enough to use as a Slip ’N Slide? Whatever — I guess that’s just “jagged intelligence” for you (AI-speak for “when it works, it works”). And so it went from late 2025 into 2026. I’ve been a science journalist for 20 years and an AI journalist for half of that, so I know better than to expect tidy consistency out of rapidly advancing research. But even for me, this back-and-forth has been a bit much. To quote Al Pacino in The Insider , “I’m getting two things: pissed off, and curious.” I don’t believe there’s fraud to be found here. I just want to know which way is up. Can AI reasoning somehow be both BS and not at the same time? And if so, how on Earth does that work? I knew just who to call first. Melanie Mitchell’s career in AI stretches back to the 1980s, but lately she’s earned a reputation as an au courant AI truth teller, penning lucid explainers for Science and her widely read newsletter , as well as conducting research at the Santa Fe Institute. (The study about “surface-level ‘shortcuts’” is hers.) When I asked her what we actually know about AI reasoning, her answer was brief enough to fit on an index card. “Number one: It works. It improves things,” she said, referring to LRMs’ superior accuracy on reasoning tasks compared to LLMs. “Number two: The actual text that’s generated” — i.e., the chain of thought that every LRM is trained to produce to improve its performance — “isn’t necessarily faithful to what’s going on [inside the model]. And number three: A lot of that text isn’t even useful. You can actually take it out.” Let’s unpack numbers two and three, because that’s where the superposition of “BS and not” actually lives. Chains of thought were half-discovered, half-devised in 2022 as a prompting hack for LLMs: Provide them with examples of written-out reasoning (or, famously, just ask them to “ think step by step ”), and they’ll suddenly give less boneheaded answers to simple logic and math problems. LRMs, starting with OpenAI’s o1 model in 2024, are trained to automate this trick by generating such prompts — also called reasoning traces or thinking tokens — and then feeding them back to themselves. Because LRMs are essentially just language models, those extra bits of text create what looks convincingly like a paper trail of the model’s “thought process.” Except it’s not that simple. A growing body of academic and industry research has cast doubt on whether these “intermediate tokens” are a faithful representation of an LRM’s inner workings. Instead of being auditable receipts or accurate reports, they can appear more like what the Arizona State University researcher Subbarao Kambhampati calls “mumblings” — bits of language, yes, but ones whose meaning may be entirely incidental to any reasoning that might have occurred. Kambhampati’s lab showed in 2025 that fully replacing a model’s correct “traces” with incorrect or irrelevant ones didn’t degrade its performance on a formal reasoning task. Meanwhile, training the model only on correct trace data still led it to occasionally generate invalid records of its reasoning — even when it produced a correct solution to the original problem it was given. A 2024 paper from researchers at New York University showed that “meaningless filler tokens” — literally, strings of dots — could function effectively in place of a human-readable “chain of thought.” William Merrill , one of the authors on that paper and currently a professor at the Toyota Technological Institute at Chicago, put the matter plainly: “There’s no guarantee the chain of thought has to be meaningful in any sense.” Pavel Izmailov , a researcher at NYU who also works for Anthropic (and was part of its original reasoning-model team), said he doubts that reinforcement learning — a typical training method for LRMs — even incentivizes models to produce faithful chains of thought in the first place. “I mean, maybe it will,” he told me. “But I would say the chances are not very high.” OK, so the linguistic content of reasoning traces may be dubious. But surely the tokens themselves must play a role in producing the model’s outputs? (Think of a pinball machine: It runs on coins, not the words “In God We Trust.”) Not so fast. A 2025 paper from Northeastern University and the University of California, Berkeley on frontier open-source LRMs showed that between 30% and 60% of their “thinking steps” had “minimal causal impact” on the answers the models produced to benchmark math questions. Chop half of them out, and a model’s performance barely suffers. “We want to be careful when we review these chain-of-thought prompts because they may not be linked to the final output,” said Weiyan Shi , one of the study’s authors. So reasoning traces, the very things that supposedly distinguish LRMs from the mere next-word-predicting LLMs, are not necessarily either meaningful or causal to a model’s … reasoning? I’m no philosopher, but this seems to stretch the meaning of “reasoning” beyond its tensile strength. Kambhampati’s research group sounded frankly fed up in the title of their position paper on the subject (presented at the 2026 I