메뉴
BL
TechCrunch AI • 45일 전

미공개 앤스로픽 AI, 난제 '리만 가설' 증명서 진전

IMP
9/10
핵심 요약

앤스로픽의 아직 공개되지 않은 AI 모델이 150년 된 수학계 최대 난제 중 하나인 '리만 가설(Riemann hypothesis)'에 대한 하한선을 크게 높이는 성과를 냈습니다. 수학적 배경이 없는 직원의 단순 프롬프트 지시 하나로 AI가 60개의 하위 에이전트를 조율하여 약 1만 5천 달러의 연산 비용을 들인 끝에 650개의 접근법을 테스트하며 스스로 새로운 수학적 아이디어를 도출해 낸 것입니다. 이는 대형 언어 모델(LLM)이 단순 번역이나 요약을 넘어, 독자적인 탐구 및 코드 작성을 통해 과학적·수학적 발견을 이끌어내는 '연구 자동화 에이전트'로서의 잠재력이 입증된 중요한 사례입니다.

번역된 본문

150년이 넘는 기간 동안 '리만 가설(Riemann hypothesis)'은 소수의 분포에 대한 오랜 미스터리로서 수학에서 가장 크게 풀리지 않은 문제 중 하나로 자리매겨해 왔습니다. 현재 이 가설에 대한 일반적인 증명에는 100만 달러의 현상금이 걸려 있지만, 아직 주인을 찾지 못했습니다. 최신 AI 모델들 역시 이를 완전히 해결할 수는 없습니다. 하지만 예상보다 훨씬 더 큰 진전을 만들어낼 수 있으며, 이 발견은 최신 AI가 새로운 과학 및 수학적 아이디어를 발견하는 능력에 대한 오랜 논쟁을 다시 불붙일 가능성이 높습니다.

월요일, 앤스로픽(Anthropic)은 아직 출시되지 않은 모델이 리만 가설에서 중요한 진전을 이루어, 해당 가설이 참인 해의 하한선(lower bound)을 크게 높였다고 발표했습니다. 더욱 인상적인 것은 이 진전이 이루어진 방식입니다. 심층적인 수학적 훈련을 받지 않은 앤스로픽의 한 직원이 이 가설을 증명해 보라며 모델에게 '진지하게 시도해 보(take a real stab)'라고 프롬프트를 입력한 뒤, 이후 하루 반나절 동안 모델이 스스로 작업을 조정하도록 내버려 두었습니다. 총합하여 이 모델은 문제 해결을 위해 650가지의 다른 아이디어를 테스트했으며, 60개의 하위 에이전트(sub-agent)에 걸쳐 조정을 수행하고 총 31만 토큰(비용 약 1만 5천 달러로 추정)을 소비했습니다.

논문의 각주는 이렇게 설명합니다. "60개의 하위 에이전트 중 2개는 핵심 수학적 아이디어를 개발하는 역할을 담당했고, 13개는 이 에이전트들에게 아이디어를 제공했으며, 30개는 새로운 아이디어를 개발하려 시도했지만 실패했습니다. 13개는 논증의 정확성을 검증하는 검증자 역할을 했고, 마지막 2개는 초기 논문 작성을 도왔습니다."

이 발견은 앤스로픽의 내부 수학자 두 명에 의해 확인되었으며, 오픈소스 증명 보조기인 '린(Lean)'을 사용해 정형화되었습니다. 이는 대형 언어 모델(LLM)이 주도한 수학적 돌파구의 연속 중 최신 결과입니다. 올해들어 여러 에르되시(Erdos) 문제가 AI 모델에 의해 해결되었으며, 더 강력한 모델이 출시되면서 더욱 인상적인 결과가 도출되고 있습니다. OpenAI는 최근 내부 모델인 '아스트라(Astra)'가 증명한 10가지 주요 결과 세트를 공개했고, 앤스로픽의 별도 노력으로는 오랜 난제였던 야코비안 추측(Jacobian conjecture)을 반증했습니다.

이러한 결과가 계속 늘어나면서 수학계는 흥분과 우려를 동시에 느끼고 있습니다. 6월에 서명된 공개 선언에서 저명한 수학자 그룹은 AI가 이 분야의 핵심 가치를 훼손할 수 있다고 우려했습니다. 특히 진정한 수학적 증명은 '그 발견에 대한 공로를 인정받고 정확성에 대해 책임을 지는 특정 저자에게 귀속되어야 한다'는 표준을 들었습니다. 하지만 수학자들이 이 새로운 연구 기술에 어떻게 접근해야 하는지에 대해서는 여전히 학계의 의견이 엇갈립니다. 이 선언에 대한 응답으로 작성된 블로그 게시물에서 필즈 메달리스트 티모시 가워스(Timothy Gowers)는 AI의 영향이 수학을 더 복잡하고 긍정적인 방식으로 바꿀 수 있는지에 의문을 제기했습니다. 가워스는 "수학적 정리가 더 이상 수학자와 연관되지 않는 세상에 도달한다면, 아마도 별이 천문학자의 이름을 따서 명명되지 않고 대부분 아예 이름이 없다는 사실보다 더 문제가 되지는 않을 것"이라고 적었습니다.

원문 보기
원문 보기 (영어)
For more than 150 years, the Riemann hypothesis has stood as one of the major unsolved problems in mathematics, a long-running mystery about the distribution of prime numbers. There is currently a $1 million bounty for a working general proof of the hypothesis, which remains unclaimed. Contemporary AI models still can't solve it either — but they can make a lot more progress than you might expect, a finding that's likely to reopen long-standing questions about contemporary AI's ability to discover new scientific and mathematical ideas. On Monday , Anthropic announced that an as-yet-unreleased model had made significant progress on the Riemann hypothesis, significantly increasing the lower bound of solutions for which the hypothesis holds true. Even more impressive is how the progress was made: An Anthropic staff member without significant mathematical training prompted the model to "take a real stab" at proving the hypothesis, then left the model to coordinate the task across the following day and a half. All told, the model tested 650 different ideas for solving the problem, coordinating across 60 sub-agents and spending 31 million in total. "Out of the 60 subagents, two were responsible for developing the key mathematical ideas," a footnote to the paper explains, "13 contributed ideas to these agents, 30 attempted (but were unable) to develop new ideas, 13 served as validators to check the correctness of the arguments, and the final two helped to write the initial paper." The finding was confirmed by two of Anthropic's in-house mathematicians, and formalized using the open-source proof assistant Lean . This is the latest in a string of mathematical breakthroughs led by Large Language Models, or LLMs. A number of Erdos problems have been solved by AI models over the course of this year, and the release of more powerful models has led to more impressive results. OpenAI recently released a set of ten major results proved by its internal "Astra" model, while a separate effort from Anthropic disproved the long-standing Jacobian conjecture. The growing body of results has caused both excitement and concern in the mathematical field. In a public declaration signed in June , a group of prominent mathematicians raised concerns that AI could undermine critical values of the field — particularly the standard that true mathematical proofs should be "attributable to specific authors who take credit for their discovery and assume responsibility for their correctness." But the field is still split on how mathematicians should approach the new research techniques. In a blog post responding to the declaration , Fields Medal winner Timothy Gowers questioned whether the influence of AI might change mathematics in a more complex and positive way. "If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all," Gowers wrote. Topics AI , Anthropic , mathematics When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $300 toda y! REGISTER NOW Most Popular Mark Zuckerberg's AI manifesto is exactly why people don't like AI Russell Brandom YouTube now requires creators to have twice as many watch hours to start earning money Aisha Malik This ‘adversarial' pattern can prevent surveillance cameras from detecting you Zack Whittaker ChatGPT brings unlimited text chats to free users Ivan Mehta Tesla and SpaceX will invest $16.8B to start building ‘Terafab' chip factory in Texas Sean O'Kane Amid legal battles, Suno says it will start watermarking songs Ivan Mehta Ford's new electric truck, ‘Fathom,' starts at $28,350 Sean O'Kane