구글 딥마인드의 단백질 구조 예측 AI '알파폴드(AlphaFold)'의 성공 이후 데이터 중심 AI가 과학 발전의 정답으로 여겨졌으나, 방대하고 정확한 데이터 구축에는 엄청난 시간과 비용이 든다는 한계가 명확합니다. 따라서 데이터 의존도를 낮추고 불확실성 속에서도 스스로 실험을 기획하고 추론하는 'AI 에이전트(AI Agent)'가 과학 발전을 가속화하는 현실적인 대안으로 주목받고 있습니다.
번역된 본문
수십 년마다 누군가는 과학이 종말에 다다랐다고 선언하곤 한다. 1903년, 존경받는 물리학자 알버트 마이컬슨은 “물리학적 사실은 이미 모두 발견되었다”고 말했다. 1980년대에는 스티븐 호킹이 이론 물리학이 세기 말이면 완결될지도 모른다고 예측했다. 인공지능(AI)의 폭발적인 등장과 함께 이러한 느낌이 다시금 공기 중에 감돌고 있다. 이번에는 노벨상까지 동반한 채 말이다. 2024년, 구글 딥마인드의 데미스 허사비스와 존 점퍼는 수천 개의 실험적으로 측정된 형태를 학습하여 단백질의 3차원 구조를 예측하는 신경망 '알파폴드(AlphaFold)'의 성과로 노벨 화학상의 일부를 수상했다. 이 지독한 문제는 반세기 동안 체계적인 공격에도 굴복하지 않았었다. 하지만 알파폴드가 이를 단번에 해결한 것처럼 보였고, 세상은 그 접근 방식이 주는 약속에 매료되었다. 허사비스와 그의 팀은 알파폴드를 “AI가 모든 과학을 디지털 속도로 가속화할 수 있는 방법의 표본(template)”이라고 불렀다. 딥마인드의 성공에 힘입어 생물학, 화학, 신소재 발견을 위한 파운데이션 모델(foundation model)을 구축하는 수많은 스타트업들이 수십억 달러의 자금을 조달했다. 알파폴드는 AI와 충분한 데이터의 결합이 (우리가 기저 메커니즘을 이해하지 못하더라도) 획기적인 발견을 만들어낼 수 있음을 보여주었고, 남은 과학의 영역을 향한 길이 우리 앞에 펼쳐진 듯했다. 물론 AI는 과학에 엄청난 변화를 가져올 것이다. 하지만 알파폴드와 같은 사례가 그러한 혁신을 위한 최선의 표본이 아닐 수 있다는 사실이 점점 더 분명해지고 있다. 심오한 성취임에는 틀림없으나, 알파폴드와 같은 결과물을 탄생시킨 조건은 매우 드물며, 다른 분야에서 그러한 조건을 충족시키는 데 걸리는 시간은 몇 년이 아니라 수십 년 단위로 측정될 것이다. 대신, 과학의 가속화는 또 다른 접근 방식인 'AI 에이전트(AI Agent)'를 통해 이루어질 것이다. 알파폴드 성공의 주요 조건은 딥마인드 팀이 모델을 훈련시킬 수 있었던, 실험적으로 검증된 약 17만 개의 단백질 구조 데이터 세트인 '단백질 데이터 뱅크(Protein Data Bank)'의 존재였다. 단백질 데이터 뱅크의 구축은 결코 간단하지 않았다. 최근 추산치에 따르면, 이를 완성하기 위해 53년의 국제적인 과학적 협력과 약 210억 달러(약 2조 8천억 원) 상당의 실험 작업이 필요했다. 그 정도 규모의 노력은 자금을 조달하기가 어렵기로 악명 높고, 조정하는 것은 거의 불가능에 가까우며, 실행하는 데 엄청난 시간이 소요된다. 그 결과 흔히 실패로 끝나곤 했다. 하지만 필수적인 응집력과 자원이 있고 관련 데이터가 상업적 소유권으로 인해 접근 불가능하지 않은 분야라 할지라도, 너무나도 조명받지 못하는 또 다른 장벽이 있다. 바로 비교 가능한 데이터를 생성하는 것이 과학적으로 불가능하다는 점이다. 단백질 구조의 경우, 핵심 실험 기술인 '단백질 결정학(protein crystallography)'은 비정상적으로 복제하기 쉽고 신뢰할 수 있는 도구여서 25개 이상의 노벨상이 이에 의존했다. 하지만 대부분의 실험 과학에서 결과는 십중팔구 변동된다. 세포주는 표류하고, 화학물질에는 미량의 오염 물질이 있으며, 실험실 습도 또한 변한다. 생물학이나 대부분의 화학 분야에서 최신 신경망을 훈련시킬 만큼 충분히 일관성 있고, 정확하고, 정밀하며, 확장 가능한 측정 데이터 세트를 구축하려면 새로운 종류의 측정 방식과 새로운 표준화된 접근 방식이 필요할 것이다. 하지만 이러한 도구들은 결코 빠른 시일 내에 준비되지 않을 것이다. 물론 이러한 요구 조건을 충족하는 소수의 분야도 존재한다. 일기 예보, 상당 부분의 유전체학, 매우 제한된 화학 영역 등이 그것이다. 아직 그렇지 않았다면, 이 분야들은 조만간 알파폴드 방식의 획기적인 발전을 목격할 수 있을 것이다. 미국 국가안보 신생물기술위원회(US National Security Commission on Emerging Biotechnology)가 주장한 바와 같이, 이러한 데이터 세트의 생산 및 조정에 대한 정부의 지원이 매우 중요할 것이다. 하지만 과학의 대부분의 미해결 문제에 대해서는, 적어도 단기적으로는 다른 대안이 필요하다. 다행히도, 조용하고 겸손하지만 유망한 무언가가 등장하기 시작했다. 과학자들은 항상 불확실성 속에서 추론해 왔다. 신원을 파악하기 위해 노력하는 생물학자...
Every few decades, someone announces that science has reached its end. In 1903, the revered physicist Albert Michelson wrote that the “facts of physical science have all been discovered.” In the 1980s, Stephen Hawking predicted that theoretical physics might be finished by the end of the century. With the explosive arrival of artificial intelligence, the feeling is in the air again—this time accompanied by a Nobel Prize. In 2024, Demis Hassabis and John Jumper of Google DeepMind were awarded part of the Nobel in chemistry for their neural network AlphaFold, which predicts the three-dimensional structures of proteins by learning from thousands of experimentally measured shapes. This devilish problem had resisted systematic attacks for half a century; AlphaFold seemed to have solved it once and for all, and the world became fixated on the promise of its approach. Hassabis and his team called AlphaFold “the template for how AI can accelerate all of science to digital speed.” A wave of startups building foundation models for biology, chemistry, and materials discovery raised billions of dollars, buoyed by DeepMind’s success. AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries (even if we did not understand the underlying mechanisms involved), and it seemed, once again, that a path through the rest of science was laid out before us. To be sure, AI will bring extraordinary changes to science, but it has become increasingly clear that AlphaFold, and things like it, may not be the best template for that metamorphosis. Though it is a profound achievement, the conditions that produced the likes of AlphaFold are rare, and the time it will take to meet those conditions in other fields will be measured in decades, not years. Instead, the acceleration of science will come about thanks to another approach: AI agents. The primary condition for AlphaFold’s success was the existence of the Protein Data Bank, a data set of roughly 170,000 experimentally validated protein structures on which DeepMind’s team could train its model. The creation of the Protein Data Bank was not simple: It took 53 years of international scientific cooperation and, by a recent estimate , roughly $21 billion worth of experimental work to assemble. Efforts of that scale are infamously difficult to fund, next to impossible to coordinate, and hugely time-consuming to execute; they have often been unsuccessful as a result. But even in fields with the requisite cohesion and resources, and where the relevant data are not rendered inaccessible by commercial ownership, another barrier is too little discussed: the scientific impossibility of generating comparable data. In the case of protein structures, the key experimental technique—protein crystallography—is an unusually replicable and dependable tool, so much so that over 25 Nobel Prizes have relied on it. But in most of experimental science, results vary more often than not. Cell lines drift. Chemicals have trace contaminants. Lab humidity changes. The creation of measured datasets that will be consistent enough, accurate enough, precise enough, and scalable enough to train a modern neural network in biology or most of chemistry would require new kinds of measurement and new standardized approaches—none of which will be ready anytime soon. Of course, there are a handful of fields where these requirements are met: weather forecasting, much of genomics, very limited areas of chemistry. These may see AlphaFold-style breakthroughs soon, if they haven’t already . Government support for the production and coordination of those datasets will be critical, as the US National Security Commission on Emerging Biotechnology has argued . But for most open questions in science, we will need a different plan, at least in the short term. Luckily, something quieter and more modest has begun to show promise. Scientists have always reasoned under uncertainty. Biologists working to identify new drug targets have never had perfect datasets. Instead, they combine docking calculations and known structures, factor in molecular dynamics, run a handful of binding assays, and use their judgment to weigh each method according to its particular strengths and points of failure. The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in. This is how most working research actually proceeds. But until very recently, no software could do it. Agents now can. Simply put, an agent is an AI reasoning engine that has been given access to tools—digital or physical—and the capabilities to use them. Over the last few years, a fundamental architectural shift in AI has enabled the rapid proliferation of these programs, which are powered by large language models, dramatically reducing the need for scientifically specialized datasets. For science, this technological advancement represents a foundational change: it has allowed us to create digital tools that can mimic the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery. Consider Google’s AI Co-Scientist , announced in May. Researchers gave it a one-page brief and a goal: Figure out how antibiotic resistance spreads between bacterial species, a key driver of drug-resistant infections. The system spun up sub-agents. One drafted hypotheses from the literature. Another picked them apart like a peer reviewer. A third ran tournaments to rank the strongest candidates. A fourth refined the winning hypothesis. The agent concluded that resistance genes were hitching rides on bacterial viruses, borrowing whichever virus could ferry them into a new host. The hypothesis was correct. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking wet-lab work; their paper, previously unseen by Co-Scientist, was still in peer review. Agents like Co-Scientist are still novel tools, and there are real challenges to overcome before they become a ubiquitous part of the scientific process: They are still liable to hallucinate, their judgment is not consistent, and they have memory and input constraints that limit the time they can run autonomously. But these technical barriers will fall away, and as they do we will begin to notice the compounding effects of scientific agents on the reliability, consistency, and velocity with which science is done. Perhaps most notably, agents offer a structural fix for science’s “reproducibility crisis,” the widespread problem of researchers’ inability to replicate each other’s results. For decades, the scientific community has begged researchers to share their raw data and exact code in an effort to standardize experimental processes. But researchers have long resisted this tedious administrative work, which happens after the interesting science is already done. Agents, in contrast, automatically log every move they make, creating an exact record of the method that led to their results and allowing for precise replication. A second consequence will be an amplification of scientific memory. The transfer of knowledge between researchers is a famously murky process; if it isn’t done over years of training and observation, graduate students are left to pore through the messy lab notebooks kept by decades of predecessors, looking for the details that will make or break their protocol. As agents become an increasingly large part of the scientific process, though, a lab’s entire scientific history will be recorded in a central, standardized repository of institutional knowledge. But the most important impact of agents will be speed. In any field, when testing an idea takes less time than arguing about it in a meeting, people stop debati