메뉴
HN
Hacker News 45일 전

화학자가 된 AI, 클로드의 NMR 분석 능력

IMP
7/10
핵심 요약

Anthropic은 세계 최고 수준의 화학자들과 협력하여 클로드가 다양한 화학 데이터를 해석하고 분석하도록 개선하고 있습니다. 특히 화학자들이 가장 많이 사용하는 분석 데이터인 NMR 스펙트럼(NMR spectrum)을 클로드가 어떻게 해석하는지 평가하는 첫 백서를 발표했습니다. 이는 AI가 단순히 데이터베이스를 조회하는 것을 넘어, 화학 구식을 시각적으로 읽고 논리적으로 추론하여 화학자들의 실무를 실질적으로 보조하는 중요한 전환점입니다.

번역된 본문

과학 클로드를 화학자로 만들기 2026년 6월 5일

저희는 클로드의 화학 분야 성능을 향상시키기 위해 세계 최고 수준의 합성, 계산 및 분석 화학자들과 협업하고 있습니다. 이번 게시물에서는 이러한 노력의 일환으로 수행된 첫 번째 연구를 공유합니다. Anthropic의 화학자인 David Kamber는 클로드가 화학자의 가장 일반적인 분석 입력 데이터인 NMR 스펙트럼(NMR spectrum)을 어떻게 분석하는지 조사했습니다.

분자를 다룰 때, 화학자들은 화이트보드에 그려진 손그림 구조, 기기 측정 결과, 데이터베이스 쿼리 문자열, 특허 및 출판물의 기술적 표기법 사이를 오가며 작업합니다. 이러한 각각의 표현 방식은 동일한 기본 화학적 의미를 담고 있지만, 각기 다른 종류의 이해도와 숙련도를 요구합니다. 예를 들어, 카페인의 스케치를 보면 화학자는 이것이 신체의 졸음 신호인 아데노신(adenosine)과 닮았다는 것을 알아채고, 수용체를 차단하여 우리를 각성시킨다고 예측할 수 있습니다. 하지만 그와 동일한 스케치만으로는 화학자가 거의 동일하게 생긴 다른 분자들과 카페인을 구별할 수는 없습니다.

화학자가 어떤 분자를 다루고 있는지 이해하는 것은 매우 중요합니다. 화학은 우리가 섭취하는 음식과 의약품부터 로션, 페인트, 플라스틱에 이르기까지 모든 것의 기초가 됩니다. 동일한 원자들 사이의 결합을 일부만 바꿔도 포도당은 과당이 되어, 같은 화학식을 공유하지만 완전히 다른 대사 경로를 거치게 됩니다. 분자를 거울상으로 뒤집으면, 살리도마이드(thalidomide) 참사에서 발생했던 것처럼 수면제가 기형유발물질이 될 수도 있습니다. 1

화학자의 일상적인 업무는 주어진 작업에 맞는 표현 방식에 상관없이 이러한 신호를 올바르게 읽는 것에 달려 있습니다. 표현 방식 간의 변환(그림에서 구조를 파악하고, 기기 측정 결과와 예상 생성물을 비교하고, 올바른 표기법으로 데이터베이스를 쿼리하는 작업)은 시간이 많이 걸리며 대규모로 유지하기란 불가능합니다. 가장 큰 화학 등록 데이터베이스인 CAS(CAS)는 2억 9천만 개 이상의 알려진 물질을 카탈로그화하고 있으며, 매일 약 15,000개의 새로운 물질이 추가됩니다.

AI는 이러한 연구의 부담을 덜어줄 수 있는 적합한 위치에 있지만, 화학 분야에서는 여전히 큰 기대치에 머물러 있는 실정입니다. 머신러닝 도구는 수년간 역합성(retrosynthesis - 목표 분자에서 더 간단한 전구체로 거슬러 올라가 합성 계획을 세우는 과정), 반응 예측 및 특성 추정에 혁명을 일으킬 것으로 주목받았습니다. 하지만 이러한 도구에 필요한 데이터를 얻기는 어려웠습니다. 부정 결과(null-results)에 대한 데이터가 부족하고, 형식이 일관되지 않으며, 구독형 학술지의 유료 벽과 비정형적인 부자료(supporting information) 뒤에 잠겨 있기 때문입니다. 역합성이 대표적인 예입니다. 유능한 AI 도구가 수년 전부터 존재했지만 채택은 일관되지 않았으며, 평균적인 학계나 소규모 연구실 화학자는 여전히 이를 사용하지 않습니다.

그럼에도 불구하고 AI의 발전은 마침내 화학 분야에 도달하고 있습니다. 오늘날의 최첨단 모델은 멀티모달이며 명시적인 추론이 가능합니다. 이제 AI는 미리 큐레이션된 분자 데이터베이스에 의존하는 대신, 학술지 도면이나 손그림에서 직접 화학 구조를 읽을 수 있습니다. 또한 실제 출판된 형태 그대로 실험 방법 섹션이나 부자료의 세부 내용을 읽어낼 수 있습니다. 단계별로 추론 과정을 보여줄 수도 있으므로 화학자는 AI의 출력 결과를 직접 검증할 수 있습니다. 이 모든 것이 현장에서 수년간 제기되어 온 데이터 문제를 해결해 주지는 않지만, 이러한 한계에도 불구하고 어떤 문제를 풀 수 있는지를 변화시키고 있습니다.

궁극적으로 저희의 주장은 소박합니다. 클로드는 화학자들의 판단을 보완하는 일상적인 번역, 리콜 및 통합 작업을 의미 있게 지원하기 시작했으며, 저희는 그 유용성을 계속 확장해 나갈 계획입니다. 오늘 저희는 이 작업을 가속화하기 위한 노력의 일환으로 첫 번째 백서를 발행합니다. 이 백서는 화학자들의 가장 일반적인 분석 입력 데이터인 NMR 스펙트럼을 다룹니다.

클로드 vs ChemDraw: NMR 예측 및 구조 결정 전체 버전은 여기에서 확인할 수 있습니다.

거의 모든 소분자(의약품, 살충제, 염료, 향료, 중합체, DNA 또는 단백질 하위 단위, 기능성 무기물 및 고체 재료)는 화학자가 그 구조를 결정했기 때문에 존재할 수 있습니다. 이러한 분자는 육안으로 볼 수 없기...

원문 보기
원문 보기 (영어)
Science Making Claude a chemist Jun 5, 2026 We’re working with world-class synthetic, computational, and analytical chemists to make Claude better at chemistry. In this post, we share our first work as part of this effort, in which Anthropic chemist, David Kamber, examines how Claude performs on a chemist’s most common analytical input, an NMR spectrum. When working with molecules, chemists move between hand-drawn structures on a whiteboard, instrument readouts, database query strings, and the technical notations of patents and publications. Each of these representations encodes the same underlying chemistry, but each demands a different kind of fluency. A sketch of caffeine, for example, allows a chemist to spot its resemblance to adenosine, the body’s drowsiness signal, and predict that it keeps us alert by blocking the receptor. However, that same sketch cannot help a chemist tell it apart from other near-identical looking molecules. Understanding what molecule a chemist is working with is critical. Chemistry undergirds everything from the foods and medicine we ingest to our lotions, paints, and plastics. Reroute a handful of bonds among the same atoms, and glucose becomes fructose, molecules sharing a formula but processed through entirely different metabolic pathways. Flip a molecule into its mirror image, and a sedative becomes a teratogen, as happened in the thalidomide disaster. 1 Chemists’ everyday work depends on reading these signals correctly across whichever representation befits a given task. Translating between these representations (chasing down a structure from a figure, reconciling an instrument readout against a proposed product, querying a database in the right notation) is time consuming and impossible to keep up with at scale—CAS, the largest chemistry registry, catalogs over 290 million disclosed substances and grows by roughly 15,000 new ones every day. AI is well-positioned to take on this research burden, yet it still remains largely aspirational in the context of chemistry. Machine-learning tools have been positioned for years as transformative for retrosynthesis—the process of working backward from a target molecule to simpler precursors to plan how to build it—reaction prediction, and property estimation, but the data those tools need have been hard to come by—sparse on null-results, inconsistent in format, and locked behind paywalls at subscription journals (and in unstructured supporting information). Retrosynthesis is a case in point—capable AI tools have existed for years, but adoption is uneven, and the average academic or small-lab chemist still doesn't use them. Even so, advancements in AI are finally reaching chemistry. Today’s frontier models are multimodal, and capable of explicit reasoning. They can read a chemical structure directly from a journal figure or hand sketch rather than depending on a pre-curated molecular database. And they can read the experimental detail of a methods section or supporting information in the form it is actually published. They can also show their reasoning step by step, which means a chemist can audit the outputs. None of this eliminates the data problem the field has been describing for years, but it changes which problems are tractable despite it. Ultimately, our claim is a modest one: Claude is starting to meaningfully assist chemists with the daily translation, recall, and integration work that complements their judgment, and we plan to keep extending its helpfulness. Today we are publishing the first white paper in the effort to accelerate this work. It tackles a chemist's most common analytical input: an NMR spectrum. Claude vs. ChemDraw on NMR prediction and structure elucidation Full version can be found here Nearly every small molecule—drug, pesticide, dye, fragrance, polymer, DNA or protein subunit, and functional inorganic or solid-state material—exists because a chemist determined its structure. Given that these molecules cannot be seen with microscopes, chemists must rely on spectral analysis, probing a molecule with light, radio waves, or magnetic fields. The way a given molecule absorbs, emits, or deflects this energy gives chemists a pattern, or spectrum, with which they can elucidate its structure. NMR spectroscopy—one of the canonical techniques chemists rely on for this—is one of the most time-consuming steps in synthetic chemistry; for every compound, a chemist has to match each peak in the spectrum to an atom in the proposed structure by hand. For this white paper, we tested how Claude fared against the dedicated NMR software chemists rely on today. We measured three Claude models (Opus 4.7, Opus 4.6, Sonnet 4.6) against ChemDraw and MestReNova on 20 compounds drawn from synthetic chemistry preprints published after the models’ training cutoff so as to avoid selection bias. Both ChemDraw and MestReNova do forward prediction, using a drawn structure to simulate what NMR spectrum will be produced. In addition to forward prediction, we also wanted to see whether Claude could go the other direction—starting from an experimental spectrum and proposing the structure behind it. This is the harder task, and the one existing software currently leaves to the chemist. To set up our assessment, we pulled 20 compounds from ChemRxiv preprints 2 posted after the models’ training cutoff, taking the first fully characterized novel molecules from each paper. The 20 span four structural families, five compounds each, with each family selected because it involves a different category of NMR challenge. Each tool was given the structure encoded as a SMILES string—the line-of-text notation chemists use to input a molecule to software—and was asked to predict where every hydrogen and carbon peak would fall along a 1D NMR spectrum (a horizontal axis measuring chemical shifts in ppm, parts per million). Given that NMR samples are dissolved in a liquid, and that the choice of solvent (chloroform, DMSO, etc.) moves the peak positions slightly, each tool was told to predict the spectrum in whatever solvent the chemists had used in the published paper. Because a language model’s output varies between runs, each Claude model was queried three times per compound and averaged; ChemDraw and MestReNova return the same answer every time and were run once. We then paired each predicted peak with its experimental counterpart and measured the gap in ppm. These landed within the window a chemist would call correct—±0.20 ppm for hydrogen or ±1.0 ppm for carbon. On hydrogen, Opus 4.7 was most accurate, with an average error of ±0.079 ppm—well under half the tolerance window—and the highest share of peaks landing inside it. On carbon, Opus 4.7 and MestReNova were effectively tied, at ±1.37 and ±1.48 ppm; the remaining tools kept the same rank order on both elements. Opus 4.6 was predictably middling, and Sonnet 4.6 was the weakest. The gap between them was most evident on a single notoriously difficult hydrogen—an NH proton in the chloropyridazine family whose true position falls in a narrow band between 6.8 and 7.9 ppm. Opus 4.7 placed it slightly low but consistently so; Opus 4.6 scattered its guesses across several ppm; Sonnet 4.6 put it in the 10–13 range, well outside where it actually appears. While Opus 4.7 performed fairly comparably to ChemDraw and MestReNova, the gap was wider on predicting the shape taken by a hydrogen’s NMR peak and how far apart the peaks sit, features which also contain structural information a chemist reads alongside position. Opus 4.7 matched the experimentally reported splitting pattern more often than any other tool, and all three Claude models predicted the sub-peak spacing to within half a hertz roughly 80% of the time—against 26 to 35% for ChemDraw and MestReNova. Opus 4.7 was also the most consistent across its three repeat runs: its average error varied less from run to run than the margin separating it from the next-best tool. From there, we evaluate