Anthropic이 Claude 모델(Mythos Preview, Opus 4.8)을 활용해 신약 개발 초기 단계인 미니바인더 설계 실험을 진행했고, 26.8%의 히트율로 업계 평균(10~15%)을 웃돌았다고 발표했습니다. Claude는 자체 단백질 모델 없이 기존 오픈소스 도구들을 직접 설치·구동하며 설계부터 검증까지 수행했으며, 독립적인 검증은 아직 진행 중입니다.
번역된 본문
Anthropic은 모든 연구실이 이제 언어모델 에이전트가 단백질 설치 전체 스택을 운영하게 할 수 있다고 밝혔습니다.
두 차례의 실험에서 Anthropic은 자사 Claude 모델을 신약 개발 초기 단계 작업에 투입했습니다. 회사에 따르면 단백질 설계 결과는 업계 통상 히트율을 능가했습니다. 다만 결과에 대한 독립적 검토는 아직 이뤄지지 않았습니다.
Anthropic은 Claude 모델이 신약 개발 초기 단계 작업을 수행한 두 건의 실험을 발표했습니다. 첫 번째 실험은 미니바인더(minibinder) 설계에 초점을 맞췄습니다. 미니바인더는 표적 단백질에 단단히 결합해 그 기능을 차단하거나 변경하는 작은 단백질로, 이 원리는 많은 약물의 기반이 됩니다. 자연에서 탐색하는 대신 컴퓨터로 처음부터 이러한 바인더를 설계하는 것을 '데 노보(de novo) 설계'라고 합니다. 기술 보고서에 따르면 이 작업은 여러 전문가적 판단과 함께 전문 소프트웨어 및 컴퓨팅 자원을 며칠간 조율하는 과정이 필요합니다.
Mythos Preview와 Opus 4.8 모델은 16개 표적 단백질에 대한 바인더를 설계했으며, 이 중 15개에서 측정 가능한 결과가 나왔습니다. Claude는 15개 중 14개 표적에서 성공했습니다. 실험실에서 테스트된 1,320개 설계 중 354개가 실제로 표적에 결합해 26.8%의 히트율을 기록했습니다. Claude가 자체 목록에서 1위로 평가한 설계만 보면 49%가 결합했습니다.
모든 표적을 48시간 내에 동시에 처리하는 멀티타깃 모드에서 모델은 26.7%(Mythos Preview)과 22.6%(Opus 4.8)의 히트율을 달성했습니다. Mythos Preview가 각 표적을 개별적으로 다룰 때는 히트율이 35.1%로 상승했지만, 표적당 컴퓨팅 예산은 2.8배나 들었습니다. 저자들이 인정하듯 집중도와 예산은 분리할 수 없습니다. 비교를 위해 Anthropic은 proteinbase.com 데이터베이스의 공개 문서화된 캠페인에서 도출된 현재 업계의 전형적인 범위인 10~15%를 인용했습니다.
언어모델이 12개의 전문 도구를 운영
Anthropic은 자체 단백질 모델을 만들지 않았습니다. Claude는 이 분야에서 이미 사용되는 오픈소스 전문 소프트웨어만을 설치하고 실행했습니다. 단백질의 공간 구조 골격(백본)은 PXDesign(358개 설계), RFdiffusion3(267개), Genie 3(185개), FreeBindCraft(135개), BoltzGen(134개), RFdiffusion(118개), Proteina-Complexa(100개) 등에서 나왔습니다. 이 골격을 형성하는 아미노산 사슬은 대부분 ProteinMPNN의 변형인 SolubleMPNN이 계산했습니다. 후보를 필터링하고 순위를 매기기 위해 Claude는 ESMFold2, ESMFold2-Fast, Protenix v2를 혼합해 사용했습니다. 이 프로그램들은 바인더와 표적이 함께 어떻게 접히는지 예측하고, 결합이 작동할지에 대한 신뢰도 점수도 제공합니다. AlphaFold-3 가중치, Rosetta, ESM3는 라이선스 문제로 제외되었습니다.
설정에는 모든 에이전트의 시스템 프롬프트로 실행된 약 16,000 단어 분량의 프로토콜 프롬프트가 포함됐습니다. 이 중 약 3분의 1만이 과학적 지침과 참고문헌 목록이고, 나머지 3분의 2는 일정 관리, 서브 에이전트 위임, 검증, 예산 규율에 관한 내용이었습니다. 프롬프트는 어떤 표적에 대해서도 단백질 표면의 어떤 부위를 공격할지(이른바 에피토프) 지정하지 않았습니다. 컴퓨팅 예산은 멀티타깃 캠페인당 5만 달러, 단일 표적당 1만 달러였으며, 클라우드 제공업체 Modal을 통해 실행됐습니다. 인간은 표적을 선정하고, 프롬프트를 작성하고, 합성을 주문하고, 측정 데이터를 해석했습니다. 그 사이에는 인프라 장애 후 재개를 위한 짧고 비기술적인 지시만 필요했다고 보고서는 전합니다.
동일한 검정 플레이트에서 대회 우승자를 능가
검증은 유료 계약 연구소인 Adaptyv Bio와 Twist Bioscience가 담당했습니다. 이들은 각 설계를 변형 없이 생물학적으로 제작하고, 표적에 결합하는지, 얼마나 단단히 결합하는지를 독립적으로 측정했습니다. 결합 강도는 나노몰(nM) 단위로 측정된 KD 값으로 표시되며, 숫자가 작을수록 결합이 강합니다. 약물이 더 낮은 용량으로 작동할 수 있기 때문에 작은 값이 중요합니다.
표적별 결과는 크게 달랸습니다. 면역 수용체 TREM2의 경우 90개 설계 중 72개가 결합했습니다. 성장 인자의 경우...
Anthropic says any lab can now let a language model agent run the whole protein design stack Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 19, 2026 Nano Banana Pro prompted by THE DECODER In two experiments, Anthropic put its Claude models to work on early-stage drug discovery. The company says its protein design results beat the usual industry hit rates. An independent review of the results is still pending. Anthropic has published two experiments in which its Claude models took on tasks from the early phase of drug development. The first focused on designing so-called minibinders , small proteins that lock tightly onto a target protein and block or change its function. This principle underlies many drugs. Designing such binders from scratch on a computer, rather than searching for them in nature, is called de novo design. According to the technical report, it still takes a series of expert decisions plus days of orchestrating specialized software and compute. The models Mythos Preview and Opus 4.8 designed binders against 16 target proteins, 15 of which produced usable measurements. Claude succeeded on 14 of those 15 targets. Of 1,320 designs tested in the lab, 354 actually bound to their target, a hit rate of 26.8 percent. Looking only at the designs Claude ranked first on its own list, 49 percent bound. In multi-target mode, where all targets were handled at once within 48 hours, the models reached 26.7 percent (Mythos Preview) and 22.6 percent (Opus 4.8). When Mythos Preview worked each target on its own, the rate rose to 35.1 percent, but with 2.8 times the compute budget per target. Focus and budget can't be separated, as the authors acknowledge. For comparison, Anthropic cites today's typical range of 10 to 15 percent, drawn from publicly documented campaigns in the proteinbase.com database. A language model runs a dozen specialized tools Anthropic didn't build its own protein model. Claude installed and ran only open-source specialty software that the field already uses. The spatial scaffolds of the proteins, known in the jargon as backbones, came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), RFdiffusion (118), and Proteina-Complexa (100), among others. The amino acid chain that eventually forms this scaffold was mostly computed by SolubleMPNN, a variant of ProteinMPNN. To filter and rank the candidates, Claude used a mix of ESMFold2, ESMFold2-Fast, and Protenix v2. These programs predict how binder and target fold together. They also give a confidence score for whether the binding should work. AlphaFold-3 weights, Rosetta, and ESM3 were excluded for licensing reasons. The setup included a protocol prompt of about 16,000 words that ran as the system prompt in every agent. Only about a third of it is scientific guidance plus a reading list. The other two thirds cover scheduling, delegation to sub-agents, verification, and budget discipline. The prompt didn't specify which spot on the protein surface to attack, the so-called epitope, for any target. The compute budget was $50,000 per multi-target campaign and $10,000 per single target, run through the cloud provider Modal. Humans picked the targets, wrote the prompt, ordered the synthesis, and interpreted the measurement data. In between, the report says, only short, non-technical instructions were needed to resume after infrastructure outages. A contest winner beaten on the same assay plate Validation was handled by the paid contract labs Adaptyv Bio and Twist Bioscience . They produced each design biologically, unchanged, and measured independently whether it binds to its target and how tightly. Binding strength is given as a KD value, measured in nanomolar (nM). The smaller the number, the tighter the binding. Small values matter for a drug because it then works at a lower dose. Results varied wildly by target. For the immune receptor TREM2, 72 of 90 designs bound. For the growth factor VEGF-A, 54 of 90. The most telling case is RBX1, part of an enzyme complex that controls the targeted breakdown of other proteins in the cell. In an open design contest, only 9 of 245 newly designed candidates bound there. With Claude, it was 28 of 90. Anthropic had the contest's winning design rebuilt and checked on the same assay plate. It bound at 45 nM, while Claude's best design bound at 3.9 nM, roughly ten times tighter. TNFα is considered especially hard. It's an immune system messenger that triggers inflammation and is the target of five approved biotech drugs, including Humira. Several earlier design approaches reported zero hits there. Claude produced 12 binders from 150 designs, all from Opus 4.8, none from Mythos Preview. But those twelve trace back to just four different scaffolds, so they're less independent of each other than the number suggests. Binding to animal versions of the target proteins was only a secondary goal in the prompt, yet 130 of 233 tested binders also bound the mouse counterpart of their target. That's practically relevant because drugs are tested in animals before humans. Two targets showed clear limits. BBF-14 is a computer-invented, barrel-shaped protein that doesn't occur in nature and therefore has no evolutionary history a design method could draw on. The bacterial maltose-binding protein MBP, in turn, has a smooth, water-loving surface that gives a binder little to grab onto. Against BBF-14, only three weakly binding designs worked, and against MBP, none of 90. Anthropic says the confidence scores from the folding prediction warned of neither failure. The designs against MBP and BBF-14 got almost the same scores as those against successful targets. Chemistry analysis in under 25 minutes In the second experiment, Opus 5 interpreted raw data from two standard measurements from a contract lab. Nuclear magnetic resonance spectroscopy (NMR) shows chemists whether they actually made the substance they meant to make. Liquid chromatography with mass spectrometry (LC-MS) reveals how pure the sample is. Both instruments output files in proprietary formats that are normally analyzed by hand in the manufacturer's software. From these raw files and prompts of one to three sentences, Claude delivered results in 23 and 19 minutes. For the LC-MS file, the model found no suitable reader program and decoded the format itself. As a cross-check, it exactly reproduced the summary values the instrument had stored for all 2,664 measurement points. In the NMR analysis, Claude proposed the same follow-up experiment the lab had independently run three days after the first measurement, then corrected its own mistake. It first reported four missing signals, but the internal check found two. The much-cited purity values of 96.4 versus 96.33 percent rest on different baselines, according to the report. Calculated by the lab's method, Claude would land at 98.8 percent. What's actually new here Predicting protein structures and generating new proteins has been established since AlphaFold , RFdiffusion , and BindCraft. The actual design work here was done by these tools too. What's new is the layer above them. A general language model researched the biology of each target, chose the docking site, installed the programs itself from their public code repositories, combined them across 24 different workflows, and delivered a finished ranking without any human touching a single design decision. Since all the models used are open-source, the report argues, such campaigns are within reach for any lab. The authors spell out the limits clearly themselves: There was no parallel campaign by human experts as a control, and they don't claim Claude's designs are better than what specialists would achieve with the same tools and budget. For four of the six contest targets, the contest results were in Claude's reading list. The only thing measured was whether the designs bind, not their actual spatial shape and not their biologica