메뉴
BL
The Decoder • 16일 전

딥마인드, 유전체 모든 변이 예측하는 '알파게놈 아틀라스' 공개

IMP
9/10
핵심 요약

구글 딥마인드가 인간 유전체에서 가능한 약 90억 개의 단일 염기 변이가人体에 미칠 영향을 AI로 예측한 '알파게놈 아틀라스'를 공개했습니다. 1페타바이트 규모의 이 데이터셋은 단백질 구조 데이터베이스인 알파폴드보다 30배 이상 크며, 수십만 개에 달하는 예측값을 단일 점수(AVI)로 요약해 질병 관련 변이를 신속히 선별할 수 있게 합니다. 실제로 진단되지 않던 소아 간질 사례에서 원인 변이를 찾아내는 성과를 거두었습니다.

번역된 본문

딥마인드, 인간 유전체의 모든 가능한 DNA 변이를 매핑한 알파게놈 아틀라스 공개

막시밀리안 슈라이너 (2026년 9월 9일)

구글 딥마인드가 인간 유전체에서 가능한 약 90억 개의 단일 염기(글자) 변이 각각이 몸속에서 어떤 영향을 미칠 가능성이 있는지 예측해냈습니다. 알파게놈 아틀라스(AlphaGenome Atlas)는 연구자들이 수많은 유전적 변이 중 소수의 의미 있는 변이를 찾아내도록 돕는 것을 목표로 합니다.

인간 유전체는 약 30억 개의 DNA 글자로 구성되어 있습니다. 모든 사람은 기준 서열과 비교해 수백만 개의 미세한 차이를 가지고 있으며, 대개 글자 하나가 바뀐 형태입니다. 이런 변이 대부분은 무해하지만 일부는 질병을 일으킵니다. 어떤 변이가 유해한지는 DNA만 보고 직접 알 수 없으며, 가능한 변이가 약 90억 개에 달하므로 이를 모두 실험실에서 검증하는 것은 사실상 불가능합니다.

새로 공개된 알파게놈 아틀라스는 예측으로 이 간극을 메우고자 합니다. 90억 개 변이 각각에 대해 수백 가지 세포 유형과 조직에서 해당 변이가 분자 과정에 미칠 것으로 예상되는 영향의 추정치를 제공합니다. 이 데이터셋은 1페타바이트 규모로, 단백질 구조 데이터베이스인 알파폴드(AlphaFold)보다 30배 이상 큽니다.

이는 2025년에 공개된 AI 모델 알파게놈(AlphaGenome)을 기반으로 합니다. 이 모델은 100만 글자 길이의 DNA 구간을 읽고, 유전자가 얼마나 강하게 발현되는지, 조절 단백질이 DNA에 결합할 수 있는지, 유전자의 전사체가 어떻게 스플라이싱되는지 예측합니다. 지금까지는 각 변이마다 모델에 개별적으로 질의해야 했지만, 이제는 답이 미리 계산되어 제공됩니다.

논문에 따르면 각 변이에는 평균 약 27,000개의 개별 예측값이 따라옵니다. 이는 단백질 설계도를 담고 있지 않은 유전체의 약 98%에 특히 중요합니다. 이러한 비부호화(논코딩) 영역은 유전자가 언제, 어떤 조직에서 활성화될지 결정하는 스위치와 다이얼처럼 작동합니다. 질병 관련 변이 대부분이 바로 이곳에 위치하며, 동시에 그 영향을 읽어내기가 가장 어려운 곳이기도 합니다.

돌연변이마다 하나의 숫자를 연구진은 변이당 수천 개의 예측값은 일상적인 사용에는 너무 많다고 말합니다. 그래서 딥마인드는 이를 단일 숫자로 압축한 알파게놈 변이 영향 점수(AVI, AlphaGenome Variant Impact Score)를 만들었습니다. 작은 신경망이 알파게놈 예측값에 단백질 모델 알파미센스(AlphaMissense)와 수백만 년의 진화 동안 DNA 부위가 얼마나 변하지 않았는지에 대한 두 가지 지표를 결합합니다. AVI는 18개의 입력 특성으로 작동하는 반면, 기존 벤치마크 도구인 CADD는 150개 이상을 사용합니다.

어떤 변이가 유해한지 확실히 알려진 경우는 거의 없습니다. 연구진은 이를 우회했습니다. 집단 내에서 매우 희귀한 변이는 유해할 가능성이 높은 것으로, 흔한 변이는 무해할 가능성이 높은 것으로 간주했습니다. 유해한 돌연변이는 세대를 거치며 퍼지는 경우가 적기 때문입니다. 논문에 따르면 이러한 간접적 학습에도 불구하고 AVI는 이미 임상적으로 분류된 변이에 대한 테스트에서 특히 비부호화 영역에서 기존 도구들을 능가했습니다. 일부 과제에서는 경쟁 도구가 앞서기도 했습니다. 아틀라스는 각 변이의 점수를 어떤 과정이 주도하는지, 예를 들어 유전자 전사체의 스플라이싱인지 스위치인지 등도 세분화해 보여줍니다.

간질 사례가 보여주는 성과 해결되지 않은 희귀질환을 연구하는 GREGoR 컨소시엄의 사례가 이것이 실제로 어떻게 도움이 되는지 보여줍니다. 중증 간질을 앓는 아이가 유전체 시퀀싱에도 불구하고 진단을 받지 못했습니다. AVI는 이전에 '불명확'으로 분류되었던 DNM1 유전자의 변이를 후보 목록 최상위로 올렸습니다. 알파게놈 예측은 메커니즘도 제공했습니다. 해당 변이는 유전자 전사체 처리 과정에서 잘못된 스플라이싱 부위를 만들어 단백질을 아미노산 13개만큼 길어지게 합니다. 그러나 이는 뇌에서만 발현되는 유전자 버전에서만 일어나기 때문에 이전의 혈액 검체 연구에서는 아무것도 발견되지 않았던 것입니다. 실험실 실험이 예측을 확인했으며, 연구진은 이 변이를 '질병 원인 가능성 있음'으로 분류할 것을 권고했습니다. 컨소시엄이 이미 해결한 사례들을 돌아본 결과, AVI는 29.5%의 사례에서 원인 변이를 상위 50개 후보 안에 포함시켰습니다.

원문 보기
원문 보기 (영어)
Deepmind's AlphaGenome Atlas maps every possible DNA change in the human genome Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 9, 2026 Google Google Deepmind has predicted what each of the roughly nine billion possible single-letter changes in the human genome would likely do inside the body. The AlphaGenome Atlas aims to help researchers spot the few relevant variants among a flood of genetic ones. The human genome runs to about three billion DNA letters. Every person carries millions of tiny deviations from the reference sequence, usually a single swapped letter. Most of these variants are harmless. A few cause disease. Which ones can't be read directly from the DNA, and testing every one in the lab is basically impossible when there are roughly nine billion possible swaps. The newly released AlphaGenome Atlas tries to fill that gap with predictions. For each of those nine billion changes, it offers an estimate of how the change would likely affect molecular processes across hundreds of cell types and tissues. The dataset spans one petabyte, making it more than 30 times the size of the AlphaFold database for protein structures. It builds on the AI model AlphaGenome , introduced in 2025. The model reads DNA stretches one million letters long and predicts how strongly a gene gets read, whether regulatory proteins can bind to the DNA, and how a gene's transcript gets spliced. Until now, the model had to be queried for each variant one at a time. Now the answers are precomputed. According to the paper , each variant comes with about 27,000 individual prediction values on average. That matters most for the roughly 98 percent of the genome that holds no blueprints for proteins. These noncoding regions act like switches and dials that decide when and in which tissue a gene is active. That's where most disease-linked variants sit, and it's also where their effects have been hardest to read. One number for every mutation Thousands of prediction values per variant are too much for everyday use, the team says. So Deepmind built the AlphaGenome Variant Impact Score (AVI), which boils it all down to a single number. A small neural network combines the AlphaGenome predictions with the protein model AlphaMissense and two measures of how unchanged a DNA site has stayed across millions of years of evolution. AVI works with 18 input features. The established benchmark tool CADD uses more than 150. For almost no variant is it known for sure whether it causes harm. The team worked around this. Variants that are very rare in the population are treated as likely harmful, common ones as likely harmless, because harmful mutations spread less often across generations. Despite this indirect training, AVI beat existing tools in tests on variants that had already been clinically classified, especially in noncoding regions, according to the paper. On some tasks, the competition edged ahead. The atlas also breaks down for each variant which process drives its score, such as whether the splicing of a gene's transcript or a switch is affected. An epilepsy case shows the payoff A case from the GREGoR consortium, which studies unsolved rare diseases, shows how this helps in practice. A child with severe epilepsy had gone without a diagnosis despite genome sequencing. AVI pushed a variant in the gene DNM1, previously classed as unclear, to the top of the candidate list. The AlphaGenome predictions also supplied the mechanism. The variant creates a wrong splice site during the processing of the gene's transcript, which lengthens the protein by 13 building blocks. But this happens only in a gene version that is read exclusively in the brain. That's why earlier work on blood samples had found nothing. A lab experiment confirmed the prediction, and the researchers recommend classifying it as likely disease-causing. Looking back at cases the consortium had already solved, AVI ranked the causal variant among the top 50 candidates in 29.5 percent of cases, compared with 12.5 percent for CADD. More signal in the noise The atlas is also meant to push population studies forward. To find out whether rare variants in a genome region affect something like a blood value, you have to analyze many of them together, because each one alone is too rare for statistics. If harmless and effective variants get mixed together, the signal disappears in the noise. Gareth Hawkes of the University of Exeter used the atlas to group only those variants predicted to act the same way, drawing on genome data from more than 54,000 UK Biobank participants. That turned up 22 percent more links between noncoding variants and protein levels in the blood than conventional filters did. From the predictions, the team also derived 2,601 recurring short DNA patterns, essentially the "words" of the genome where regulatory proteins latch on. A research tool, not a diagnosis AlphaGenome has limits too. It doesn't know every cell type, and it misses effects that work through the amount of other regulatory proteins. The atlas and AVI are research tools, Deepmind says, and can only be one link in the chain of evidence behind a diagnosis. The atlas is available for noncommercial use through a web portal , an API , and as a skill in Google Antigravity . A commercial version is set to follow through Google Cloud. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->