런던의 AI 스타트업 Basecamp Research가 엔비디아와 앤스로픽의 Anthology Fund 등으로부터 1억 4천만 달러를 투자받았다. 이 회사는 전 세계 환경 샘플에서 수집한 미생물 DNA로 EDEN이라는 AI 모델을 학습시켜 새로운 항생제와 세포 치료용 유전자 삽입 도구를 설계한다. 초기 실험실 연구와 마우스 실험에서 다약제내성균에 효과를 보였지만, 인체 임상시험을 통한 안전성·유효성 검증은 아직 남아 있다.
번역된 본문
Basecamp Research는 엔비디아와 앤스로픽의 Anthology Fund 등 투자자들로부터 1억 4천만 달러를 조달했다. 이 런던 기반 기업은 열대우림, 해양, 온천 등에서 채취한 유전 물질로 AI 모델을 학습시켜 항생제와 세포 치료용 도구를 설계한다. THE DECODER와의 인터뷰에서 CTO 필립 로렌츠(Philip Lorenz)는 생물학이 언어보다 훨씬 더 큰 AI 문제인 이유, 그리고 서류상 좋은 점수가 좋은 분자를 보장하지 않는 이유를 설명했다.
회사에 따르면 이번 투자 라운드는 S32가 주도했으며, 엔비디아, 앤스로픽의 Anthology Fund, NATO 혁신 기금, Redalpine 등이 참여했다. 2020년에 설립된 Basecamp는 이 자금을 생물학 AI 모델인 EDEN을 지속적으로 개발하고 자체 치료 후보 물질을 임상 개발 단계로 추진하는 데 사용할 계획이다. 세포 치료가 우선이다. 세포 치료는 환자의 세포를 유전적으로 변형해 암과 싸울 수 있도록 하는 치료법으로, Basecamp는 이러한 변형을 체내에서 직접 수행하고자 한다.
이를 위해 회사는 연구자들이 거의 연구하지 않은 서식지의 미생물에 주목한다. 모델은 이 미생물들의 게놈에서 의학적으로 활용 가능한 패턴을 학습하게 된다.
"생물학의 복잡성을 고려하면, 우리는 지금보다 몇 자릿수 더 많은 데이터가 필요합니다"라고 Basecamp의 CTO 필립 로렌츠는 THE DECODER에 말했다. 그에게 가장 중요한 것은 이 대규모 데이터 수집을 소규모의 정밀한 표적 실험과 결합하는 일이라고 한다. 물론 이것이 실제 치료법이 되기까지는 아직 많은 개발 작업이 남아 있다. 현재까지 Basecamp가 확보한 것은 실험실 결과와 마우스 실험 데이터뿐이다. 이것만으로는 이 접근 방식이 인간에게 안전하고 효과적인 치료법을 만들어낼 수 있는지 알 수 없다.
공공 게놈 데이터베이스가 생물학 AI를 감당하지 못하는 이유
월스트리트 저널은 최근 제약회사들이 AI에 수십억 달러를 쏟아붓고 있지만, 임상 개발의 성공률이 눈에 띄게 높아졌다는 확실한 증거는 여전히 드물다고 보도했다. 로렌츠는 이것이 기술을 의심할 이유가 된다고 보지 않는다. 스케일링은 언어 모델에 매우 빠르게 큰 성과를 가져다주었다. 생물학도 진전하고 있지만, 임상시험 가속화 같은 큰 미해결 문제는 훨씬 더 어렵다고 그는 말한다. 문제로서 생물학은 "훨씬, 훨씬, 훨씬 더 복잡하고 훨씬 더 크다."
데이터도 근본적으로 다르다. 언어 모델은 방대한 텍스트 컬렉션으로 학습하지만, 제약회사는 종종 개별 연구에서 나온 작고 분절된 데이터셋을 다룬다.
로렌츠는 Epoch AI의 추정치를 인용해 이 격차가 얼마나 큰지 보여준다. 존재하는 모든 단어의 상한선은 대략 5조(quadrillion) 토큰으로 추산된다. 반면 Basecamp는 지구상의 뉴클레오타드 수를 10의 37제곱으로 추정한다. "10의 37제곱 장의 포커 카드를 쌓으면, 그 카드 더미는 관측 가능한 우주를 백만 번 둘러쌀 것입니다"라고 로렌츠는 말한다. 그만큼 큰 데이터셋이 필요한 것은 아니지만, 현재 존재하는 것보다는 훨씬 큰 데이터셋이 필요하다고 한다.
공공 게놈 데이터베이스 역시 자연의 불완전한 모습만을 보여준다. Basecamp의 BaseData 데이터베이스 연구 논문에 따르면, Sequence Read Archive의 염기서열 볼륨 중 약 68%가 단 5개 종에서 나온 것이며, 인간만이 전체의 약 54%를 차지한다. 이러한 편중은...
Inside Basecamp Research, the AI startup turning evolution into training data Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 23, 2026 Basecamp Research / GPT-Image-2 prompted by THE DECODER Key Points London-based Basecamp Research has raised $140 million from investors including Nvidia and Anthropic's Anthology Fund to build AI models for developing drugs and cell therapies. The models, called EDEN, train on DNA from microorganisms collected in environmental samples around the world. Basecamp uses them to design new antibiotics and biological tools that insert genes at targeted spots inside the body. Early lab work and mouse studies show that individual drug candidates work against multidrug-resistant bacteria. Human clinical trials still have to prove the therapies are safe and effective. Ask about this article… Search Basecamp Research has raised $140 million from investors including Nvidia and Anthropic's Anthology Fund. The London company trains AI models on genetic material from rainforests, oceans, and hot springs to design antibiotics and tools for cell therapies. In an interview with THE DECODER, CTO Philip Lorenz explains why biology is a far bigger problem for AI than language, and why good scores on paper don't guarantee good molecules. According to the company, investor S32 led the round. Nvidia, Anthropic's Anthology Fund, the NATO Innovation Fund, and Redalpine also took part, among others. Founded in 2020, Basecamp plans to use the money to keep developing its biological AI models, called EDEN , and to push its own therapy candidates toward clinical development. Cell therapies come first. These treatments genetically modify a patient's cells so they can fight cancer, for example. Basecamp wants to make that change directly inside the body. Ad To get there, the company looks to microorganisms from habitats researchers have barely studied. The models are supposed to learn patterns from their genomes that can be put to medical use. Ad "Given the complexity of biology, we just need orders of magnitude more data," Philip Lorenz, Basecamp's CTO, tells THE DECODER. What matters most, he says, is pairing that massive data collection with smaller, targeted experiments. There's still a lot of development work ahead before any of this becomes a treatment. So far, Basecamp has lab results and data from mice. That doesn't show whether its approach can produce safe and effective therapies in humans. Ad Why public genome databases can't carry biological AI The Wall Street Journal recently reported that pharma companies are pouring billions into AI, but solid evidence of noticeably higher success rates in clinical development is still scarce . Lorenz doesn't see that as a reason to doubt the technology. Scaling delivered big gains for language models very quickly. Biology is making progress too, he says, but the big open questions, like speeding up clinical trials, are far harder. As a problem, biology is "way, way, way more complicated and way larger." The data is also fundamentally different. Language models learn from huge text collections, while pharma companies often work with small, siloed datasets from individual studies. Ad Lorenz uses an estimate from Epoch AI to show how wide the gap is. It puts the upper limit of all existing words at roughly five quadrillion tokens. Basecamp estimates the number of nucleotides on Earth at 10 to the power of 37. "If you took a stack of cards, like poker cards, of 10 to the power of 37 cards, that stack would surround the observable universe a million times," Lorenz says. You don't need a dataset that big, he says, but you do need one far bigger than what exists today. Ad Public genome databases also paint an incomplete picture of nature. According to Basecamp's research paper on its BaseData database , about 68 percent of the sequence volume in the Sequence Read Archive comes from just five species. Humans alone account for about 54 percent. That focus makes sense for medical research. For a model meant to learn as many different biological processes as possible, though, Basecamp sees it as a limitation, since many other life forms barely show up in the data. "If you were to train LLM only on newspaper articles from 1975, it would be a really, really bad model," Lorenz said in a case study from Microsoft , the company's cloud partner. "That is kind of where we are in biology." So Basecamp collects its own samples with local research partners, from rainforest soil, volcanic ground, and the deep sea off Antarctica. According to the latest funding announcement, the network now spans more than 30 countries and all seven continents. Microsoft puts the number of participating organizations at 208 across 31 countries. At the time of the interview, the dataset held about 15 trillion tokens, according to Lorenz. That already puts it in the same range as the text datasets used to train models like Claude or GPT. Here, though, a token is a single DNA building block. The AI doesn't process words like a chatbot does. It processes strings of characters that describe genetic material. According to Microsoft, the first EDEN generation trained on 9.7 trillion DNA building blocks from more than a million newly sequenced species. Basecamp ran the training with Microsoft researchers on Azure, and the compute reportedly matched GPT-4's. OpenAI has never officially disclosed that figure, however. Over the next year and a half, the dataset is supposed to grow roughly a hundredfold, passing one quadrillion tokens. The framework for that is the Trillion Gene Atlas announced in March, in which Basecamp, Anthropic, Nvidia, PacBio, and Ultima Genomics plan to assemble genetic data on the scale of one trillion genes. In its latest funding announcement, Basecamp already calls the Trillion Gene Atlas the training foundation for its EDEN models. It doesn't say how far the planned expansion has gotten. Bacterial arms races offer a blueprint for new drugs Lorenz sees evolution, which drives all biological processes, as the link between environmental samples and medicine. Whether a bacterium in a hot spring adapts to heat or a cancer cell spreads, similar selective pressures are at work. The genomes of humans, mammals, and nearly all vertebrates are largely sequenced, while the rest of biodiversity has barely been cataloged. Many drugs also originally came from plants, bacteria, and fungi. He points to competing microorganisms as a concrete example. Some of the most powerful therapeutic molecules come from "biological warfare" between organisms, he says. When a bacterial species tries to dominate an ecosystem, it sometimes develops substances that inhibit or kill rival species, securing food and habitat for itself. This happens billions of times in countless places on Earth, especially where nutrients are scarce. In medicine, such substances can be useful as antibiotics. In the Microsoft case study, Lorenz names phages as another example. These viruses infect bacteria and inject their DNA into the cells. That back-and-forth produced tools that can be used for gene editing. Basecamp wants to learn from a wide range of these evolved solutions. The models aren't just supposed to rediscover known compounds. They're meant to design new candidates from the patterns they've learned. To do this, the company records the chemical, physical, and ecological conditions at each site along with the genetic material. It also focuses on long, continuous stretches of DNA. These show which genes sit next to each other and might work together. Short, isolated fragments often lose that context. An AI-designed antibiotic clears its first animal test Lorenz sees EDEN's antibiotic designs as early proof of medical value. "We prompt on a pathogen and then the model designs an antibiotic that kills it," he says. In a research paper on the EDEN model family that hasn't been peer-reviewed yet, the authors report on a small, curated set of