메뉴
BL
The Decoder • 28일 전

구글 딥마인드 AI 공동과학자, 실험 계획부터 논문 작성까지 수행

IMP
8/10
핵심 요약

구글 딥마인드가 멀티에이전트 시스템 'Co-Scientist'를 가설 생성기를 넘어 실험을 계획하고 실험실 장비를 직접 제어하며 과학 논문까지 작성하는 통합 연구 파트너로 확장했습니다. 재료과학, 생물학, 컴퓨터과학 세 분야에서 실험적으로 검증된 결과를 제시했지만, 벤치마크 점수와 실제 인간 평가 간 괴리 등 한계도 확인되었습니다.

번역된 본문

구글 딥마인드의 AI Co-Scientist, 이제 실험을 계획하고 실험실 장비를 작동하며 과학 논문까지 작성한다

구글 딥마인드는 멀티에이전트 시스템 'Co-Scientist'를 가설 생성기에서 실험실과 통합된 연구 파트너로 확장했다고 발표했다. 구글에 따르면 이 시스템은 세 개의 학문 분야에서 실험적으로 검증된 결과를 제출했다.

최신 Gemini 모델을 기반으로 구축된 Co-Scientist는 이제 단순히 가설을 생성하는 대신 실험을 계획하고, 코드를 작성하며, 실험실 장비를 제어한다. 기술적으로 새로운 점은 폐쇄 루프(closed-loop) 연구 워크플로우다. 시스템은 연구 질문에서 가설을 도출하고, 실험 계획과 프로그램 또는 기계 판독 가능한 실험 프로토콜을 작성하며, 결과를 분석하고 과학 논문 원고를 생성한다. 검증 모듈은 텍스트 내 수치 주장을 생성된 코드의 실행 로그와 대조해 조작된 결과를 줄인다.

구글은 2025년 2월 Co-Scientist를 처음 선보였으며, 당시에는 Gemini 2.0 기반이었고 팩트체크와 문헌 검토에 결함이 있었다. 확장된 시스템은 자율성을 높여가며 세 분야에서 검증되었다. 재료과학에서는 인간이 실행할 합성 레시피를 설계했고, 생물학에서는 전문가 피드백을 받아 예측 파이프라인을 구축했으며, 컴퓨터과학에서는 완전히 독자적으로 작업했다.

Co-Scientist, 합성·예측·설계를 수행하다

재료 합성을 위해 연구진은 Co-Scientist를 반자동 고온 furnace와 연결했다. 시스템은 주로 위험한 식각(etching) 공정으로만 생산되던 수요가 많은 2차원 물질에 대해 더 안전한 합성 경로를 찾았고, 해당 실험실 장비에 맞춘 완전한 성장 레시피를 생성했다. 인간의 다듬기를 거친 25차례의 시도 끝에 팀은 목표 물질과 유사한 특성을 지닌 층상 구조를 만들어냈지만, 원자 구조의 확정적 확인은 아직 남아 있다. 두 번째 실험에서는 세 개의 반도체 박막을 첫 시도에서 합성했다. Co-Scientist는 Gemini 3 Deep Think를 사용해 장비를 직접 제어하며 레시피 개발 시간을 며칠에서 몇 분으로 단축했다. 다만 시료와 전구체 물질은 여전히 인간이 수동으로 장전해야 했고, 고속 모드는 정교하게 최적화된 레시피보다 더 작고 균일성이 떨어지는 결정을 만들어냈다. 제1저자 사무엘 슈미트갈(Samuel Schmidgall)은 이 레시피가 다른 실험실로 이전 가능한지는 열린 문제라고 밝혔다.

생물학에서 Co-Scientist는 유전자 조작 대장균(E. coli) 집락이 서로 다른 화학 물질 농도에서 어떤 패턴을 형성하는지 예측하는 이미지 분석 파이프라인을 자율적으로 구축했다. Gemini 3 Pro Image로 생성된 예측은 네 가지 형태 특징 중 세 가지에서 미공개 실험 결과와 일치했다. 연구진은 이 시스템이 알려진 조건 사이에서만 추론할 수 있으며 완전히 새로운 시스템에서의 동작은 예측할 수 없다고 인정했다. 만약 가능하다면 훨씬 더 주목할 만한 일일 것이다.

컴퓨터과학 실험은 초기 설정 외에는 인간의 개입 없이 진행되었다. Co-Scientist는 들어오는 질의를 분류하고 수십 개의 응답 후보를 병렬로 생성한 뒤 이를 다듬는 의료 AI 아키텍처 'Agent_H'를 설계했다. 과도하게 긴 응답을 보정한 후 Agent_H는 GPT-5와 Claude Opus 5를 포함한 6개 프런티어 모델을 건강 벤치마크에서 능가했다. 그러나 벤치마크 결과는 인간 평가에서는 유지되지 않았다. 세 명의 board 인증 의사가 아홉 개 범주에서 응답을 평가한 결과, Agent_H는 기준 모델인 Gemini 3.1 Pro에 비해 단 하나의 범주, 즉 잠재적으로 유해한 응답의 위험이 낮다는 점에서만 통계적으로 유의한 우위를 보였다. 자동 벤치마크 평가자들의 판단도 의사들의 평가와 약한 상관관계만 보였다. 연구진은 높은 벤치마크 점수가 시스템이 임상적 관점에서 실제로 더 나은 답변을 제공한다는 의미는 아니라고 말하며, 이러한 벤치마크가 정확히 무엇을 측정하는지에 대한 의문을 제기했다.

신뢰성 모듈로 환각률을 4%까지 낮추다

LLM 기반 자율 연구 시스템의 핵심 문제는 (후략)

원문 보기
원문 보기 (영어)
Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 28, 2026 Nano Banana Pro prompted by THE DECODER Google Deepmind has expanded its multi-agent system Co-Scientist from a hypothesis generator into a lab-integrated research partner. According to Google, the system has delivered experimentally validated results across three disciplines. Built on current Gemini models, Co-Scientist now plans experiments, writes code, and controls lab equipment instead of just generating hypotheses. What's technically new is the closed-loop research workflow: the system derives hypotheses from a research question, creates experimental plans, programs, or machine-readable lab protocols, analyzes results, and generates scientific manuscripts. Verification modules cross-check numerical claims in the text against the execution logs of the generated code to cut down on fabricated results. Google first introduced Co-Scientist in February 2025 , then based on Gemini 2.0 and with shortcomings in fact-checking and literature review. The expanded system was validated across three disciplines with increasing autonomy. Co-Scientist designed synthesis recipes for humans to execute in materials science, built a prediction pipeline with expert feedback in biology, and worked entirely on its own in computer science. Co-Scientist synthesizes, predicts, and designs For material synthesis, the researchers paired Co-Scientist with a semi-automated high-temperature furnace. The system found a safer pathway for a sought-after 2D material previously produced mainly through hazardous etching and generated complete growth recipes tailored to the lab's equipment. After 25 rounds with human refinement, the team produced layered structures whose properties resemble the target material, but definitive confirmation of the atomic structure is still pending. In a second experiment, three semiconductor thin films were synthesized on the first try. Co-Scientist used Gemini 3 Deep Think for direct equipment control, cutting recipe development from days down to minutes. Humans still had to load samples and precursor materials manually, and the fast mode produced smaller, less uniform crystals than carefully optimized recipes would. Whether the recipes transfer to other labs remains open, lead author Samuel Schmidgall writes . In biology, Co-Scientist autonomously built an image analysis pipeline that predicts which patterns genetically engineered E. coli colonies form at different chemical concentrations. Predictions generated with Gemini 3 Pro Image matched unpublished lab results for three out of four shape features. The researchers acknowledge the system only reasons between known conditions and can't predict behavior in entirely new systems. That would be far more remarkable . A computer science experiment ran without any human involvement beyond the initial setup. Co-Scientist designed "Agent_H," a medical AI architecture that classifies incoming queries, generates dozens of response candidates in parallel, and refines them. After correcting for overly long responses, Agent_H outperformed six frontier models on health benchmarks, including GPT-5 and Claude Opus 5. But the benchmark results didn't hold up against human evaluation. Three board-certified physicians scored responses across nine categories, and Agent_H showed a statistically significant advantage over the baseline Gemini 3.1 Pro in just one, a lower risk of potentially harmful responses. The automated benchmark evaluators also correlated only weakly with the physicians' judgments. High benchmark scores don't mean a system actually delivers better answers from a clinical perspective, the researchers say, which raises questions about what these benchmarks are really measuring. Reliability modules push hallucination rate down to 4 percent A core problem with LLM-based autonomous research systems is AI bullshit . When an AI agent is rewarded for good results, it has an incentive to make things up. Previous analyses documented fabrication rates of 80 to 100 percent in existing systems. Co-Scientist addresses this two ways. The system is penalized for fabricated or plagiarized content, and a separate verification module cross-checks every numerical claim against the actual results of the executed code. In a double-blind study with 30 domain experts and 450 independent reviews, the researchers evaluated 150 autonomously generated papers. With reliability modules active, Co-Scientist fabricated key results in 4 percent of cases. Without them, the rate hit 46 percent. The comparison system reached 90 percent. Completely fabricated data never appeared in Co-Scientist's output but showed up in 44 percent of the comparison system's papers. Near-plagiarized content dropped from 60 percent to 16 percent. An integrated safety architecture rejected 98.7 percent of potentially harmful research directions. Despite the gains, the researchers still observed leftover errors. The system tends toward selective reporting and, according to Schmidgall, writes "highly plausible methods in the paper that did not match its actual code." The gap between lab assistant and autonomous researcher is still wide The researchers see the results as progress toward closed-loop multi-agent AI systems that improve through experimental feedback and could speed up scientific research. "There is a long journey ahead before AI systems can navigate the physical realities of science. But we are deeply excited about the potential of LLMs to help people and accelerate real-world progress," Schmidgall writes . Automated research is currently one of the most hyped-up AI applications. OpenAI plans to unveil an AI agent system this fall that can conduct research at least at intern level. But there's an ongoing debate about whether current LLM-based systems can truly discover new knowledge or are just making explicit what's already buried in their training data . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->