메뉴
BL
The Decoder • 39일 전

JAMA 논문 "AI가 의사를 능가하는데 규제가 의사 개입을 강제해선 안 된다"

IMP
8/10
핵심 요약

의학 저널 JAMA에 실린 논문이 자율적 AI가 의사-AI 협업보다 의학적 추론에서 뛰어날 것이라 전망하며, 규제기관이 '휴먼 인 더 루프'를 의무화하는 것에 반대했다. 저자들은 2024년 이후 연구에서 AI가 단독으로 진단 등 핵심 과제에서 의사와 맞먹거나 더 나은 성과를 보였다며, AI가 명확히 앞설 때 인간 검토자는 오히려 오류 원인이 된다고 주장한다. 다만 저자 중 일부는 AI 원격진료 회사 관계자로 이해충돌 우려가 있고, 대부분의 근거가 시뮬레이션에 기반한 한계가 있다.

번역된 본문

AI가 의사를 능가하는 상황에서 규제기관은 인간 개입을 강제해서는 안 된다는 JAMA 논문

의학 저널 JAMA에 실린 오피니언 글은 자율적 AI가 어떤 의사-AI 조합보다도 의학적 추론에서 더 뛰어난 성과를 낼 것이라 전망한다. 저자들은 규제기관이 규정에 '휴먼 인 더 루프(human in the loop)'를 명문화하는 것을 막고자 한다.

누가 이 글을 썼는지 알면 그 관점을 이해하는 데 도움이 된다. 제1저자 에제키엘 엠마누엘(Ezekiel Emanuel)은 펜실베이니아대 생명윤리학자이자 오바마 행정부 의료개혁의 설계자 중 한 명으로, 미국 보건정책계의 중진이다. 공동저자 닐 코슬라(Neil Khosla)는 AI 원격진료 회사 커라이 헬스(Curai Health)의 CEO이며, 그의 아버지 비노드 코슬라(Vinod Khosla)는 OpenAI와 커라이 헬스 양쪽에 투자한 투자자다. 저자 중 두 명은 이 글이 옹호하는 미래로부터 직접 이득을 볼 수 있는 입장이다.

이러한 미래는 미국의사협회(AMA)나 미국내과의사학회(ACP) 같은 의사 단체의 입장과 충돌한다. 이들은 AI가 의사를 보조할 수는 있어도 결코 대체해서는 안 된다고 주장해왔다. 의학교수 로버트 웩터(Robert Wachter)는 저서에서 더 나아가 AI 단독 진료를 의학의 '이코노미 클래스'라고 불렀다. JAMA 저자들은 이런 서열 매기기는 입증되지 않았다며 두 가지 반론을 제시한다.

첫 번째는 연구 결과에 관한 것이다. 2024년 이후 대부분의 연구에서 AI는 단독으로 의학의 5가지 핵심 추론 과제, 즉 환자 병력 청취, 진단, 검사 선택, 가이드라인에 따른 치료, 만성질환 관리에서 의사와 맞먹거나 더 나은 성과를 보였다. 구글의 대화형 시스템 AMIE는 모의 환자 대화에서 거의 모든 범주에서 1차 진료 의사보다 높은 점수를 받았다. 377건의 복잡한 사례에서 ChatGPT o3는 60%의 경우 정확한 진단을 첫 번째로 제시했지만, 20명의 내과 의사들은 15.9%에 그쳤다. 마이크로소프트의 진단 오케스트레이터는 예산 제약 하에서 의사보다 약 4배 더 자주, 더 낮은 비용으로 정확한 진단을 찾아냈다. 저자들은 반대 결과를 보인 대부분의 연구는 시대에 뒤떨어졌거나 최고 성능의 모델을 제외하는 등 방법론적으로 취약하다고 일축한다.

두 번째 논거는 향후 전망에 관한 것으로, 글의 핵심을 이룬다. 모델은 빠르게 발전하는 반면, 의사들은 AI 사용으로 인해 자신의 실력을 잃어간다. 대장내시경에 관한 랜싯(Lancet) 연구가 이를 시사한다. 격차는 앞으로 더 벌어질 수밖에 없다. 기계가 명확히 앞서게 되면, 검토하는 의사는 안전망이 아니라 오류의 원천으로 변한다. 저자들에 따르면 106개 실험을 분석한 메타분석이 이를 뒷받침한다. 인간이 더 뛰어날 때는 결합이 도움이 되지만, AI가 더 뛰어날 때는 인간이 잘못된 부분에서 시스템을 번복함으로써 결과를 악화시킨다. 실제 환자 사례를 활용한 한 연구에서 GPT-4 단독은 진단 추론에서 92%를 기록한 반면, 같은 모델에 접근할 수 있었던 의사들은 겨우 76%에 그쳤다.

저자들은 체스를 역사적 선례로 든다. 딥블루(Deep Blue)가 1997년 카스파로프를 이긴 후 수년간은 인간-기계 팀이 우세했지만, 2017년부터 AI가 그 팀마저 이기기 시작했다.

규정 입안자들에 대한 경고

이로부터 이 글은 진짜 메시지를 도출한다. 의사가 최종 결정을 내리도록 요구하는 가이드라인을 고착시키면 곧 뒤처질 형태의 진료를 굳히는 결과가 될 수 있다는 것이다. 저자들은 2030년까지 인지 과제에 한정되지만 자율적 AI가 일부, 어쩌면 많은 업무 흐름에서 준비될 것으로 예상한다. 책임, 보수, 규제, 의학 교육 모두 지금 다시 생각해야 한다고 주장한다.

다만 저자들 스스로 인정하듯 그 주장에는 한계가 있다. 거의 모든 근거가 실제 환자 진료가 아닌 단일 과제 시뮬레이션에서 나왔고, 인간과 모델 간 정보 전달은 약점으로 꼽힌다. 수술, 분만, 대장내시경 같은 신체적 시술은 로봇공학이 따라오지 못하는 한 당분간 인간의 몫으로 남는다. 또한 자율 시스템은 환각, 인터넷 장애, 사이버 공격 등 의사와는 다른 방식으로 실패한다. 이러한 위험들도 함께 저울질해야 한다.

원문 보기
원문 보기 (영어)
As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 18, 2026 Nano Banana Pro prompted by THE DECODER An opinion piece in the medical journal JAMA predicts that autonomous AI will outperform any doctor-AI pairing at medical reasoning. The authors want to stop regulators from writing a human in the loop into the rules. Knowing who wrote the JAMA piece helps explain its angle: Lead author Ezekiel Emanuel is a bioethicist at the University of Pennsylvania and one of the architects of Obama's healthcare reform, making him a heavyweight in US health policy. Coauthor Neal Khosla is CEO of the AI telemedicine company Curai Health, and his father, Vinod Khosla, is an investor in both OpenAI and Curai Health. Two of the authors stand to gain directly from the future the piece champions. That future clashes with the position held by physician groups like the American Medical Association and the American College of Physicians , which say AI should support doctors, never replace them. Medical professor Robert Wachter goes further in his book, calling AI-only care the "economy class" of medicine. The JAMA authors consider this ranking unproven and offer two arguments against it. The first concerns the research. In most studies since 2024, AI on its own matches or beats doctors at the five core reasoning tasks of medicine, namely taking a patient history, making diagnoses, choosing tests, treating according to guidelines, and managing chronic disease. Google's conversational system AMIE scored higher than primary care doctors in almost every category during simulated patient conversations. Across 377 complex cases, ChatGPT o3 named the correct diagnosis first 60 percent of the time, compared with 15.9 percent for 20 internists. Microsoft's diagnostic orchestrator found the correct diagnosis under budget constraints about four times as often as doctors, at lower cost. The authors dismiss most studies with the opposite finding as outdated or methodologically weak, for example because they left out the best models. The second argument is a forecast about where things are headed, and it forms the heart of the piece. The models are improving fast, while doctors lose their own skills through AI use, as a Lancet study on colonoscopies suggests. The gap should only widen from here. Once the machine is clearly ahead, the doctor doing the checking turns from a safety net into a source of error. According to the authors, a meta-analysis of 106 experiments backs this up. When the human is better, the combination helps. When the AI is better, the human makes the result worse by overruling the system in the wrong places. In a study using real patient cases , GPT-4 alone scored 92 percent on diagnostic reasoning, while doctors with access to the same model scored just 76 percent. The authors use chess as a historical parallel: After Deep Blue beat Kasparov in 1997, human-machine teams dominated for years, until AI began beating the teams too starting in 2017. A warning to the rule writers From this the piece draws its real message: Locking in guidelines that require a doctor to make the final call could cement a form of care that will soon fall behind. By 2030, the authors expect autonomous AI to be ready for some, maybe many, workflows, limited to cognitive tasks. Liability, payment, regulation, and medical training all need a rethink now, they argue. There are however limitations to their message, as they admit. Almost all the evidence comes from simulations of single tasks, not real patient care, and the handoff of information between human and model counts as a weak spot. Physical procedures like surgery, childbirth, and colonoscopies will also stay with humans for now, since the robotics aren't up to it. And autonomous systems fail in ways doctors don't, through hallucinations, internet outages, or cyberattacks. Those risks have to be weighed against the higher accuracy, the authors concede. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->