메뉴
BL
The Decoder 10일 전

방사선 AI, 오진인데도 확신해 환자 위협

IMP
8/10
핵심 요약

방사선 전용 AI 성능을 평가하는 'RadLE 2.0' 벤치마크에 따르면, 최신 AI 모델들은 오답을 내놓으면서도 매우 확신하는 경향을 보여 의료 현장에서 위험할 수 있습니다. 의학에서 정확도 자체는 빠르게 향상되고 있으나, AI가 자신의 한계를 인지하지 못하고 무리하게 진단을 시도하는 것이 문제로 지적되고 있습니다.

번역된 본문

X-ray를 판독하는 AI 챗봇은 틀렸을 때도 위험할 정도로 확신에 차 있는 경우가 많습니다.

방사선학 분야의 AI 시스템이 언제 인간 의사에게 진단을 넘겨야 하는지 알고 있는지 평가하는 두 번째 버전의 RadLE 벤치마크가 공개되었습니다. 많은 모델이 완전한 확신을 가지고 잘못된 소견을 내놓으며, 이것이 환자 치료에 있어 위험한 이유입니다. 'Radiology's Last Exam'의 약자인 RadLE 2.0은 인도 아쇼카 대학교의 CRASH Lab에서 개발되었습니다. 이는 연구진이 2025년 9월에 처음 공개했던 테스트의 개정판입니다. 이 새로운 버전은 모델이 진단을 정확히 내렸는지, 그 대답에 얼마나 확신하는지, 그리고 자신의 능력 한계를 인정할 수 있는지를 측정합니다. AI는 0부터 4까지의 척도로 자신의 답변에 대한 확신도를 평가해야 하며, 명시적으로 "모르겠다"고 대답하는 것이 허용됩니다.

이 테스트는 16개의 모델에 걸쳐 200건의 사례를 진행했으며, 패널로 구성된 방사선 전문의들의 성과와 비교했습니다. 인간 전문의들은 2,000점 만점 중 988.7점을 받았습니다. 가장 성능이 좋은 AI 모델은 758점을 받았습니다.

지나친 추측보다 솔직한 침묵이 낫습니다

채점 시스템은 솔직함에는 보상을 주고 과도한 자신감에는 벌점을 줍니다. 높은 확신을 가지고 정답을 맞히면 만점을 받습니다. 반면 높은 확신을 가지고 오답을 내면 상응하는 점수를 잃게 됩니다. "모르겠다"고 대답하면 0점을 받지만 감점은 당하지 않습니다. 자신 있게 추측만 하는 모델은 기본 정답률이 괜찮아 보이더라도 순위에서 떨어지게 됩니다. 이 연구는 최근 크게 인용된 논문에서 제기된 문제를 다룹니다. 즉, 벤치마크가 정확도에만 보상을 주는 한, AI 모델은 무작정 추측하도록 학습된다는 것입니다. 의학 분야에서는 확신에 찬 오진이 불확실함을 솔직하게 인정하는 것보다 훨씬 더 위험합니다.

모든 영역을 석권하는 단일 모델은 없습니다

전반적인 1위 모델은 없었습니다. Anthropic의 Claude Fable 5는 신뢰할 수 있고 안전한 답변 부문에서 가장 우수하여 주요 지표에서 선두를 차지했습니다. Google의 Gemini 3 Pro는 순수 정확도가 가장 높았습니다. Meta의 Muse Spark 1.1은 언제 사례를 인간에게 넘겨야 하는지 아는 능력에서 가장 뛰어났습니다. Meta는 최근 모델이 오답을 내놓기보다 대답을 거부하는 빈도가 늘어남에 따라 Muse Spark 1.1의 환각(Hallucination) 비율을 거의 절반으로 줄였습니다. 반면 다른 최신 모델들은 반대의 양상을 보입니다. 예를 들어, Grok 4.5는 이전 버전보다 환각이 훨씬 심한데, 더 많은 것을 알게 되었지만 동시에 자신의 잘못된 대답에 더 큰 확신을 가지기 때문입니다.

연구진에 따르면, 여러 모델이 추측하는 대신 더 자주 침묵했다면 점수가 훨씬 더 높았을 것입니다. 이는 특히 오픈 웨이트(Open-weight) 모델과 의료 전용으로 학습된 모델들에서 두드러졌습니다. 이들은 거의 모든 사례에 대답하려고 시도했으며, 주로 매우 높은 확신을 가지고 잘못된 답을 제시했습니다. 테스트의 첫 번째 버전은 훨씬 더 심각한 그림을 보여주었습니다. 방사선 전문의는 83%의 정확도에 도달한 반면, 최고 모델은 약 30% 수준에 그쳤습니다. 단 3개월 만에 Gemini 3 Pro는 이미 전공의 수준을 넘어섰습니다. 정확도는 빠르게 향상되고 있지만, 여전히 모델들은 자체적인 한계에 대한 감각이 부족합니다.

환자들은 이미 자신의 MRI를 챗봇에 보내고 있습니다

점점 더 많은 사람들이 X-ray나 MRI 스캔을 챗봇에 업로드하고 그 대답을 믿고 있습니다. 최근 'npj Digital Medicine'에 발표된 연구에 따르면, 널리 사용되는 챗봇들이 의학적 질문에 대해 종종 신뢰할 수 없는 답변을 제공하는 것으로 나타났습니다. 연구진은 경영진과 투자자들이 AI 모델이 할 수 있는 일에 대해 공개적으로 과장하고 있다고 비난합니다. AI 시스템이 이미 99%의 의사보다 더 잘 진단한다는 주장은 대부분 일화나 시뮬레이션에 기반하고 있습니다. 최근 4월에 당시 최고 수준으로 평가받던 21개 모델을 대상으로 한 연구에 따르면, 이들은 감독 없는 임상 적용을 할 준비가 되지 않았음이 확인되었습니다.

RadLE 2.0은 새로운 모델을 포함하여 지속적으로 확장될 예정입니다. 비용 분석과 오류 분류가 포함된 전체 과학 논문의 출판이 발표되었습니다. 자율 주도형 의료 AI 에이전트에 관한 다른 두 최근 연구들은 다른 방향을 가리키고 있었습니다. 전자 건강 기록 시스템인 MIRA와 AMIE는 [원문 생략된 부분] 유지할 수 있었습니다.

원문 보기
원문 보기 (영어)
AI chatbots reading X-rays can be dangerously confident even when they're wrong Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jul 19, 2026 Nano Banana Pro prompted by THE DECODER The second version of the RadLE benchmark tests whether AI systems in radiology can tell when they should leave a diagnosis to a human. Many models produce wrong findings with full confidence, and that's what makes them dangerous for patient care. RadLE 2.0, short for "Radiology's Last Exam," was developed by the CRASH Lab at Ashoka University in India. It's the revised follow-up to a test the team first released in September 2025 . The new version measures whether a model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it's out of its depth. The AI has to rate its answers on a confidence scale from 0 to 4 and is explicitly allowed to say "I don't know." The test ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758. Honest silence beats overconfident guesswork The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer "I don't know," and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent. The study tackles a point recently raised by this highly cited paper : as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty. No single model wins across the board There's no overall winner. Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy. Meta's Muse Spark 1.1 was the best at knowing when to hand a case off to a human. Meta had recently cut Muse Spark 1.1's hallucination rate nearly in half because the model more often refuses to answer rather than giving a wrong one. Other frontier models trend the opposite way. Grok 4.5 , for example, hallucinates significantly more than its predecessor because while it knows more, it's also more convinced of its wrong answers. According to the research team, several models would have scored much better if they had stayed quiet more often instead of guessing . This was especially obvious among open-weight models and those trained specifically for medical use. They tried to answer nearly every case and were often wrong, usually with high confidence. The first version of the test painted an even starker picture . Radiologists hit 83 percent accuracy, while the best model managed only about 30 percent. Within three months, Gemini 3 Pro had already surpassed the level of resident radiologists. Accuracy is growing fast, but the models still lack any sense of their own limits. Patients are already sending their MRIs to chatbots More and more people are uploading X-rays or MRI scans to chatbots and trusting the responses. A recent study in npj Digital Medicine showed that widely used chatbots frequently give unreliable answers to medical questions. The research team accuses executives and investors of publicly overstating what AI models can do. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations. As recently as April, a study of 21 models that were then considered state-of-the-art showed they aren't ready for unsupervised clinical use. RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced. Two other recent studies on autonomous medical AI agents pointed in a different direction. MIRA, a system for electronic health records, and AMIE were able to keep pace with general practitioners in simulated consultations. Both fueled expectations that AI could soon make diagnoses on its own. The RadLE 2.0 authors push back: before an AI makes decisions independently, it has to know when it's better off not doing so. Then there's the problem of skill loss. A Polish observational study from 2025 found that doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions without the tool. Detection rates dropped from 28.4 to 22.4 percent. The authors call it the "Google Maps effect": without the navigation aid, users are lost. Radiology has seen this AI hype before Radiology already went through one AI hype cycle. In 2016, AI researcher Geoffrey Hinton declared that we should stop training radiologists because deep learning would soon take over the job. Colleagues like Richard Sutton agreed. Nearly ten years later, radiologists are still overburdened, and Hinton had to walk back his prediction. He had reduced the profession to image analysis and overlooked the complexity of the entire field. The fact that these systems can confidently produce wrong diagnoses means humans remain indispensable. OpenAI CEO Sam Altman spent years predicting that AI would replace human jobs at a scary pace , then recently walked it back, suggesting AI may have actually created more jobs . So far, research doesn't support either claim. AI specialists may understand their models, but they routinely overestimate how fast entire professions can be replaced. Those kinds of predictions are back in fashion right now . Much like the AI they build, even people don't always know when they'd be better off staying quiet because they're outside their own expertise . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->