메뉴
BL
The Decoder 5일 전

무료 챗GPT, 유료 버전보다 낮은 수준의 건강 조언 제공

IMP
8/10
핵심 요약

OpenAI가 미국 사용자를 대상으로 'Health in ChatGPT' 기능을 본격 출시하며, 애플 건강 앱 및 의료 기록과의 연동을 지원합니다. 이 기능은 유료 구독자에게는 더 우수한 최신 모델(GPT-5.6)을 제공하고, 무료 사용자에게는 성능이 낮은 모델(GPT-5.5)을 제공하는 이원화된 시스템을 채택했습니다. 전문가들은 AI가 여전히 의료적 한계를 인정하지 않은 채 확신에 찬 잘못된 정보를 제공할 위험이 크기 때문에, 의사의 진료를 완전히 대체할 수는 없다고 경고합니다.

번역된 본문

원문 제목: 챗GPT, 돈을 내지 않으면 더 낮은 수준의 건강 조언 제공 소스: 블로그 (THE DECODER)

핵심 요점:

  • OpenAI는 만 18세 이상 미국 사용자를 대상으로 'Health in ChatGPT(챗GPT 내 건강)' 기능을 출시했습니다. 이를 통해 애플 건강(Apple Health), 의료 기록, 웰니스 앱을 연동하여 검사 결과를 검토하고, 진료 예약을 준비하며, 건강 데이터를 분석할 수 있습니다.
  • 무료 사용자는 GPT-5.5 Instant 기반으로 품질이 낮은 건강 조언을 받는 반면, 유료 구독자는 더 강력한 GPT-5.6 Sol 모델에 접근할 수 있습니다.
  • 매주 3억 명 이상이 챗GPT에게 건강 관련 질문을 던지고 있지만, 리스크는 여전히 큽니다. AI 챗봇은 자신의 불확실성을 인정하기보다 높은 확신을 가지고 잘못된 의료 소견을 제공하는 경향이 확인되었습니다.

본문: OpenAI는 1월에 이 기능을 발표하고 테스트한 후, 만 18세 이상의 미국 사용자에게 'Health in ChatGPT'를 본격적으로 출시했습니다. 이제 구독료를 지불하는 것이 문자 그대로 당신의 생명을 구할 수도 있습니다. 사용자들은 애플 건강, 의료 기록 및 웰니스 앱을 연동하여 검사 결과를 검토하고, 의사 진료를 준비하며, 수면 및 활동 데이터를 분석할 수 있습니다. OpenAI는 연동된 건강 데이터를 모델 학습이나 광고에 사용하지 않을 것이라고 밝혔습니다.

무료 사용자, 더 낮은 수준의 건강 조언 받아 무료 버전의 챗GPT 사용자들은 더 낮은 수준의 건강 조언을 받게 됩니다. OpenAI는 이 기능을 GPT-5.5 Instant로 구동하는데, 이 모델은 유료 사용자를 위해 예약된 새로운 최고급 모델인 GPT-5.6 Sol보다 건강 벤치마크 점수가 낮습니다. OpenAI는 두 모델 모두 'HealthBench Professional' 테스트에서 의사들의 답변보다 뛰어난 성과를 보였다는 점을 들며 윤리적 근거를 통해 이러한 이원화된 시스템을 방어할 가능성이 높습니다.

하지만 벤치마크 결과가 결정적으로 보일지라도, 이는 지식을 측정하기 위해 설계된 인위적인 테스트 환경에서 나온 결과이며, 의사들은 여러 가지 이유로 낮은 점수를 받을 수 있습니다. 의사들은 시간 압박을 받거나 피로에 시달릴 수 있고, 환자 기록이나 동료의 의견과 같은 도구 없이 테스트를 진행할 수도 있습니다. 또한 벤치마크는 실제 의료 검진 중에 일어나는 많은 일들을 반영하지 못합니다. 실제 진료에서 의사는 환자를 직접 대면하여 검사하고, 비언어적인 단서를 파악하며, 수년간의 경험을 바탕으로 환자의 전반적인 상태를 평가합니다. OpenAI 자체도 공지에서 챗GPT는 여전히 실수를 할 수 있으며 의료 조언을 대체할 수 없다고 거듭 강조하고 있습니다. 이 회사는 260명 이상의 의사들이 이 건강 기능 개발을 도왔다고 밝혔습니다.

OpenAI의 초기 테스트에 따르면, 참가자의 70% 이상이 전용 건강 섹션으로 이동하는 것이 너무 번거로워 일반 대화에서 건강 질문을 했습니다. 이에 따라 회사는 데이터 관리 및 과거 건강 대화 기록에 액세스할 수 있는 별도의 건강 섹션은 유지하면서도, 모든 대화에서 건강 기능을 사용할 수 있도록 변경했습니다.

'구글 닥터'보다는 나을지 몰라도, 의사를 대체할 수는 없어 OpenAI에 따르면 현재 매주 3억 명 이상이 챗GPT에게 건강 질문을 던지고 있으며, 이는 1월의 2억 3천만 명에서 증가한 수치입니다. 회사는 이 건강 기능이 유럽에 제공될지 여부나 시기에 대해서는 언급하지 않았습니다. OpenAI가 1월에 이를 발표했을 때, 유럽 경제 지역, 스위스, 영국은 명시적으로 제외되었습니다. 더 엄격한 EU 데이터 개인정보 보호 규칙과 EU AI 법(AI Act)에 따라 이 기능이 고위험군으로 분류될 가능성이 그 이유일 수 있습니다.

이러한 위험은 가상의 것이 아닙니다. 최신 방사선 벤치마크인 RadLE 2.0에서 테스트된 16개의 AI 모델 중 인간 방사선 전문의만큼 우수한 성능을 보인 모델은 없었습니다. 가장 큰 문제는 챗봇이 한계에 도달했을 때 이를 인정하는 대신 높은 확신을 바탕으로 잘못된 소견을 제공했다는 점이며, 반면 인간 방사선 전문의들은 불확실성을 인정하는 데 훨씬 능했습니다. 과도한 자신감, 설득력, 그리고 사용자의 그릇된 가정을 지적하기보다 동조하는 '아첨(sycophancy)'이 혼합된 양상은 심각한 정신 건강 피해를 초래하기도 했습니다. 동시에, 일부 보고서에 따르면 AI가 의료 전문가가 놓친 건강 데이터의 패턴을 발견하여 때로는 놀라운 결과를 보여주기도 합니다.

원문 보기
원문 보기 (영어)
ChatGPT will give you worse health advice if you don't pay Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 23, 2026 GPT-Image-2 prompted by THE DECODER Key Points OpenAI is rolling out its "Health in ChatGPT" feature to U.S. users aged 18 and older, allowing them to connect Apple Health, medical records, and wellness apps to review lab results, prepare for appointments, and analyze health data. Free users receive lower-quality health advice powered by GPT-5.5 Instant, while paying subscribers get access to the stronger GPT-5.6 Sol model. Despite more than 300 million people asking ChatGPT health questions each week, risks remain significant: AI chatbots have been shown to give incorrect medical findings with high confidence rather than admitting uncertainty. Ask about this article… Search OpenAI is rolling out "Health in ChatGPT" to U.S. users aged 18 and older after unveiling and testing the feature since January. Paying for a subscription could literally save your life. Users can connect Apple Health, medical records, and wellness apps to review lab results, prepare for doctor's appointments, and analyze sleep or activity data. OpenAI says it won't use connected health data for model training or advertising. Free users get worse health advice Users on the free version of ChatGPT receive lower-quality health advice . OpenAI powers the feature with GPT-5.5 Instant , which scores lower on health benchmarks than the new flagship model, GPT-5.6 Sol , reserved for paying users. OpenAI will likely defend this two-tier system on ethical grounds by pointing out that both models beat doctors' answers on the HealthBench Professional test. Ad Even when benchmark results appear decisive, they come from artificial test environments designed to measure knowledge, and doctors may score lower for several reasons. They may be under time pressure, dealing with fatigue, or taking the test without tools such as patient records or input from colleagues. Ad DEC_D_Incontent-1 Benchmarks also can't capture much of what happens during an actual medical exam, where doctors can examine patients in person, pick up on nonverbal cues, and draw on years of experience to assess their overall condition. OpenAI itself repeatedly states in the announcement that ChatGPT can still make mistakes and can't replace medical advice. The company says more than 260 physicians helped develop the Health features. OpenAI's early tests also found that more than 70 percent of participants asked health questions outside the dedicated Health section because switching to it was too cumbersome. The company has since made Health available in any conversation while keeping the separate Health section for managing data and accessing past health chats. Ad A better Dr. Google, maybe, but no substitute for a doctor OpenAI says more than 300 million people now ask ChatGPT health questions each week, up from 230 million in January . The company hasn't said whether or when the Health feature will be available in Europe. When OpenAI announced it in January, it specifically excluded the European Economic Area, Switzerland, and the United Kingdom. Stricter EU data privacy rules and the possibility that the feature could be classified as high-risk under the EU AI Act are likely reasons. Those risks aren't hypothetical. In the latest radiology benchmark, RadLE 2.0, none of the 16 AI models tested performed as well as human radiologists . The main problem was that chatbots gave incorrect findings with high confidence instead of admitting when they had reached their limits, while human radiologists were far better at acknowledging uncertainty. That mix of overconfidence, persuasion, and sycophancy, where chatbots validate users rather than challenge false assumptions, has also contributed to serious mental health harms . Ad DEC_D_Incontent-2 At the same time, some reports show AI spotting patterns in health data that medical professionals miss, sometimes with striking results . MIRA, a system for electronic health records, and AMIE both performed about as well as primary care doctors in simulated consultations . One researcher compared AI agents like these to an airplane's autopilot: "These systems can support and relieve medical professionals by taking over routine tasks, but ultimate responsibility will always remain with the physicians." Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI