메뉴
HN
Hacker News 27일 전

관찰적 증거에 대한 옹호

IMP
6/10
핵심 요약

이 글은 무작위 대조 시험(RCT)이 의학 등에서 증거의 '황금 표준'으로 추앙받고 있지만, 막대한 비용과 시간이 든다는 점을 지적하며 관찰적 증거(Observational Evidence)의 가치를 재조명합니다. 역사적인 사례를 통해 방대한 샘플 크기를 확보할 수 있는 관찰 연구가 오히려 미세한 차이를 매우 효율적으로 밝혀낼 수 있음을 보여줍니다.

번역된 본문

모두가 무작위 대조 시험(RCT)이 증거의 황금 표준이라는 것을 알고 있습니다. ….맞죠? 1710년, 스코틀랜드의 의사 존 아버스넛(John Arbuthnot)은 신의 존재에 대한 새로운 증거를 제시했습니다. 그는 82년 연속으로 런던에서 여아보다 더 많은 남아의 세례가 기록되었다는 사실을 관찰했습니다. 여아가 태어날 확률과 남아가 태어날 확률이 같고, 연도에 따라 독립적으로 변한다고 가정할 때, 이 결과가 우연히 발생할 확률은 0.5^82입니다. 따라서 이 성비는 우연이 아니라 신성한 통합 원리에 의해 지배된다고 결론지었습니다.

피에르시몽 라플라스(Pierre-Simon Laplace)는 1781년에 발표된 분석에서 이 데이터를 재검토했고, 남아가 태어날 확률은 단순히 절반보다 약간 높다는 더 차분한 결론을 내렸습니다. 유럽 여러 도시의 출생 성비 차이에 더 관심을 가진 라플라스는 런던의 이 비율이 파리보다 0.38% 더 높다는 것을 발견했고, 이를 유의미하다고 판단했습니다. 파리와 나폴리 간의 동일한 비교에서는 1/100의 확률이 나왔지만, 라플라스는 이를 '최종적인 결론을 내리기에는 충분히 극단적이지 않다'고 보았습니다.

이러한 정량적 가설에 대한 초기 테스트는 관찰적 증거의 위험성과 장점을 모두 보여줍니다. 아버스넛이 보여주었듯이, 데이터가 우리가 이미 사실이라고 믿고 싶었던 것을 증명해주는 것을 발견하기란 충분히 쉬운 일입니다. 그럼에도 불구하고 아버스넛과 라플라스 모두 지역 기록을 확보하여 책상 앞에서 바로 과학적 연구를 시작할 수 있었다는 점은 놀라운 일입니다. 이 과정에서 두 사람은 데이터를 얻기 위해 많은 노동이나 자본을 투자할 필요 없이 중요한 현상을 정확하게 문서화했습니다.

지식을 얻는 이러한 방법의 효율성, 단순성 및 아름다움은 특히 의학 및 공중 보건 분야에서 과소평가되어 왔습니다. 신약과 같은 개입의 경우 무작위 대조 시험(RCT) 형태로 데이터를 얻는 것이 이상적인 것으로 간주되지만, 이러한 형태의 데이터 수집이 두 분야 모두에서 항상 가능한 것은 아닙니다. 물론 라플라스가 테스트한 '파리 대 런던' 가설은 비교적 단순했지만, 그의 표본 크기는 193만 명 이상으로 의학 역사상 대다수의 중재적 시험보다 많았습니다. 이진 확률 변수에서 0.3%의 차이를 감지할 수 있을 만큼 충분히 크고 잘 분포된 표본을 가진 무작위 시험은 매우 드물지만, 라플라스는 200년이 넘는 시간 전에 이를 해냈습니다.

RCT의 장점은 이를 의학 분야 중재 시험의 황금 표준으로 확고히 했으며, 오늘날에도 많은 일반인들이 과학을 하는 유일하고 진정한 방법으로 생각하고 있습니다. 그러나 일단 우리가 이러한 장점이 어디서 오는지, 이들이 표본 수집의 경제학과 어떻게 상호작용하는지, 그리고 대안책의 장점을 이해하게 되면, 관찰적 증거가 우리가 생각하는 것보다 더 자주 승자로 떠오른다는 것을 알게 될 것입니다.

RCT의 느린 발명

통제된 시험에 대한 최초의 서면 언급은 기원전 7세기로 거슬러 올라가는 다니엘서 1장 11-16절에서 찾을 수 있습니다. 예언자 다니엘은 바빌론의 느부갓네살 왕의 시종장에게 왕이 먹는 화려하고 코셔(kosher) 식단이 아닌 채식 식단을 먹게 해달라고 허락을 구합니다. 왕의 시종장은 다니엘과 그의 동료들이 쇠약해질까 봐 걱정했지만, 10일 동안 식단을 테스트해본 뒤 두 식단을 따르는 사람들의 모습이 얼마나 '건강하고 살이 찐(fair and fat)' 상태인지 평가하기로 동의했습니다. 채식주의자들이 승리했고, 시종장은 납득했습니다.

전근대시대의 다른 통제된 시험은 찾아보기 힘들지만, 11세기 페르시아의 철학자 이븐 시나(Ibn Sina)의 '의학의 법칙(Canon of Medicine)'은 실험 설계를 위한 규칙을 정립했으며, 여기에는 환자를 '복합적인 상태가 아닌 단일 조건'으로 테스트해야 한다는 규정이 포함되어 있었습니다. 이는 특이한 동반 질환이나 다른 특이한 상황을 가진 환자를 시험에서 제외시키는 현대 RCT의 원칙과도 맞닿아 있습니다. 오늘날처럼 거의 천 년 전에도 민감했던 주제 중 하나는 바로 표본 크기였습니다. 의학에서 대규모 표본 비교를 제안한 가장 초기의 기록 중 하나는 실제 시험에 대한 묘사가 아니라 실현되지 못한 제안에서 찾아볼 수 있습니다.

원문 보기
원문 보기 (영어)
Everyone knows RCTs are the gold standard of evidence. ….Right? In 1710, Scottish doctor John Arbuthnot presented a new proof for the existence of God. 1 He had observed that for 82 years in a row, London counted more christenings of baby boys than girls. Assuming that the probability of birthing a girl is equal to that of birthing a boy, and that it varies independently over years, the odds of this outcome occurring by chance are 0.5^82. It follows that the ratio must be governed not by random chance but by a divine unifying principle. Pierre-Simon Laplace revisited the data in an analysis published in 1781, and concluded more soberly that the probability of birthing a boy is simply a bit higher than one in two. Further interested in the difference in male-to-female birth proportions in various European cities, Laplace found that this proportion was 0.38% higher in London than in Paris, which he found significant. 2 The same comparison between Paris and Naples yielded a probability of 1/100, which Laplace didn’t consider “sufficiently extreme for an irrevocable pronouncement.” These early tests of quantitative hypotheses illustrate both the risks and merits of observational evidence. To be sure, as Arbuthnot showed, it’s easy enough to find that the data proves something you wanted to be true all along. Yet it is remarkable that both Arbuthnot and Laplace could obtain local records and start doing science right from their desks. In doing so, both correctly documented important phenomena without the need to invest much labor or capital to get the data. The efficiency, simplicity, and beauty of this method of gaining knowledge has been underappreciated, especially in medicine and public health. Although it is considered ideal to obtain data in the form of a randomized control trial for an intervention like a drug, that form of data collection is not always possible in either field. Granted, the Paris versus London hypothesis Laplace was testing is relatively simple, but his sample size was over 1.93 million, more than the vast majority of interventional trials in the history of medicine. It is a rare randomized trial (studying a particular intervention with a sufficiently large and well-distributed sample population) that can detect an 0.3% difference in a binary random variable — but Laplace could, more than 200 years ago. The advantages of the RCT have cemented it as the gold standard for interventional trials in medicine, and it remains what many laypeople think of as the one true way to do science. Yet once we understand where these advantages come from, how they interact with the economics of collecting samples, and the merits of the alternative, observational evidence emerges as the winner more often than one might think. The slow invention of the RCT The first written mention of a controlled trial can be found as early as the 7th century BC in the Book of Daniel 1:11-16. The eponymous prophet asks the steward of Nebuchadnezzar, the king of Babylon, for permission to eat a vegetarian diet instead of the king’s rich and possibly non-kosher food. The king’s steward worries that Daniel and his companions will waste away, but agrees to let them test the diet for 10 days, after which adherents of both diets will be evaluated by how “fair and fat” they looked. The vegetarians win, and the steward is convinced. While other pre-modern controlled trials are thin on the ground, the 11th century Persian philosopher Ibn Sina’s Canon of Medicine does set out rules for designing experiments, including the prescription to test patients with a “single, not a composite condition.” This is echoed in modern RCTs, in which patients who have unusual comorbidities or other unusual circumstances are excluded from trials. A subject as touchy nearly a millenium ago as it is today is sample size. One of the earliest mentions of a proposed large sample comparison in medicine is found not in a description of a real trial, but from an unrealised proposal to conduct one. In one of his letters from the 14th century Italian poet Petrarch declared the following as part of a polemic against physicians: I solemnly affirm and believe, if a hundred or a thousand men of the same age, same temperament and habits, together with the same surroundings, were attacked at the same time by the same disease, that if one half followed the prescriptions of the doctors of the variety of those practicing at the present day, and that the other half took no medicine but relied on Nature’s instincts, I have no doubt as to which half would escape. 3 Flemish doctor Jan Baptist van Helmont was somewhat more optimistic about the practice of medicine when he wrote a provocative letter in 1648, casually proposing mentioning what could be the first explicit randomization with an external source of entropy: Let us take from the itinerants’ hospitals, from the camps or from elsewhere 200 or 500 poor people with fevers, pleurisy etc. and divide them in two: let us cast lots so that one half of them fall to me and the other half to you. I shall cure them without blood-letting or perceptible purging, you will do so according to your knowledge (nor do I even hold you to your boast of abstaining from phlebotomy or purging) and we shall see how many funerals each of us will have: the outcome of the contest shall be the reward of 300 florins deposited by each of us. 4 The sample size of what may be one of the first earliest recorded attempts at a clinical controlled trial was a bit smaller: 12. In 1747, the British naval doctor James Lind wanted to test the most promising cures for scurvy. He selected 12 patients with similar symptoms, all fed the same diet and all housed together in the same part of the ship. The only difference was in their treatment: cider, vinegar, “elixir vitriol” (a medical extract of alcohol and sulfuric acid), sea water, citrus fruit, or a medicinal paste recommended by the ship’s surgeon. This wasn’t quite a modern controlled trial — for one thing, there was no control group — but at the time, it hardly mattered. The citrus group were unambiguously the first to recover. Better design trials remained rare until the 20th century.A positive exception is an 1898 diphtheria trial conducted by Danish doctor Johannes Fibiger. With 484 participants, this trial was larger than Lind’s. Unlike most other doctors at the time, Fibiger intentionally sought a large sample size and ensured that his participants actually had the disease they sought to treat. Most importantly, he produced comprehensive documentation of what happened in the trial. 5 Fibiger’s trial also featured a true control group and the first attempt at explicit randomization in a clinical trial. Fibiger decided whether a patient would get their experimental serum based on the day they were admitted to the hospital, instead of preferentially treating sicker patients or anything else which might bias the results. One major source of bias in clinical trials comes from the researchers themselves: Even careful and well-intentioned researchers like Fibiger have an incentive to find that their own treatments work and seek all kinds of avenues, intentional or not, to nudge results in their favor. Another source of bias is the placebo effect, where dummy treatments that a patient believes are real can sometimes have clinical effects. Today, this is addressed by double-blind trials, where neither the doctors conducting the study nor the patients themselves know who is getting a placebo or the real treatment. In 1946, the United Kingdom’s Medical Research Council started planning our last candidate for the first-ever RCT: a trial of streptomycin for treating tuberculosis. This experiment brought together many of the elements of the RCT as we find it today: attempted randomization and concealment in treatment assignments, systematic enrollment criteria, randomization, and quantitative hypothesis testing. Taking all the improvements in methodolog