메뉴
HN
Hacker News 22일 전

AI 슈퍼포캐스터의 시대가 도래하다

IMP
8/10
핵심 요약

최근 예측 시장과 금융 시장에서 인간 전문가를 뛰어넘는 수익률을 기록하는 'AI 슈퍼포캐스터(superforecaster)'들이 등장했습니다. 이들은 프롬프트 엔지니어링과 다수의 하위 에이전트를 구동하는 스캐폴드(scaffold) 기술을 통해 복잡한 문제를 분석하고 미래를 고도로 정확하게 예측합니다. 예측 시장의 판도를 완전히 바꾸고 자본 시장의 효율성에 큰 변화를 가져올 핵심 AI 응용 분야로 주목받고 있습니다.

번역된 본문

예측 시장 연례 컨퍼런스가 지난달 초 열렸습니다. 올해는 예측 시장이 소규모 취미 생활에서 수십억 달러 규모의 산업으로, 반쯤 불법이던 상태에서 대통령의 아들이 자문으로 참여하는 수준까지 성장한 해였습니다. 하지만 컨퍼런스에서 누가 그런 이야기를 했는지 기억나지 않을 정도로 사람들의 관심은 온통 'AI 슈퍼포캐스터(AI superforecaster)'에게 쏠려 있었습니다.

저는 한 AI 슈퍼포캐스터 스타트업 창업자를 만났는데, 그는 자신의 AI가 칼시(Kalshi) 플랫폼에서 7개월 만에 35달러를 200만 달러로 불렸다고 말했습니다. 또 다른 창업자는 마켓 중립 포트폴리오를 이용해 주식 시장 수익률을 25%나 압도하고 있다고 했습니다. 물론 이것이 운일 수도 있지만, 그들은 칼시(Kalshi)와 폴리마켓(Polymarket)에서도 비슷한 차이로 승리하고 있었습니다. 사실 저는 이들의 말을 믿습니다. 그래프의 추세선을 연장해 온 AI 커뮤니티에서는 2026년~2027년경 AI가 최고의 인간 예측자를 뛰어넘을 것이라고 오래전부터 예측해 왔습니다. '미래를 예측하는 분야에서 마침내 봇이 인간을 이기는 순간'이 어떤 모습일 것이라 예상하셨나요? 분위기? 논문? 에세이? 돌이켜 보면 그것은 AI가 예측 시장에서 엄청난 수익을 올리고, 주식 시장을 안정적으로 이기는 모습으로 나타날 것이 분명합니다. 하지만 앞으로 어떤 일이 벌어질까요?

AI 슈퍼포캐스터 사용하기

자세한 내용으로 들어가기 전에, 이것이 정확히 무엇인지 알아봅시다. AI 슈퍼포캐스터는 미래 예측에 뛰어나도록 수정된 AI를 의미합니다. 일반적으로 챗GPT(ChatGPT)나 클로드(Claude) 같은 최첨단 모델이 사용됩니다. 이는 보통 '스캐폴드(scaffold, 뼈대)'를 의미하는데, 이는 다양한 프롬프트와 도구, 서브에이전트(subagent)를 언제 생성할지에 대한 조언 등을 통해 AI가 긴 연구 과정을 무사히 마치도록 돕는 프로그램입니다. 전반적인 경험은 다른 AI를 사용하는 것과 비슷하지만, 훨씬 더 많은 작업을 수행하기 때문에 속도는 느리고 비용은 더 많이 듭니다.

예시를 들면 더 이해하기 쉬울 것입니다. 주식 시장 수익률을 압도하고 있다고 주장하는 '퓨처서치(FutureSearch)'라는 회사에서 친절하게도 저에게 그들의 AI 슈퍼포캐스터를 사용해 보고 이곳에 글을 써볼 기회를 제공했습니다. 테스트 질문으로, 실리콘밸리의 일부 자선가들이 최근 일반적인 감기와 같은 호흡기 감염을 끝내기 위한 프로젝트를 시작했다는 사실을 떠올렸습니다. 저는 그들의 성공 가능성에 대해 물어보기로 했습니다. 예측자들에게는 매우 정밀한 질문이 필요하기 때문에, 2040년까지 감기 발병률이 절반으로 줄어들 가능성이 얼마나 되는지 물어보았습니다.

2분이 지나자, 이 AI는 3개의 서브에이전트를 배포하고 16개의 웹사이트를 읽었으며, (제가 스크린샷을 찍은 바로 그 순간) '2040년까지 일반 주거 환경에 ASHRAE Standard 241 공기 정화 기술이 광범위하게 채택될 가능성'을 조사하고 있었습니다. 5분 후, AI는 답을 도출했습니다. 2040년까지 미국의 호흡기 감염이 절반으로 줄어들 확률은 7%였습니다.

AI는 다음과 같이 설명했습니다: 에이전트의 추론은 동시에 모두 성공해야 하는 까다로운 복합적 요구 조건들을 중심으로 전개됩니다: 1: 생물학적 장벽이 너무 높습니다. 감기를 일으키는 200개 이상의 바이러스와 150개 이상의 라이노바이러스 혈청형이 존재합니다. 지난 50년간의 백신 개발 노력은 실패했으며, 2025년 말의 전문가들조차 5년 이내에 감기 백신이 나올 것으로 기대하지 않았습니다. 전례 없는 행동 변화가 있었던 코로나19 봉쇄 기간 동안에도 라이노바이러스 발생률은 아주 약간 감소하는 데 그쳤고 즉시 반등했습니다. 2: 프로젝트의 타임라인 자체가 빠듯합니다. 그들의 로드맵은 상용화까지 57년 이상이 소요될 것으로 예상하며, 이는 2040년까지 3상 임상시험을 마치고 FDA 승인을 받아 대중적으로 배포하기에는 너무 적은 시간입니다. 게다가 이 모든 것에 5억 달러 이하의 비용(해당 규모를 고려할 때 턱없이 부족한 금액)이 들 것으로 예상했습니다. 3: 채택 여부는 알 수 없습니다. 가벼운 질병의 경우 정기적인 비강 스프레이나 예방약에 대한 순응도가 낮을 것으로 예상됩니다. 공기 정화 인프라는 막대한 설치 비용과 불확실한 근거(2025년 살균 UV에 대한 무작위 대조 시험에서 감염 감소에 유의미한 효과를 발견하지 못함)에 직면해 있습니다. 4: 측정이 불가능할 수 있습니다. 미국에는 인구 전체를 대상으로 하는 일상적인 감기 감시 시스템이 없으며, 20202026년의 기준선은 팬데믹 시기의 행동 양식으로 인해 왜곡되었습니다. 즉, 실제로 진전이 있더라도 50% 이상의 감소를 확인된 데이터로 증명하기는 어렵습니다.

이 이니셔티브는 진정으로 혁신적이고 신뢰할 수 있지만, 혁신적인 AI의 가속화 없이 14년 안에 감기 빈도를 절반으로 줄이는 것은 비정상적으로 높은 장벽입니다. 이에 대해 후속 예측을 진행해 볼 수도 있을 것입니다.

원문 보기
원문 보기 (영어)
The AI Superforecasters Are Here ... Scott Alexander Jul 02, 2026 455 292 59 Share The annual prediction market conference was earlier this month. This was the year prediction markets went from an obscure hobby to a multi-billion dollar industry; from semi-illegal to having the President’s son as an advisor. I can’t remember if anyone talked about any of that. It didn’t even register. All eyes were on the AI superforecasters. I met an AI superforecaster startup founder who told me his AI had turned $35 into $2 million on Kalshi over seven months. I met another who said they were beating the stock market by 25% with a market-neutral portfolio - of course this could be luck, but they’d beaten Kalshi and Polymarket by similar margins. In fact, I believe all of these people. The extending-lines-on-graphs community has long predicted that AIs would beat the best human forecasters sometime in 2026 - 2027. What did you expect the bots-finally-beat-humans-at-predicting-the-future moment to look like? Vibes? Papers? Essays? In retrospect, sure: it will look like AIs making crazy profits on prediction markets and beating the stock market by some comfortable amount. But what happens next? Using An AI Superforecaster Before getting into details, what exactly are we talking about? An AI superforecaster is an AI - usually a frontier model like ChatGPT or Claude - which has been modified to be good at forecasting. This usually means a “scaffold” - a program that handholds it through a long research process with various prompts, tools, advice about when to create subagents, etc. The overall experience is a lot like using any other AI, but slower and more expensive, because it’s doing more work. This might make more sense with an example. FutureSearch - the company that claims to be beating the stock market - kindly offered to let me try their AI superforecaster and write about it here. For a test question - some Silicon Valley philanthropists recently started a project to end respiratory infections like the common cold . I decided to ask about their chances of success. Since forecasters need very precise questions, I asked how likely it was that the rate of colds would be cut in half by 2040: By two minutes in, the AI had deployed three subagents, read 16 websites, and (at the exact moment I took this screenshot) was “investigating the scalability of ASHRAE Standard 241 air cleaning technology for widespread residential adoption by 2040.” After five minutes, it had its answer: the chance of US respiratory infections halving by 2040 was 7% : It explained: The agents’ reasoning centers on a tough conjunctive chain of requirements that all must succeed simultaneously: 1: The biology is brutal. 200+ cold-causing viruses, 150+ rhinovirus serotypes alone — 50 years of vaccine efforts have failed, and experts as recently as late 2025 didn’t expect a cold vaccine within 5 years. During COVID lockdowns with unprecedented behavioral changes, rhinovirus incidence dropped only modestly and rebounded immediately. 2: Intercept’s own timeline is tight. Their roadmap projects commercialization 5–7+ years out, leaving little room to clear Phase 3 trials, FDA approval, and achieve mass deployment by 2040 — all for under $500M (likely far too little for that scope). 3: Adoption is a wildcard. For a mild illness, compliance with regular nasal sprays or prophylactics is expected to be low. Air-cleaning infrastructure faces steep retrofit costs and a mixed evidence base ( a 2025 RCT of germicidal UV found no significant reduction in infections). 4: Measurement may be impossible. The US lacks routine population-wide cold surveillance, and the 2020–2026 baseline is distorted by pandemic-era behavior — making a verified >50% reduction hard to demonstrate even if real progress occurs. The initiative is genuinely novel and credible, but halving cold frequency in 14 years without transformative AI acceleration is an extraordinarily high bar. You could follow up by forecasting a softer threshold (e.g., 25% reduction by 2035), or by ranking the specific Intercept technical approaches (nasal sprays vs. vaccines vs. air cleaning) by their likelihood of meaningful impact. …and, in case I wanted to learn more, cited its 212 sources. The forecast had taken five minutes and cost me $8 in credits. But is it true? Ideally we would wait until 2040 and see. For now, I started by comparing its answer to another superforecaster AI. Preseen is the company that claimed to 100,000x their seed money on Kalshi. Here’s their answer : 8.8% compared to FutureSearch’s 7%, not bad! Are either of these true? I asked a human superforecaster to predict this question, to see if she got the same as the AI. She said that depending on an ambiguity in the wording, she would give it 5-10%. Again, not bad! Man Vs. Machine Of course, it would be even better to do the same experiment at scale and figure out how AIs compare to humans once and for all. But measuring forecasting ability is hard. You can’t say something like “it gets 85% of questions right”, because that depends entirely on question difficulty. If the questions are things like “will the sun rise tomorrow morning”, then even a 100% hit rate is unimpressive. Instead, we can only match different forecasters against each other and determine who is better or worse. Any anchoring in an absolute space will come from the inclusion of groups whose predictive abilities we intuitively understand (eg the average member of the public, CIA analysts, etc). The forecasting website Metaculus matches AIs against humans and each other on a common metric. Here are their results over time : The Metaculus Community Prediction is a “wisdom of crowds” style aggregation of all the forecasters on Metaculus. The Metaculus Pro Forecasters are top professional superforecasters. This graph makes it look like - as of May 2026 when Gemini 3.1 was state of the art - AI was approaching the Community Prediction. This is no mean feat, but it’s still far from the professional superforecaster level. But in a recent blog post , Metaculus adds context. The graph above only measures out-of-the-box brand-name AIs like GPT and Claude. It doesn’t count forecasting-focused scaffolds like FutureSearch. A different investigation by Metaculus finds that these efforts are “worth 9 months of base model progress”, eg a well-scaffolded AI today is already as good at forecasting as base models will be in nine months. If you extend the dotted green line on the graph to July 2026, then add nine months for the extra scaffolding, it looks like the best AIs should be around 31, compared to top pro forecasters’ 36. So in theory, the absolute best forecasters in the world are still beating the top AIs, but the margin of victory is less than the graph suggests, and we should expect human-AI parity in about six months. But the claim that scaffolded AIs are nine months ahead of base models is itself ~9 months old. Several people in the field told me that they thought this underestimated true progress. Claims by the AI startups themselves may be treated skeptically, but even a few top human superforecasters said they were no longer confident they could beat the bots. Seems like time for a head-to-head matchup. The Metaculus Cup - the World Cup of forecasting! - is on the case. Once a season, top humans and AIs compete on about fifty questions like “Who will win the upcoming Nepali elections?” and “Will the US attack Iran?” Here are the winners of the most recent tournament: Humans took the top two spots, but Preseen’s AI came in third. Every forecasting competition involves a heavy dose of luck, so realistically at this point humans and AIs are in a statistical dead heat. We can confirm by looking at the intermediate results of the ongoing summer Metaculus Cup: Of humans who placed in the top ten during spring, 2/10 - benshindel and MarcosO - repeated their performance in summer. So did two top-ten AIs - manticAI and Laertes (