메뉴
BL
The Decoder • 1일 전

최고 AI 전문가들도 AI 발전 속도를 크게 과소평가했다

IMP
7/10
핵심 요약

예측연구소(FRI)의 중간 보고서에 따르면, 상위 대학 교수와 최다 피인용 연구자를 포함한 AI 전문가들이 최근 AI의 벤치마크 성과와 상용화 지표를 상당히 과소평가해 왔습니다. AI는 국제수학올림피아드 금메달 등 주요 마일스톤을 전망보다 수년 앞서 달성했으며, 경제 전망 역시 지나치게 보수적이었습니다. 이에 FRI는 빠른 AI 발전 속도에 맞춰 지속 갱신되는 LLM 예측 등 더 신속한 예측 방법을 도입할 계획입니다.

번역된 본문

연구 결과, 최고 AI 전문가들조차 AI 분야의 발전 속도를 크게 과소평가한 것으로 나타났다.

AI는 얼마나 빠르게 발전하고 있을까? 이 질문은 보통 최고 대학의 전문가, 피인용이 많은 AI 연구자, 경험이 풍부한 경제학자들에게 던져진다. 그러나 예측연구소(Forecasting Research Institute, FRI)의 중간 보고서에 따르면, 바로 이 전문가들이 최근 벤치마크와 일부 도입 지표에서 나타난 진전을 상당히 과소평가했다.

FRI는 2022년 중반 이후 여러 연구와 프로젝트를 통해 AI 진전에 관한 예측을 수집해 왔다. 그 표본에는 수석 전문가들이 포함되어 있다. LEAP(장기 전문가 AI 패널) 1차 조사에는 339명의 전문가가 참여했으며, 여기에는 76명의 컴퓨터 과학자, 76명의 산업계 전문가, 68명의 경제학자, 119명의 AI 정책 전문가가 포함되었다. 컴퓨터 과학자 중에는 상위 20개 대학의 교수 30명과 피인용 상위 200인 AI 저자 10명이 있었다. 패널에는 입증된 정확한 예측 기록을 가진 일반주의자인 '슈퍼포캐스터'들도 포함되었다.

AI, 전망보다 수년 앞서 주요 마일스톤 달성

가장 큰 격차는 수학 분야에서 나타났다. AI는 2025년 7월 국제수학올림피아드에서 금메달 수준에 도달했는데, 이는 전문가 중앙 예측보다 5년, 슈퍼포캐스터 중앙 예측보다 10년 앞선 것이다. 이 예측들은 ChatGPT 출시 이전인 2022년에 수집되었지만, FRI에 따르면 이후에도 같은 패턴이 이어졌다.

AI가 밀레니엄 문제(밀레니엄 상 문제) 하나를 풀었을 가능성도 있지만, 그 해법이 평가 기준을 충족하는지는 아직 불분명하다. 2025년 8~9월 설문에서 전문가들은 2027년 말까지 이런 해법이 나올 확률의 중앙값을 겨우 10%로 봤고, 슈퍼포캐스터는 5.4%로 봤다.

바이러스학 분야의 AI 역량 연구에서 전문가들은 AI 모델이 문제해결 벤치마크에서 최고 수준의 바이러스학자 팀과 맞설 수 있게 되는 시점을 2030년으로 예측했고, 슈퍼포캐스터는 2034년으로 봤다. FRI는 이것이 이미 2025년 4월에 일어났을 가능성이 높다고 밝혔다. 사이버보안 벤치마크에서도 비슷한 과소평가가 나타났다.

경제 전망도 지나치게 보수적이었다. 전문가들은 2026년 말 AI 기업의 최고 연간반복매출(ARR) 중앙값을 200억 달러로 봤다. 경제학자들은 160억 달러, 슈퍼포캐스터는 250억 달러로 봤다. FRI는 2026년 9월 기준 Anthropic의 매출이 약 1,000억 달러로, 이미 이 수치를 초과했을 가능성이 크다고 인용했다.

현실 세계 영향은 판단이 더 어려워

모든 예측이 낮게 잡힌 것은 아니었다. 바이오보안 전문가들은 언어모델을 사용하는 참가자의 22.5%가 생물학 실험 과제를 완수할 것으로 예측했다. 바이러스학자들은 40%, 슈퍼포캐스터는 16.2%로 봤다. 통제된 실험에서는 언어모델과 인터넷 접근을 갖춘 그룹의 5.2%만이 성공했으며, 인터넷만 사용한 그룹은 6.6%였다. 언어모델은 측정 가능한 차이를 만들지 못했지만, 실험 규모가 작았다.

전문가들이 자율주행차에서는 오히려 과대평가했을 수도 있다. 2027년 미국 이동서비스(라이드헤일링) 중 자율주행 비중에 대한 전문가 중앙 예측은 7.3%였으나, LLM 기반 전망은 2.5%로 보고 있다.

FRI는 경제 성장, 고용, 주요 AI 피해에 관한 예측은 아직 신뢰할 수 있게 판단할 수 없다고 밝혔다. 동시에 응답자들은 기대치를 상향 조정하고 있다. 두 차례 설문에 모두 응한 응답자들 중, AI가 '세기의 기술'이 될 확률의 평균값은 9개월 동안 전문가는 31%에서 36%로, 슈퍼포캐스터는 28%에서 35%로 상승했다.

FRI, AI 속도에 맞추기 위해 더 빠른 방법 도입

앞으로 FRI는 2040년까지 매우 빠른 AI 진전을 예상하는 응답자 하위 표본을 부각하고, 인간 예측과 함께 지속적으로 갱신되는 LLM 예측을 함께 발표할 계획이다. ForecastBench에 따르면 일부 모델은 특정 유형의 질문에서 이미 슈퍼포캐스터 수준에 도달했다. FRI는 또한 충분한 데이터가 쌓이면 가장 정확한 LEAP 패널 참가자를 선별해 그들의 예측을 주목할 계획이다.

FRI는 자체 데이터의 함정도 지적한다. 현실이 예측을 앞지르는 순간 과소평가는 분명해진다.

원문 보기
원문 보기 (영어)
Top AI experts badly underestimated how fast the field is moving, study finds Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 24, 2026 Nano Banana Pro prompted by THE DECODER How fast is AI improving? That question usually goes to experts at top universities, heavily cited AI researchers, and seasoned economists. Yet these same specialists significantly underestimated recent progress on benchmarks and some adoption metrics, according to an interim report from the Forecasting Research Institute (FRI). Since mid-2022, FRI has collected forecasts on AI progress across several studies and projects . Its samples include senior specialists. The first round of LEAP (Longitudinal Expert AI Panel) drew 339 experts, including 76 computer scientists, 76 industry experts, 68 economists, and 119 AI policy specialists. The computer scientists included 30 professors at top-20 institutions and 10 of the 200 most-cited AI authors. The panels also included superforecasters, generalists with a proven record of accurate predictions. AI hit major milestones years ahead of forecasts The widest gap involves math . AI reached gold-medal level at the International Mathematical Olympiad in July 2025 , five years before the median expert forecast and ten years before the median superforecaster forecast. Those predictions were gathered in 2022, before ChatGPT launched, but the pattern held afterward too, according to FRI. AI may also have solved a Millennium Prize Problem , though it's still unclear whether the solution meets the evaluation criteria. In a survey from August and September 2025, experts had put the median odds of such a solution by the end of 2027 at just 10 percent, and superforecasters at 5.4 percent. In a study of AI capabilities in virology, experts predicted AI models wouldn't match a top team of virologists on a troubleshooting benchmark until 2030. Superforecasters said 2034. FRI says that likely happened as early as April 2025 . A cybersecurity benchmark showed similar underestimates. Economic forecasts were also far too conservative. Experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion. Economists said $16 billion, and superforecasters said $25 billion. FRI cites roughly $100 billion for Anthropic in September 2026 as a figure that has likely already been reached. Real-world impact is harder to call Not every forecast ran too low. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks. Virologists expected 40 percent, superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone. The language model made no measurable difference, though the trial was small. Experts may also have overshot on self-driving cars. Their median forecast for the share of autonomous US ride-hailing trips in 2027 was 7.3 percent, while an LLM projection puts it at 2.5 percent. FRI says forecasts on economic growth, employment, and major AI harms can't be reliably judged yet. At the same time, respondents are revising their expectations upward. Among those who completed both surveys, the average probability assigned to AI becoming a "technology of the century" rose from 31 to 36 percent for experts and from 28 to 35 percent for superforecasters over nine months. FRI is adding faster methods to keep pace with AI Going forward, FRI will highlight a subsample of respondents who expect very rapid AI progress through 2040 and publish continuously updated LLM forecasts alongside the human ones. According to ForecastBench , some models already match superforecasters on certain question types. FRI also wants to find the most accurate LEAP panelists and feature their forecasts once enough data is in. RI does flag a catch in its own data: underestimates become obvious as soon as reality overtakes a prediction, but overestimates only become clear once a deadline passes. That makes the interim report naturally tilted toward finding cases where forecasters were too cautious. Some of FRI's own assessments also rely on LLM projections that use information the original forecasters didn't have. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->