메뉴
BL
The Decoder • 1일 전

AI 성능 비용, 역대 어떤 기술보다 빠르게 하락

IMP
7/10
핵심 요약

연구기관 Epoch AI는 특정 벤치마크 점수에 도달하는 비용이 분기당 평균 47%, 연간 약 13배씩 떨어지고 있다고 분석했습니다. 그러나 MIT 연구진은 하드웨어 저렴화와 경쟁 요인을 제거하면 순수 알고리즘 효율 개선은 연간 약 3배에 그친다고 지적합니다. 이는 가격 하락이 전부 실질적 기술 진보를 반영하는 것은 아니며, 모델 선택 시 벤치마크 점수만으로 판단할 수 없음을 시사합니다.

번역된 본문

AI 성능 비용, 역대 어떤 기술보다 빠르게 하락

특정 AI 벤치마크에서 고정된 성능 수준에 도달하는 가격이 2023년 이후 급격히 하락했다. AI 트렌드를 추적하는 연구기관 Epoch AI에 따르면 비용은 분기당 평균 약 47%, 즉 연간 약 13배씩 감소하고 있다. 이 그룹은 그 어느 변혁적 기술도 이렇게 빠르게 가격이 떨어진 적은 없다고 밝혔다. 다만 이 수치는 순수한 알고리즘·아키텍처 발전이나 실제 생산성 비용이 아니라 고정된 벤치마크 점수에 대한 시장 가격을 반영한 것이다.

MIT 연구진은 비교 가능한 데이터에서 연간 5~10배의 비용 하락을 확인했다. 이들이 더 저렴해진 하드웨어와 경쟁에 따른 가격 압박을 제거하면 실제 알고리즘 효율 개선은 연간 약 3배에 그친다고 추정한다. 반면 작업당 최고 성능을 발휘하는 실행은 오히려 비싸질 수 있는데, 최신 추론 모델은 작업당 훨씬 많은 컴퓨팅을 소모하기 때문이다. 두 연구는 서로 다른 질문을 던지고 있다. 작년의 최고 성능을 따라잡는 것은 극적으로 저렴해졌다. 하지만 현재 최고 모델을 실행하는 것은 쿼리당 더 비쌀 때가 많다.

o3 수준의 정확도가 이제는 새 발의 피

Epoch는 OpenAI의 o3를 예로 든다. 2025년 초 o3는 박사 과정 수준의 과학 시험인 GPQA Diamond에서 75%를 기록했으며 질문당 약 30센트의 비용이 들었다. 18개월 후 GPT-5.6 계열 모델은 0.04센트로 같은 점수를 달성했다. Epoch는 이것이 원래 가격의 1/725라고 말한다. 자동차 가격이 이 속도로 떨어졌다면 5만 유로짜리 차가 70유로 미만이 될 것이다. OpenAI는 며칠 전 더 저렴한 GPT-6 Sol과 Luna 모델을 출시했으므로 격차는 더 벌어졌을 것이다.

Epoch는 수학, 과학, 논리 퍼즐을 아우르는 5개 벤치마크에 분석을 기반을 두고 있다. 표본이 좁기 때문에 이 기관은 자신들의 결과를 '가용한 최선의 데이터에 기반한 합리적이지만 대략적인 측정'이라고 부른다.

알고리즘 개선은 하락의 일부만 설명

한스 군들라크(Hans Gundlach)와 MIT 동료들은 비교 플랫폼 인공 분석(Artificial Analysis)의 2024년 4월부터 2025년 11월까지 가격 데이터를 사용하며, 기존 연구보다 테스트당 훨씬 많은 모델을 평가했다. 이들의 수치가 Epoch보다 낮은 이유 중 하나는 최신 추론 모델이 어려운 문제에 추가 테스트 시간 컴퓨팅을 투입해 토큰당 가격이 계속 하락하는데도 정답당 비용을 끌어올릴 수 있기 때문이다. 토큰은 공급자가 과금하는 텍스트 단위로, 이것만 비교하면 전체 그림을 놓치게 된다.

MIT가 가격 하락의 요인을 분해했을 때 저렴해진 하드웨어가 일부를, 경쟁이 그보다 더 많은 부분을 차지했다. 연구진은 오픈 모델을 별도로 분석해 경쟁 요인을 통제했고, 남는 것이 순수 알고리즘 효율 개선으로 연간 약 3배로 산출됐다. Epoch의 13배 수치가 더 높은 것은 하드웨어와 경쟁 효과를 제거하지 않았기 때문이다.

더 높은 벤치마크 점수가 항상 더 나은 효율을 뜻하진 않는다

MIT는 일부 성능 향상이 단순히 더 많은 컴퓨팅을 소비한 결과라는 점도 발견했다. 새 모델이 GPQA Diamond에서 이전 모델을 능가하면 겉으로는 발전처럼 보이지만, 저자들은 개선의 상당 부분이 질문당 더 많은 처리 능력을 사용한 데서 나온다고 추정한다. 점수는 높지만 실행 비용도 더 크다. 코딩과 수학 벤치마크에서도 같은 패턴이 소규모로 나타난다.

모든 벤치마크 진보가 효율 진보인 것은 아니다. 새 모델은 더 나은 학습, 더 나은 데이터, 더 나은 아키텍처, 더 많은 테스트 시간 컴퓨팅이 뭉쳐 있는데, 단일 점수는 이 모든 것을 섞어버린다. '벤치맥싱(benchmaxxing)' 문제도 있다. AI 기업들이 잘 알려진 테스트에 최적화해 실제 성과 없이 점수를 부풀릴 수 있는 것이다. Epoch는 비밀리에 유지된 게임을 기반으로 한 '미스터리 게임 퍼즐'이라는 테스트를 포함해 이에 대비한다. 이 테스트에서 비용 하락이 가장 더뎌, 벤치맥싱 가설과 부합하지만 태스크 형식이나 데이터 노이즈를 반영하는 것일 수도 있다.

가격만으로 어떤 모델을 선택할지 알 수는 없다

원문 보기
원문 보기 (영어)
AI performance costs are falling faster than those of any previous technology Manuel Uth Sep 24, 2026 Nano Banana Pro prompted by THE DECODER The price of reaching a fixed performance level on select AI benchmarks has dropped sharply since 2023. Epoch AI, a research organization that tracks AI trends, says costs are falling by about 47 percent per quarter on average, or about 13x per year. No other transformative technology declined that fast, the group says. The number reflects market prices for a fixed benchmark score, though, not pure algorithmic or architectural progress or even real-life productivity costs, which is a whole other story . MIT researchers looking at comparable data see costs dropping 5x to 10x annually. Once they strip out cheaper hardware and competitive pricing pressure, they put the actual gain in algorithmic efficiency at about 3x per year. Peak performance per run can actually get pricier, because newer reasoning models burn through a lot more compute per task. Both studies are asking different things, though. Matching last year's top-of-the-line capability? Dramatically cheaper. Running the current best model? Often significantly more per query. Matching o3's accuracy now costs a fraction of the price Epoch uses OpenAI's o3 as an example. In early 2025, o3 scored 75 percent on GPQA Diamond, a PhD-level science test, at an estimated 30 cents per question. Eighteen months later, a GPT-5.6 family model hit the same score for four hundredths of a cent. Epoch says that's 1/725 of the original price. If cars dropped that fast, a 50,000-euro vehicle would cost less than 70 euros. OpenAI launched the even cheaper GPT-6 Sol and Luna models just days ago, so the gap has likely widened further. Epoch bases its analysis on five benchmarks spanning math, science, and logic puzzles. Because that's a narrow sample, the organization calls its findings "reasonable but rough measurements based on the best available data." Algorithmic gains only explain part of the drop Hans Gundlach and his MIT colleagues use pricing data from the comparison platform Artificial Analysis, covering April 2024 through November 2025, and evaluate far more models per test than previous research. Their numbers are lower than Epoch's partly because newer reasoning models can throw extra test-time compute at hard problems, driving up the cost per correct answer even while per-token prices keep falling. Tokens are the text units providers bill for, and comparing them alone misses the full picture. When MIT breaks down what's pushing prices lower, cheaper hardware accounts for some of it, competition for more. The researchers control for competition by looking at open models separately. What's left is the pure algorithmic efficiency gain, which comes out to about 3x per year. Epoch's 13x figure is higher because it doesn't strip out hardware and competition effects. Better benchmark scores don't always mean better efficiency MIT also found that some performance gains simply come from spending more compute. A new model that beats its predecessor on GPQA Diamond looks like progress from the outside, but the authors estimate a big chunk of the improvement comes from using more processing power per question. It scores higher and costs more to run. Coding and math benchmarks show a smaller version of the same pattern. Not all benchmark progress is efficiency progress. New models roll together better training, better data, better architecture, and more test-time compute. A single score blends all of that. There's also the "benchmaxxing" problem : AI companies could optimize for well-known tests, inflating scores without real-world payoff. Epoch tries to guard against this by including one test, "Mystery Game Puzzles," based on a game that's been kept secret. Costs drop slowest on that test, which fits the benchmaxxing theory but could just as easily reflect the task format or noise in the data. Price alone won't tell you which model to pick Broad cost averages don't help much when you're choosing a model for a specific job. Platforms like Artificial Analysis rank models across quality, price, latency, context window, and output speed, and the cheapest option rarely wins on every dimension. A low-cost model with high latency is useless for a real-time chatbot. A powerful reasoning model might be too slow for automated workflows. A pricier frontier model could still save money if it gets things right more often and cuts down on retries. None of that shows up in a simple price-per-token comparison. We dug into this token economics question in Frontier Radar #3 . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->