메뉴
BL
The Decoder 50일 전

에이전트 AI가 토큰을 새로운 비즈니스 지표로 만드는 법

IMP
8/10
핵심 요약

자율적으로 작동하는 에이전트 AI의 등장으로 기존 정액제 구독 모델이 붕괴하고, 실제 사용량인 '토큰' 기반의 과금 모델이 빠르게 자리 잡고 있습니다. 기업들은 막대한 인프라 투자비를 회수해야 하며, 단순한 채팅과 달리 수십~수백만 개의 토큰을 소모하는 에이전트 워크플로우에 맞춰 비용 구조를 세분화하고 있습니다. 하지만 토큰 사용량이 단순한 '활동량'의 지표일 뿐 실제 비즈니스 '가치'를 나타내지 못한다는 점이 새로운 산업의 과제로 떠오르고 있습니다.

번역된 본문

Frontier Radar #3: 에이전트 AI가 토큰을 비즈니스 지표로 바꾸는 방법 Maximilian Schreiner & Matthias Bastian / 2026년 6월 8일

월간 구독을 하고, 채팅창을 열어 질문을 던지는 것. 이것이 지금까지 생성형 AI가 작동하던 방식이었습니다. 하지만 에이전트 워크플로우(Agentic workflows)는 이 모델을 완전히 깨버립니다. 에이전트는 훨씬 더 많은 토큰을 소모하고, 수시간 동안 자율적으로 실행되므로 서비스 제공업체는 더 이상 정액제(Flat rate)를 유지할 수 없게 되었습니다. 동시에 토큰 가격은 처리 속도, 전문성, 그리고 경제적 가치에 따라 세분화되고 있습니다. 비용은 점점 더 정확해지지만, 그 혜택은 여전히 모호한 경우가 많습니다. 그 결과, 토큰 사용량은 실제 성과가 아닌 단순한 활동량만 측정함에도 불구하고 가치 창출을 대변하는 대체 지표가 되어가고 있습니다.

THE DECODER의 편집팀은 1년에 6차례 뉴스레터와 구독자 전용 웹사이트를 통해 심층적인 AI 주제를 다루는 'Frontier Radar'를 발행합니다. 이번 3호에서는 생성형 AI의 떠오르는 토큰 경제(Token economy)를 조명합니다. (1호에서는 에이전트 AI의 현재 상태를, 2호에서는 AI가 생산성에 미치는 측정 가능한 영향을 다룬 바 있습니다.)

오랫동안 생성형 AI는 일반적인 소프트웨어처럼 느껴졌습니다. 월간 요금제에 가입하고, 채팅창을 열며, 질문하고 답변을 얻는 방식이죠. 파워 유저들은 항상 API를 통해 개별 요청의 실제 비용을 확인할 수 있었습니다. 그래서 많은 사용자가 과도한 사용 시 훨씬 저렴한 정액제를 선택했습니다. 하지만 대부분의 사용자에게 비용은 보이지 않았습니다. 인간의 사용에는 자연스러운 한계가 있었기 때문에 정애제가 널리 통용될 수 있었습니다. 사람들은 천천히 타자를 치고, 답변을 읽고, 휴식을 취하며, 회의에 참석하고 퇴근합니다.

에이전트(Agent)는 그런 한계를 모릅니다. 파일을 읽고, 도구를 호출하고, 코드를 작성하고, 중간 결과를 확인하고, 오류를 수정한 뒤 다시 시도합니다. 사용자가 원한다면 작업이 완료될 때까지 계속 진행합니다. 서비스 제공업체 측에도 압박이 존재합니다. 거대 AI 기업들은 데이터센터, 칩, 모델 학습에 수천억 달러를 쏟아부었습니다. 이러한 투자는 정액제로는 결코 감당할 수 없는 규모로 수익을 창출해야 합니다.

이번 Frontier Radar 호에서는 이러한 맥락에서 형성되는 새로운 토큰 경제를 조명합니다. 과금 방식이 구독에서 사용량 기반으로 어떻게 전환되고 있을까요? 토큰 자체가 어떻게 세분화된 상품이 되고 있을까요? 그리고 왜 토큰 사용량이 여전히 AI의 가치를 측정하는 데 부적합할까요?

왜 제공업체들은 정액제를 버리고 있을까? 가장 눈에 띄는 변화는 사용량 증가에 대응하여 가격 모델을 전면 개편하고 있다는 점입니다. 2026년 6월 1일부터 GitHub Copilot은 'GitHub AI 크레딧(Credits)'을 도입하며 사용량 기반 모델로 점진적으로 전환하고 있습니다. 이 크레딧은 실제 토큰 사용량 및 각 모델의 API 가격과 연동됩니다. 이는 Copilot이 단순한 코드 제안을 넘어, 주로 채팅, CLI, 에이전트 기능을 수행할 때 적용됩니다. 단, 유료 요금제에서 기본적인 코드 자동완성 기능은 계속 무료로 제공됩니다.

GitHub가 밝힌 이유가 문제의 핵심을 찌릅니다. 과거에는 짧은 채팅 질문과 몇 시간 동안 자율적으로 실행되는 코딩 세션이 거의 같은 비용으로 취급되었습니다. 이런 방식은 더 이상 지속될 수 없습니다.

Anthropic 역시 일반적인 사용과 에이전트 워크플로우 사이의 경계를 명확히 하고 있습니다. Claude Code, Claude Cowork, Managed Agents는 Claude를 하나의 디지털 노동자로 만듭니다. Anthropic은 최대 100만 개의 컨텍스트 토큰(Context tokens)과 최대 부하 상황에서 발생한 Claude Code의 병목 현상을 지적했습니다. 기존 요금제는 채팅을 많이 사용하는 사용자에게는 적합했지만, 항상 켜져 있는(Always-on) 에이전트 워크플로우에는 맞지 않았던 것입니다.

분야에 따라 사용량이 얼마나 크게 차이 나는지는 Anthropic의 공개 API 분석에서도 나타납니다. 모든 에이전트 도구 호출의 거의 절반이 소프트웨어 개발 분야에 집중되어 있습니다. 이 분야는 Claude Code와 같은 에이전트 모델과 스캐폴딩(Scaffolding)의 혜택을 가장 먼저 받은 곳입니다. 반면 고객 서비스, 영업, 재무 및 전자상거래 분야는 각각 불과 몇 %에 불과합니다. 이들 분야에서는 여전히 단순한 채팅 요청이 주를 이룹니다. 그러나 오피스, 리서치, 재무 및 법률 도구에서 에이전트 워크플로우가 성숙해지면서 이러한 격차는 더욱 벌어질 가능성이 높습니다. 이와 함께 토큰 비용은 현재까지는 크게 체감되지 않았던 영역으로까지 확장될 것입니다.

토큰 가격만으로는 오해하기 쉬운 이유 이러한 변화는 비용에 대한 근본적인 의문을 제기합니다. AI가 주로 채팅 도구로 사용되던 시절에는 토큰당 가격이 단순한 지표로 작용했습니다.

원문 보기
원문 보기 (영어)
Frontier Radar #3: How agentic AI is turning tokens into a business metric Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner & Matthias Bastian Jun 8, 2026 Nano Banana Pro prompted by THE DECODER Monthly subscription, open chat, ask a question: that's how generative AI worked until now. Agentic workflows blow up this model. They burn through far more tokens, run autonomously for hours, and make flat rates untenable for providers. At the same time, token prices are splitting along axes of speed, specialization, and economic value. But while costs get more precise, the benefits often stay vague. The result: token usage becomes a stand-in metric for value creation, even though it only measures activity, not outcomes. Six times a year, THE DECODER's editorial team takes an in-depth look at a fundamental AI topic in its "Frontier Radar," as a newsletter and exclusively here on the site for THE DECODER subscribers . Issue #3 covers the emerging token economy of generative AI. Issue #1 looked at the current state of agentic AI. Issue #2 examined the measurable impact of AI on productivity. For a long time, generative AI felt like classic software. Sign up for a monthly plan, open a chat, ask a question, and get an answer. Power users could always see through APIs what individual requests actually cost. That's why many of them went with flat rates, which were much cheaper under heavy use. For most users, though, the costs stayed invisible. Flat rates worked broadly because human usage has natural limits. People type slowly, read answers, take breaks, go to meetings, and clock out. An agent doesn't know those limits. It reads files, calls tools, writes code, checks intermediate results, fixes errors, and tries again. If the user wants, it keeps going until the task is done. There's also the pressure on the provider side: The big AI companies have poured hundreds of billions of dollars into data centers, chips, and model training. Those investments have to pay off, at a scale that flat rates simply can't support. This issue of the Frontier Radar maps out the emerging token economy along these lines. How is billing shifting from subscription to usage? How is the token itself becoming a segmented product? And why is token usage still a poor measure of AI value? Why providers are walking away from flat rates The most visible change is the overhaul of pricing models in response to growing usage. Starting June 1, 2026, GitHub Copilot is gradually moving to a usage-based model with " GitHub AI Credits ." The credits are tied to actual token usage and the API prices of each model. They kick in wherever Copilot does more than just suggest code, mainly in chat, CLI, and agent features. Standard completions stay free of these rules in paid plans. GitHub's reasoning nails the problem: a short chat question used to be treated about the same as an autonomous coding session running for hours. That can't last. Anthropic is also drawing a sharper line between normal use and agentic workflows. Claude Code, Claude Cowork, and Managed Agents turn Claude into a digital worker. Anthropic blamed bottlenecks at Claude Code on peak loads and contexts of up to one million tokens. The older plans fit heavy chat use but not always-on agent workflows. How sharply usage differs between fields shows up in Anthropic's own analysis of its public API : nearly half of all agentic tool calls go to software development, the area that first benefited from agentic models and scaffolding like Claude Code. Customer service, sales, finance, and e-commerce each sit at just a few percent. Simple chat requests still dominate there. That spread will likely widen as agentic workflows mature in office, research, finance, and legal tools. With it, the token bill moves into areas where it isn't yet felt today. Why the token price alone is misleading This development shifts the cost question: As long as AI was mainly used as a chat tool, the price per token could feel like a technical footnote. In agentic workflows, it becomes a business metric. The most obvious mistake in the new token economy is a flat price comparison. GPT-5.5 costs $30 per million output tokens, DeepSeek V4 Pro 87 cents. That says little about actual costs in use. Beyond price per token, what matters is consumption per task. Like with a car, the price of gas alone tells you nothing about what a drive from Berlin to Munich costs. You also have to know the distance and the mileage. A cheap model can get expensive if it needs more tries, fails more often, or requires more cleanup. A pricier model pays off when it gets to the goal with fewer loops and needs less human oversight. Benchmarks and other analyses make this clear. GPT-5.5, for instance, was supposed to offset part of its higher list price with shorter answers. An analysis of real-world usage by OpenRouter still showed cost increases of 49 to 92 percent over its predecessor , depending on input length. Of course, both can rise: the token price and the number of tokens consumed, as with Google's Gemini 3.5 Flash . Here, the token price jumped threefold over the predecessor Gemini 3 Flash. In Artificial Analysis's evaluation, the model also needed more steps in the Intelligence Index run. The result: in that test, it ended up more expensive than Google's current flagship, Gemini 3.1 Pro. Pushing the other way is the price pressure from providers like DeepSeek . Behind the rock-bottom prices is a bet of its own: if you pay only a fraction per token, you can run the same job four or five times and still come out cheaper. As long as the final result holds up, that's attractive. Where it doesn't, rework quickly eats the price advantage. How the token market is splitting by performance class The more the market splits, the less sense it makes to talk about "the" token price. The price per million tokens still matters but only says something within a clear performance class. A fast token in a coding agent, a cheap token in a mass-market app, and a specialized token in security analysis can be billed in similar technical fashion, but they're different economic products. Different model tiers and subscription levels have existed for a while. What's new is that the differentiation now spans more axes: latency, processing mode, context size, agent runtime, specialization, and increasingly the economic value of the output. Providers aren't just selling compute time in token form anymore. They're selling different inference services. The scarcer, faster, or more valuable that service is, the further the price can drift from raw compute costs. Nvidia CEO Jensen Huang spelled this out in two recent interviews. On Dwarkesh Patel's show, he explains why Nvidia recently licensed the inference architecture of startup Groq and folded it into its own CUDA ecosystem. The reason is economic: the value of a token has risen so much that different prices for different token types now make sense. Back in the old days, just a couple of years ago, Tokens were either free or barely expensive. But now you can have different customers, and those customers want different answers. Because the customers make so much money - for example, our software engineers - if I can give them much more responsive Tokens so that they're even more productive than they are today, I would pay for it. Jensen Huang, Nvidia Huang is describing the technical side of this segmentation. Premium inference with lower latency pays off because tokens at the top of the market can command much higher prices. Nvidia talks about expanding the Pareto front: multiple optimal points of price and speed, depending on customer segment. Where the value comes from the possible outcome, there is more segmentation possible. According to The Information, Palo Alto Networks tested Anthropic's security model Mythos to scan its own source code for vulnerabilities. The model reportedly found more than two dozen critical vulnerabilities in a