메뉴
HN
Hacker News • 49일 전

데이터브릭스가 AI 코딩 비용을 70% 절감한 방법

IMP
8/10
핵심 요약

데이터브릭스(Databricks)는 AI 코딩 도구를 도입하여 개발 속도를 크게 향상시켰으나, 기하급수적으로 증가하는 비용 문제에 직면했습니다. 이를 해결하기 위해 지속적으로 효율성이 높은 모델(오픈소스 및 저비용 모델)로 트래픽을 전환하는 등의 인프라 최적화 기법을 통해 사용자당 비용을 일정 수준으로 억제하면서도 AI 도구의 혜택을 극대화하는 전략을 공유했습니다.

번역된 본문

본문으로 건너뛰기

AI 코딩 도구는 엄청난 가치를 제공합니다. 데이터브릭스(Databricks)에서 에이전틱 코딩(agentic coding)은 우리가 추적하는 모든 속도 지표를 눈에 띄게 개선했으며, 일부 팀에서는 산출물을 한 자릿수 이상(10배 이상) 향상시켰습니다. 하지만 대규모로 AI 도구를 도입하는 거의 모든 기업이 동일한 장벽에 부딪혔습니다. 바로 기하급수적으로 증가하는 비용입니다. 이러한 증가 추세는 지속 불가능하며, 방치할 경우 결국 수익을 압도하게 될 것입니다.

비용 폭발은 기업들을 딜레마에 빠지게 했습니다. 한편으로는 AI 전환을 최대한 추진하고 직원들에게 강력한 도구를 쥐여주고 싶어 하지만, 다른 한편으로는 AI가 제공하는 효율성 이득을 훼손하거나 심지어 역전시킬 위협을 가진 총비용 프로필을 타협해야만 했습니다.

다행히도, 가장 먼저 대규모로 도입한 몇몇 기업들은 이 퍼즐을 해결하고 '이중 목표'를 달성하는 일련의 접근 방식에 수렴했습니다. 그 두 가지 목표는 (a) 마찰을 최소화하며 AI 도구에 대한 폭넓은 접근성을 제공하는 것, 그리고 (b) 총비용을 사용자당 대략 고정된 한도 내로 유지하는 것입니다. 이 글은 데이터브릭스에서의 경험과 스트라이프(Stripe), 코인베이스(Coinbase), 우버(Uber), 램프(Ramp) 등 여타 디지털 네이티브 기업들과의 대화를 바탕으로 검증된 비용 관리 기법을 간략히 설명합니다. 아래 표는 개발팀을 대상으로 비공식적으로 설문 조사한 방향성 수치를 바탕으로 현재의 기법과 그에 따른 절감액을 요약한 것입니다.

이러한 기법 중 일부는 많은 기업이 이미 사용 중인 소프트웨어로 쉽게 구현할 수 있습니다. 하지만 엔드 유저 클라이언트를 수정하거나 모델 간에 트래픽을 전환하는 등의 기법은 새로운 인프라가 필요합니다. 데이터브릭스에서는 핵심 인프라 구성 요소인 엔드 유저 메타-하네스(Omnigent)와 AI 게이트웨이(Unity AI Gateway)를 오픈소스화하거나 무료로 공개했습니다. 완전성을 위해 이 글에서는 우리가 대화한 다른 기업들이 사용하는 소프트웨어도 다룹니다.

코딩 모델을 위한 '효율성 프론티어(Efficiency Frontier)'

코딩 비용을 새롭게 출시되는 더 효율적인 모델로 이전하는 것이 비용을 줄이는 가장 큰 단일 요인입니다. '더 저렴한 모델'이라는 단순한 설명이 모델 비용과 품질 간의 미묘한 관계를 감추고 있기 때문에, 이 부분은 약간의 논의가 필요합니다. 일반적으로 '프론티어 모델(frontier model)'이라는 용어는 '가장 지능적인 모델'을 의미하며, 프론티어 AI 연구소들은 주로 최고 수준의 지능을 발전시키는 데 집중합니다. 프론티어 모델은 이제 수학이나 사이버 보안 분야의 새로운 문제를 해결할 수 있습니다. 하지만 AI가 대규모로 배포될 때는 다른 유형의 프론티어가 더 중요해집니다. 바로 '효율성 프론티어'입니다.

효율성 프론티어는 주어진 지능 수준에서 최적의 가격대를 제공하는 모델들의 집합으로 정의됩니다. 일상적인 코딩의 대부분은 수학적 증명이나 새로운 보안 통찰력을 요구하지 않으므로, 전체적으로 볼 때 중요한 것은 일반적인 소프트웨어 엔지니어링 작업에 필요한 품질 기준을 충족하는 모델의 비용입니다. 이러한 '효율성 프론티어'는 지능 프론티어보다 훨씬 빠르게 발전하고 있으며, 거의 매주 기존 모델보다 단위 가격당 더 나은 지능을 제공하는 새로운 모델이 출시되고 있습니다.

비용 제어 수단 #1: 오픈소스 및 저비용 모델로의 전환

더 새롭고 효율적인 모델을 빠르게 도입하면 다른 어떤 기술보다 큰 비용 절감 효과를 거둘 수 있습니다. 하지만 이러한 이득을 얻으려면 기업은 먼저 어떤 모델이 실제로 기존 모델보다 더 나은지 알아야 합니다. 공개 벤치마크는 코딩 작업에 대한 실제 성능을 제대로 보여주지 못하기 때문에 이를 파악하는 것은 어려울 수 있습니다. 새로운 모델을 평가하기 위해 많은 기업들은 자사의 내부 개발 환경을 더 잘 반영한다고 판단하는 자동화된 평가 시스템을 구축했습니다. 데이터브릭스는 최근 이러한 벤치마크의 예시를 게시했는데, 여기서 우리는 GLM 모델의 매우 경쟁력 있는 가성비를 관찰했습니다. 이 벤치마크를 바탕으로 우리는 내부 개발자들에게 GLM 모델을 롤아웃했습니다.

물론 새로운 모델이 항상 효율성 프론티어를 발전시키는 것은 아니며, 평가를 통해 부정적인 결과가 나오기도 합니다. 예를 들어, 스트라이프(Stripe)는 Opus 4.7이 Opus 4.6에 비해 품질을 크게 향상시키지 못하면서 비용만 증가시킨다는 것을 발견했습니다. 따라서 그들은 Opus 4.7을 내부에 도입하지 않기로 결정했습니다. 데이터브릭스에서도 비슷한 결과를 확인했습니다.

원문 보기
원문 보기 (영어)
Skip to main content AI coding tools deliver immense value: at Databricks, agentic coding has measurably improved every velocity metric we track and, in some teams, driven an order-of-magnitude gains in output. But nearly every company deploying AI tools at scale has hit the same wall: exponentially growing costs . That curve is unsustainable - left unchecked it will eventually overtake revenue. The spend explosion has left enterprises in a paradoxical situation: on the one hand, desiring to maximally push AI transformation and put powerful tools in the hands of employees, and on the other hand, having to reconcile with an aggregate cost profile that threatens to undermine or even reverse the very efficiency gains AI provides. Fortunately, several of the earliest large-scale adopters have converged on a set of approaches that solve this puzzle, achieving a “dual mandate”: (a) providing broad access to AI tooling, with minimal friction, and (b) keeping aggregate costs inside of a roughly fixed envelope per user. This post outlines proven cost management techniques, based on our experience at Databricks and conversations with several other digital-native companies, including Stripe, Coinbase, Uber, and Ramp. The table below summarizes current techniques and associated savings; the numbers are directional, based on an informal survey of development teams: Some of these techniques can be easily implemented with software many companies already use. Others require new infrastructure, particularly techniques that modify end-user clients or shift traffic across models. At Databricks, we’ve open sourced or made freely available our key infrastructure components: an end user meta-harness ( Omnigent ) and our AI Gateway ( Unity AI Gateway ). For completeness, this post also covers software used by other companies we spoke with. The “Efficiency Frontier” for Coding Models The single greatest cost lever in moving coding spend to more efficient models as they are released. This point bears some discussion, as the simple explanation of "cheaper models” in fact hides a nuanced relationship between model cost and quality. Colloquially, the term frontier model means “the highest intelligence model,” and frontier labs largely focus on advancing peak intelligence. Frontier models can now solve novel problems in math or cybersecurity. But when AI is deployed at scale, a different type of frontier matters more: the efficiency frontier . The efficiency frontier is defined by the set of models that have the best price point for a given level of intelligence . Most day-to-day coding doesn't require mathematical proofs or novel security insights, so what matters in aggregate is the cost of models that meet the quality bar for typical software engineering work. This "efficiency frontier” is advancing far faster than the intelligence frontier, with new models being released almost weekly that present better intelligence-per-unit-price than prior models. Cost Lever #1: Moving to open source and lower cost models Rapidly adopting newer, more efficient models delivers the largest cost wins of any technique. But to capture those gains, a company first needs to know which models actually beat its incumbents. This can be difficult because public benchmarks do a poor job of indicating real-world performance on coding tasks. To size up new models, many companies have built automated evaluations that they believe are more representative of their internal development mix. Databricks recently published an example of such a benchmark , in which we observed highly competitive price/performance for GLM models. That benchmark led us to roll GLM out to developers internally. Often, new models do not advance the efficiency frontier,and evaluations frequently produce negative results: Stripe found that Opus 4.7 did not meaningfully improve quality over Opus 4.6, while increasing cost. They therefore declined to make Opus 4.7 available internally. Databricks saw similar cost regressions when comparing Opus 5.0 to 4.8. Harness and Model Flexibility Since the biggest wins come from switching to new models, adopting end user tooling that allows for model flexibility is becoming a critical component of keeping costs down. The tool most commonly used in concern with a particular model is called harness . Proprietary frontier models are increasingly co-designed to work well with specific harnesses, meaning certain harnesses “work better” with certain models. If a company wants to preserve model independence there are roughly two approaches: Ask users to switch harnesses. One approach is to provide developers with a set of harnesses (Claude Code, Codex, or Cursor) and then ask them to switch between harnesses when a company wants to migrate spend to lower cost models. This lets users work in their preferred harness when possible, but the downside of this approach is that switching costs for an individual developer can be high. If switching costs become too high, the harness itself becomes a de facto lock-in to a model family, limiting the ability to move spend to more competitive models. Use a meta-harness. A new and increasingly popular approach is to use a meta-harness that surfaces a common user experience to developers while dispatching requests to underlying harnesses (both proprietary and open source). This approach allows both model/harness independence while also reducing developer switching costs. At Databricks, this is the default mode for developers who leverage Omnigent . Some companies we talked to have built custom internal meta-harnesses that integrate with their development toolchain. Cost Lever #2: Dynamic Request and Task Routing Instead of asking users to choose task-appropriate models themselves, a growing body of research suggests that automatic model and tool selection may further squeeze efficiency out of agentic coding workflows. Routing approaches roughly fall into three categories: Request Level Routing: A stateful proxy sits in between a client (such as a coding harness) and the underlying foundation models. The proxy attempts to route requests to the lowest-cost model capable of answering each inference request. Routing for agentic use cases also needs to account for server-side caching, since a cold cache hit has a very high cost for large context workloads. A new wave of products is showing early, promising results for routing. Examples are: Cursor Router , OpenRouter’s AutoRouter , Ramps Router feature and Databricks own Smart Routing feature in Unity AI Gateway . Task Level Routing (Meta Harness): A client-side process dispatches user tasks to different harnesses based on the complexity of the task. A user task might be “rename this component from X to Y” (a simple task) or an open-ended task like “Explore design considerations that would reduce latency” (a complex task). The dispatcher, often called a Meta Harness , examines which level of underlying model is required for a task and then delegates that entire end-to-end task to the model. Omnigent is an example of a Meta Harness that supports this pattern. Escalation/Delegation Patterns: A single harness pairs two models (an expensive, high-intelligence model and a cheap worker model). In some approaches, such as Claude’s Advisor Tool , the cheaper model runs the show and escalates when it thinks a task requires more horsepower. The inverse pattern also exists: In Cognition’s Devin Fusion , the higher cost model is the main loop, and it selectively outsources work to a cheaper model. Internal results at Databricks suggest that our AI Gateway Smart Router is able to consistently reduce average task cost by more than 30%, while roughly matching the quality of the most expensive model in the working set. Other companies we spoke with have seen similar results. Cost Lever #3: Giving developers visibility, tripwires, and budgets It may be surprising that this entire article did not start and end with “Give users a mont