메뉴
BL
VentureBeat AI 12일 전

AI 에이전트 평가의 함정: 검증 없이 무리하게 도입되는 기업들

IMP
8/10
핵심 요약

최근 조사에 따르면 기업들은 AI 에이전트의 자율성은 확대하고 있지만, 이를 통제할 평가 시스템에 대한 신뢰는 매우 낮은 상태입니다. 내부 테스트를 통과했음에도 실 서비스에서 잦은 오류가 발생함에도 불구하고, 인간의 개입 없이 자동화된 평가만으로 에이전트를 배포하려는 움직임이 가속화되고 있어 위험한 '평가 격차(Evaluation gap)'가 발생하고 있습니다. 이는 기업이 자율성에 걸맞은 안전장치와 검증 기술을 제대로 갖추지 못한 상태에서 무리하게 AI를 도입하고 있음을 시사합니다.

번역된 본문

157개 기업을 대상으로 한 조사 결과, 조직들은 AI 에이전트에 더 많은 자율성을 부여하면서도 정작 이러한 자율성을 통제하기 위해 마련된 평가(Evaluation) 시스템을 신뢰하지 않는 것으로 나타났습니다. 절반은 이미 내부 평가를 통과한 에이전트를 출시했다가 실제 프로덕션 환경에서 고객에게 오류를 노출한 경험이 있으며, 오늘날 자동화된 평가 시스템을 완전히 신뢰하는 곳은 20곳 중 1곳에 불과했습니다. 그리고 가장 큰 약점으로 지적된 것은 평가 결과가 실제 환경의 결과와 일치하지 않는다는 것입니다. 그럼에도 불구하고 3분의 2에 달하는 기업은 이미 인간의 개입(Human in the loop) 없이 자동화된 평가만으로 프로덕션 환경에 에이전트 변경 사항을 배포하도록 허용하거나 적극적으로 엔지니어링하고 있습니다. 그 결과 '평가 격차(Evaluation gap)'가 발생하고 있는데, 이는 기업이 에이전트에 부여하는 자율성의 크기와 장애를 방지하기 위해 마련된 테스트에 대해 갖는 신뢰도 사이의 거리를 의미합니다.

이번 벤처비트(VentureBeat) 펄스 리서치(Pulse Research) 물결은 기술 리더들이 에이전트의 성능을 어떻게 측정하는지 조사합니다. 어떤 신뢰성 및 평가 플랫폼을 사용하는지, 어떻게 선택하고 신뢰하는지, 프로덕션 환경에서 무엇이 문제를 일으키는지, 그리고 인간의 개입 없이 에이전트를 어디까지 자율적으로 실행하도록 허용할 의향이 있는지를 살펴봅니다.

핵심 발견은 바로 '평가 격차(Evaluation gap)', 즉 기업이 에이전트에 부여하는 자율성과 이를 통제하기 위한 평가 시스템에 대한 신뢰 사이의 거리입니다. 절반의 조직(50%)이 지난 1년 동안 내부 평가를 통과한 에이전트나 LLM 기능을 배포했다가 고객을 대면하는 서비스 환경에서 실패를 초래한 경험이 있으며, 4분의 1은 이런 일이 두 번 이상 발생했다고 답했습니다. 테스트 자체에 대한 신뢰도도 매우 낮아서 오늘날 자동화된 평가 시스템을 완전히 신뢰한다고 답한 비율은 단 5%에 불과했습니다. 가장 많이 지적된 한계점은 평가가 실제 환경의 결과와 제대로 일치하지 않는다는 것(29%)이었습니다. 기업들은 평가를 통과하는 것과 실제로 제대로 작동하는 에이전트라는 것이 같은 의미가 아니라는 사실을 깨닫고 있습니다.

이러한 격차를 심각하게 만드는 것은 기술 도입의 방향성입니다. 3분의 2의 조직(66%)이 이미 저위험 에이전트에 대해 인간의 개입이 전혀 없는 완전 자동화된 배포를 허용(34%)하거나, 12개월 이내에 이를 허용하기 위해 파이프라인을 적극적으로 엔지니어링(33%)하고 있습니다. 동시에, 그러한 신뢰를 얻어야 할 평가 스택은 파편화되어 있고 미성숙합니다. 가장 일반적인 주요 도구는 모델 제공업체의 기본 평가 도구였으며, 전용 도구가 아예 없는 경우와 각각 동일한 비율(각 17%)을 차지했습니다. 또한 실제 프로덕션 트래픽에 대해 실시간 품질 검사를 수행하는 기업은 약 4분의 1에 불과했습니다. 즉, 기술이 주는 자율성의 도입 속도가 안정성 확보(Assurance) 속도보다 훨씬 빠르게 진행되고 있습니다.

조사 방법론 벤처비트(VentureBeat)는 지속적인 펄스 리서치(Pulse Research) 시리즈의 일환으로 이 설문 조사를 실시했습니다. 이번 '에이전트 신뢰성 및 평가(Agentic Reliability & Evals) 트래커' 설문은 기술 리더들이 에이전트의 성능과 신뢰성을 어떻게 평가하는지에 초점을 맞췄습니다. 응답은 직원 수가 100명 이상인 조직(n=157)으로 필터링되었으며, 2026년 6월 단일 설문 조사에서 추출되었습니다. 이는 여러 달에 걸친 통합 샘플이 아닌 단일 조사이므로, 이 보고서는 횡단면으로 읽혀지며 월별 추세를 추론하지 않습니다. 질문이 다중 선택인 경우, 그 비율의 합은 100%를 초과할 수 있습니다. 역할별로 샘플은 시니어급이며 구매 결정권자로서 신뢰할 수 있습니다. 응답자 중 38%는 최종 의사 결정권자입니다.

원문 보기
원문 보기 (영어)
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap — the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures.This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop.The central finding is an evaluation gap — the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. Half of organizations (50%) have, in the past year, deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure, and a quarter have seen it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%). Enterprises are discovering that a passing eval is not the same as a working agent.What makes the gap consequential is the direction of travel. Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). At the same time, the evaluation stack that would have to earn that trust is fragmented and immature: the most common primary tools are the model providers’ native evals, tied with having no dedicated tooling at all (17% each); and only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey — the Agentic Reliability & Evals tracker — focused on how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Where questions were multiple-select, those shares can sum to more than 100%.By role the sample is senior and buyer-credible: 38% are final decision-makers fo