메뉴
BL
The Decoder 39일 전

최신 벤치마크가 폭로한 AI의 실무 한계

IMP
8/10
핵심 요약

인공지능 분석 업체가 공개한 새로운 벤치마크에 따르면, 현재 가장 성능이 좋은 AI 모델조차 복잡한 실무 환경에서는 완벽히 해결하지 못하는 것으로 나타났습니다. 수천 개의 분산된 자료를 종합해야 하는 지식 노동 작업에서 최상위 모델의 완벽한 성공률은 단 3%에 불과했으며, 이는 실제 업무 적용 시 AI의 현재 한계를 명확히 보여줍니다.

번역된 본문

새로운 벤치마크, 실제 지식 노동에서 AI가 얼마나 크게 어려워하는지 폭로하다 Maximilian Schreiner / 2026년 6월 19일

최고 성능의 AI 모델조차 현실적인 지식 노동(Knowledge work) 환경에서는 실패하여, 모든 기준을 완벽히 충족한 작업은 단 3%에 불과했습니다. 인공지능 분석 업체(Artificial Analysis)가 공개한 새로운 'AA-Briefcase' 벤치마크는 AI 모델을 슬랙(Slack) 스레드, 이메일, 회의록, 대용량 데이터 추출물과 같은 수천 개의 파편화된 소스 파일로 구성된 '수주 단위의 지식 노동 프로젝트' 과정에 투입하여 평가합니다.

가장 높은 성능을 기록한 'Claude Fable 5' 모델은 평가 기준에서 가장 높은 통과율을 기록했음에도 불구하고, 모든 기준을 완벽하게 충족한 작업은 단 3%에 그쳤습니다. 총 91개의 작업 중 31개 작업에서는 어떤 모델도 50%의 기준 통과율을 넘지 못했습니다.

모델의 성능이 향상될수록 오류의 유형도 달라집니다. 성능이 낮은 모델은 관련 파일을 누락하거나 사용할 수 없는 결과물을 내놓는 등 기본적인 실행 과정에서 난항을 겪습니다. 반면 성능이 높은 모델은 표면적인 요구 사항은 충족하지만, 여러 출처의 정보를 조립해야만 파악할 수 있는 세부 사항을 놓치면서 더 조용하게(눈에 띄지 않게) 실패하는 양상을 보입니다.

또한 가격 측면에서도 상당한 격차가 존재합니다. 작업당 비용은 DeepSeek V4 Flash의 약 0.04달러에서부터 Claude Fable 5의 31달러 이상까지 800배 이상 차이가 납니다.

원문 보기
원문 보기 (영어)
New benchmark exposes how badly AI struggles with real knowledge work Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jun 19, 2026 Even the best AI model fails at realistic knowledge work, fully solving just 3 percent of tasks. The new AA-Briefcase benchmark from Artificial Analysis puts AI models through multi-week knowledge work projects built from thousands of fragmented source files like Slack threads, emails, meeting transcripts, and large data exports. The top performer, Claude Fable 5, hits the highest rubric pass rate but still nails all criteria on just 3 percent of tasks. On 31 out of 91 tasks, no model even clears 50 percent. The types of errors shift as models get better. Weaker models choke on basic execution as they miss relevant files or spit out unusable results. Stronger models fail more quietly, as they hit the obvious requirements but miss details you'd only catch by piecing together information from multiple sources. Ad There also is a significant price gap: Per-task costs span more than 800x, from about $0.04 for DeepSeek V4 Flash to over $31 for Claude Fable 5. Ad DEC_D_Incontent-1 AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: AA Ask about this article… Search