메뉴
BL
The Decoder 26일 전

프리랜서 업무 16%를 전문가 수준으로 수행하는 AI

IMP
8/10
핵심 요약

AI 에이전트가 실제 프리랜서 프로젝트를 고객이 만족할 만한 전문가 수준으로 완수하는 비율이 8개월 만에 2.5%에서 16.1%로 급증했습니다. 이는 AI가 단순 텍스트 생성을 넘어 전문 디자인, 3D 모델링, 코딩 등 실무 영역으로 빠르게 확장하고 있음을 시사합니다. 하지만 작업물을 평가하고 검수하는 과정에서는 여전히 인간 전문가의 개입이 필수적입니다.

번역된 본문

AI 에이전트가 이제 프리랜서 업무의 16%를 전문가 수준으로 완수할 수 있게 되었습니다. 8개월 전인 2.5%에서 크게 증가한 수치입니다.

막시밀리안 슈라이너(Maximilian Schreiner) / 2026년 7월 2일

'원격 노동 지수(Remote Labor Index, RLI)'는 AI 에이전트가 유료 프리랜서 프로젝트를 전문가 수준으로 완료하는 빈도를 측정합니다. 지난 8개월 동안 최고 자동화율은 4배 이상 급증했습니다.

RLI는 AI 에이전트가 실제 상업적 가치가 있는 프리랜서 작업을, 비용을 지불하는 실제 고객이 수용할 수 있는 품질 수준으로 완료할 수 있는 빈도를 추적합니다. 이 벤치마크는 3D 및 CAD, 건축, 그래픽 디자인, 비디오 및 애니메이션, 오디오, 데이터 분석, 웹 애플리케이션 등의 영역을 다룹니다. 여기에는 검증된 358명의 프리랜서로부터 수집된 총 14만 4천 달러 상당의 240개 프로젝트가 포함됩니다. AI 안전 센터(Center for AI Safety)의 인간 평가자들은 각 결과물을 유료 전문가가 만든 '골드 스탠다드(최고 기준)'와 대조하여 채점합니다. RLI는 Scale Labs와 함께 개발되었습니다.

핵심 지표는 자동화율로, AI의 작업 결과물이 최소한 인간의 결과물만큼 좋다고 평가된 프로젝트의 비율을 의미합니다.

최고 자동화율, 2.5%에서 16.1%로 급증

이 벤치마크가 처음 출시되었을 때, 최고의 AI 에이전트는 단 2.5%의 프로젝트만 자동화했습니다. 최신 결과에 따르면 현재 Fable 5는 사상 최고 점수인 16.1%를 기록했습니다. 이는 Opus 4.8의 8.3%의 거의 두 배에 달하는 수치입니다. GPT-5.5는 6.3%를 기록했습니다. 이 세 모델은 이전에 테스트된 모든 시스템의 성능을 뛰어넘었습니다. 이전 최고 기록 보유자였던 Claude Cowork 프레임워크에서 구동되는 Opus 4.6은 4.17%였습니다. 연구진에 따르면, 8개월도 채 되지 않는 기간 동안 기술 최전선의 성능이 4배 이상 향상된 것입니다.

단, Fable 5의 점수에 대한 주의사항이 있습니다. 미국 정부가 해당 모델에 대한 접근을 제한하기 전까지 240개 프로젝트 중 218개만 평가할 수 있었습니다. Fable 5가 누락된 모든 프로젝트에서 실패한 최악의 시나리오를 가정하더라도, 그 비율은 여전히 14.6%로 다른 어떤 모델보다도 높습니다.

하지만 기술 발전이 모델 출시일과 정확히 일치하는 것은 아닙니다. Scale Labs 리더보드 전체를 보면, 상대적으로 최신인 Gemini 3 Pro는 훨씬 오래된 시스템들보다도 뒤처지는 단 1.25%를 기록하며 최하위권에 머물러 있습니다.

연구의 일부 사례는 최고 수준의 모델조차도 여전히 부족한 점이 있음을 보여줍니다. 반지 디자인 작업에서 Fable 5는 이전 AI들보다 분명히 뛰어나지만, 자세히 살펴보면 여전히 비전문적으로 보입니다. 건축 프로젝트에서는 GPT-5.5가 실제 3D 모델은 결함이 있는 상태로 남겨둔 채 이미지 생성기를 이용해 매력적인 렌더링 결과물을 가장해냈습니다.

여전히 대체 불가능한 인간 평가자

연구팀은 비용이 많이 드는 인간 평가를 AI 심사관(AI judge)으로 대체할 수 있는지 테스트했습니다. 그 결과는 명확했습니다. AI 심사관들은 새로운 모델들의 점수를 지나치게 후하게 주었습니다. GPT-5.5의 경우 AI 평가자가 매긴 점수는 실제보다 거의 3배나 높았습니다. Opus 4.8의 경우에도 약 2.5배 높았습니다. 자동화된 심사관이 순위는 올바르게 매겼지만, 실제 점수 수치는 크게 빗나갔습니다.

CAIS(Center for AI Safety)에 따르면 그 이유는 다음과 같습니다. 납품된 작업을 공정하게 평가하려면 해당 파일을 올바른 전문 소프트웨어에서 열어보고, 소프트웨어를 정확하게 조작하며, 비용을 지불하는 실제 고객의 관점에서 판단을 내려야 합니다. 바로 그런 실습적인 소프트웨어 사용이 현재 AI 에이전트가 가장 서툰 분야입니다. AI 심사관 역시 평가 대상이 되는 AI 작업자와 동일한 한계에 부딪히는 것입니다. GPT-5.5의 가짜 렌더링 사태가 좋은 예입니다. 이러한 속임수를 적발하려면 3D 모델을 직접 열어 실제 기하학적 구조를 검사해야만 합니다.

모델들이 자신의 역량을 최대한 발휘할 수 있도록, 연구팀은 Claude Code나 Codex CLI 등 개발자들이 일상적으로 사용하는 것과 동일한 도구 환경에서 모델들을 구동시켰습니다. 여기에는 그래픽 프로그램을 직접 조작할 수 있는 기능이 추가로 확장되었습니다. 이 작업 환경은 Blender, GIMP, Audacity를 포함한 30개 이상의 전문 앱이 설치된 가상 리눅스 머신입니다. 각 프로젝트에는 최대 24시간의 컴퓨팅 시간이 주어집니다. 이 설정은 또한 비판 루프(critic loop)를 사용합니다. 즉, 두 번째 AI 에이전트가 까다로운 고객처럼 결과물을 비판적으로 검토하고, 첫 번째 에이전트가 자신의 작업을 수정하는 방식입니다.

AI는 여전히 대부분의 프로젝트에서 전문가 수준의 품질을 달성하지 못하고 있습니다. 블로그 포스트에 표시된 Fable 5의 세 가지 결과물 모두 완벽하지 않습니다.

원문 보기
원문 보기 (영어)
AI agents can now complete 16 percent of freelance jobs at pro quality, up from 2.5 percent eight months ago Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jul 2, 2026 Nano Banana Pro prompted by THE DECODER The Remote Labor Index measures how often AI agents complete paid freelance projects at professional quality. In eight months, the top automation rate has more than quadrupled. The Remote Labor Index (RLI) tracks how often AI agents can finish real, commercially valuable freelance jobs at a quality level a paying client would actually accept. The benchmark covers areas like 3D and CAD, architecture, graphic design, video and animation, audio, data analysis, and web apps. It includes 240 projects worth a combined $144,000, sourced from 358 verified freelancers. Human evaluators at the Center for AI Safety score each result against a gold standard created by a paid professional. The RLI was developed together with Scale Labs. The key metric is the automation rate, meaning the share of projects where the AI's work is rated at least as good as a human's. Top automation rate jumps from 2.5 to 16.1 percent When the benchmark first launched, the best AI agent automated just 2.5 percent of projects. According to the latest results, Fable 5 now hits 16.1 percent, the highest score ever recorded. That's roughly double Opus 4.8's 8.3 percent. GPT-5.5 comes in at 6.3 percent. All three models beat every previously tested system. The prior leader, Opus 4.6 running on the Claude Cowork framework, sat at 4.17 percent. The frontier has more than quadrupled in under eight months, according to the authors. One caveat about Fable 5's score: only 218 of 240 projects could be evaluated before the U.S. government restricted access to the model. Even in the worst case, where Fable 5 failed every missing project, its rate would still be 14.6 percent, higher than any other model. Progress doesn't track neatly with release dates, though. On the full Scale Labs leaderboard , the newer Gemini 3 Pro lands near the bottom at just 1.25 percent, behind much older systems. Some examples from the study also show where even top models still fall short. On a ring design task, Fable 5 is clearly better than earlier AIs but still looks unprofessional on closer inspection. On an architecture project, GPT-5.5 faked an appealing render using an image generator while its actual 3D model remained flawed. Human evaluators still can't be replaced The team tested whether expensive human evaluation could be replaced by AI judges. The answer was clear: AI judges rated the new models far too generously. For GPT-5.5, the AI evaluator's score was almost three times too high. For Opus 4.8, about two and a half times. The automated judge did get the ranking order right, but the actual numbers were way off. The reason, according to CAIS: To fairly judge delivered work, you need to open the files in the right professional software, operate that software correctly, and form a judgment like a paying client would. That kind of hands-on software use is exactly what current AI agents are worst at. An AI judge runs into the same limits as the AI workers it's supposed to evaluate. GPT-5.5's faked rendering is a good example: catching the trick requires opening the 3D model and inspecting the actual geometry. To let the models show their full ability, the team runs them in the same tools developers use day to day, like Claude Code and Codex CLI. These were extended with the ability to operate graphical programs directly. The work environment is a virtual Linux machine loaded with over 30 professional apps, including Blender, GIMP, and Audacity. Each project gets up to 24 hours of compute time. The setup also uses a critic loop: a second AI agent reviews the output as critically as a demanding client, and the first agent then revises its work. AI still fails to hit professional quality on most projects. None of the three Fable 5 results shown in the blog post would pass as finished work. But the rise in automation rates within a single year is rapid, the authors say, and directly reflects how fast remote work automation is advancing. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->