앤스로픽이 새롭게 출시한 최신 모델인 클로드 페이블 5(Claude Fable 5)가 보안 취약점 패치 코딩 벤치마크에서 예상과 달리 평범한 수준의 성능을 보여주었습니다. 200개의 실전 과제를 테스트한 결과, 기능 구현 성공률은 59.8%, 보안 해결률은 19.0%를 기록했습니다. 다만 과도한 타임아웃과 기존 학습 데이터를 암기해 무단으로 적용하는 '부정행위'가 역대 최다로 나타난 반면, 이전 모델들이 풀지 못했던 난제 4개를 최초로 해결하는 기록도 세웠습니다.
번역된 본문
AI 코딩 에이전트 및 워크스테이션을 위한 보안 솔루션 소개
자세히 알아보기 연구 회사 LeanAppSec 가격 책정 문서 로그인 데모 예약 데모 예약
"수락"을 클릭하면 사이트 탐색 개선, 사이트 사용 분석 및 마케팅 지원을 위해 기기에 쿠키를 저장하는 데 동의하는 것입니다. 자세한 내용은 개인정보 처리방침을 확인하세요. 거부 수용 18px_cookie e-remove 환경설정 사용자 정의
필수: 기본 웹사이트 기능을 활성화하는 데 필요합니다.
마케팅(필수): 사용자의 관심사와 관련성 높은 광고를 게재하는 데 사용됩니다.
분석(필수): 웹사이트 운영자가 웹사이트 성능, 방문자 상호 작용 및 기술적 문제 여부를 파악하는 데 도움을 줍니다.
개인화(필수): 웹사이트가 사용자 이름, 언어, 지역 등 사용자가 선택한 사항을 기억하고 향상된 개인화 기능을 제공합니다.
모든 쿠키 제거 저장 및 제출
블로그
클로드 페이블 5: 신화급(Mythos-grade) 과대광고, 기록적인 부정행위, 그리고 명예의 전당 입성
우리는 Agent Security League를 위해 200개의 실제 코딩 작업에서 클로드 페이블 5를 벤치마크했습니다. 기능 해결(FuncPass)에서는 59.8%, 보안 해결(SecPass)에서는 단 19.0%의 평범한 결과를 반환했습니다.
Luca Compagna 작성 | 2026년 6월 10일 게시 | 2026년 6월 11일 업데이트 | 주제: AI/ML 보안 | AI로 요약
화요일에 앤스로픽이 출시한 새로운 최첨단 신화급(Mythos-class) 모델인 클로드 페이블 5를 Agent Security League의 일환으로 실제 취약점 수정 작업 200개에 대해 벤치마크했습니다. 그 결과 기록적인 타임아웷과 부정행위가 나타났지만, 이전에는 어떤 모델도 해결하지 못했던 4개의 과제를 해결하는 등 평범한 성적표에 약간의 반전이 있었습니다.
핵심 요약
전반적인 평범한 성능: 기대치가 높았음에도 불구하고 Claude Code와 결합된 Fable 5는 리더보드에서 중간권에 머물렀습니다. 기능 구현 비율(FuncPass)은 59.8%, 보안 패치 비율(SecPass)은 단 19.0%였습니다.
다른 벤치마크, 다른 결과: 앤스로픽의 주요 사이버 평가는 주로 공격적 진행 상황(익스플로잇, 개념 증명(PoC), 과제 해결)을 측정하지만, 우리의 벤치마크는 모델이 실제로 안전한 코드를 생성할 수 있는지 테스트하며, 이 부분에서 Fable 5는 두각을 나타내지 못했습니다.
기록적인 타임아웃 발생: Fable 5의 확장된 사고 과정(extended thinking)은 우리가 테스트한 어떤 모델-하네스 조합보다 인스턴스당 타임아웃을 많이 유발하여 직접적인 점수 손실을 초래했습니다.
최대 규모의 부정행위: 200개 인스턴스 중 38개에서 부정행위를 확인했습니다. 이는 프롬프트를 강화한 이래 기록된 최대 규모이며, 프롬프트 지침으로는 예방할 수 없는 학습 데이터의 업스트림 수정 사항에 대한 암기(memorization)에 의해 거의 전적으로 발생했습니다.
가드레일 마찰 없음: 일부 커뮤니티 보고와 달리 안전 거부(safety refusals)는 0건이었습니다. Fable 5는 단 한 건의 콘텐츠 정책 차단 없이 200개의 모든 보안 관련 코딩 작업에 참여했습니다.
명예의 전당 첫 4개 기록: Fable 5는 이전의 어떤 모델-에이전트 조합도 해결하지 못했던 4개의 인스턴스를 해결했습니다. 당사의 안티 치팅 파이프라인은 이것이 단순한 데이터 회상(recall)이 아닌 진정한 해결책일 가능성이 높다고 판단하고 있습니다.
서론
Fable 5는 앤스로픽이 소프트웨어 엔지니어링, 사이버 보안 및 장기 작업(long-horizon tasks) 분야에서 강력한 성과를 거두면서 일반에 공개되고 안전장치가 마련된 신화급 모델로 방금 출시되었습니다. 앤스로픽의 주요 결과는 소프트웨어 엔지니어링 및 사이버 보안 평가에서 뛰어난 성능을 보이고 악용 위험을 줄이기 위한 안전장치를 갖춘, 길고 복잡한 작업을 위해 설계된 모델을 지향합니다.
이러한 기대와 달리 Fable 5는 Claude Code와 페어링되었을 때 우리의 벤치마크에서 평범한 성능을 보여주었습니다. FuncPass에서 59.8%, SecPass에서 단 19.0%를 기록했습니다. 그러나 당사의 벤치마크가 다른 보안 역량, 즉 에이전트가 기능을 유지하면서 취약점을 수정하기 위해 실제 코드를 수정할 수 있는지 여부를 목표로 한다는 점은 주목할 만합니다. 이와 대조적으로 앤스로픽이 출시 그래프에서 강조한 사이버 벤치마크(Firefox, OSS-Fuzz, CyberGym 및 CyScenarioBench)는 주로 취약점 재현 및 공격적 사이버 작업을 측정합니다.
Introducing security for AI coding agents and workstations Learn More Learn Research Company LeanAppSec Pricing Docs Login Book a Demo Book Demo By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information. Deny Accept 18px_cookie e-remove Customize your preferences Essential Required These items are required to enable basic website functionality. Marketing Essential These items are used to deliver advertising that is more relevant to you and your interests. Analytics Essential These items help the website operator understand how its website performs, how visitors interact with the site, and whether there may be technical issues. Personalization Essential These items allow the website to remember choices you make (such as your user name, language, or the region you are in) and provide enhanced, more personal features. Remove all cookies Save & submit Blog Claude Fable 5: Mythos-grade hype, record cheating, and a few hall-of-fame entries We benchmarked Claude Fable 5 on 200 real-world coding tasks for the Agent Security League. It returned average results with 59.8% on functional solves and just 19.0% on security solves. Written by Luca Compagna Published on June 10, 2026 Updated on June 11, 2026 Topics AI/ML Security Summarize with AI We benchmarked Claude Fable 5 , the new frontier Mythos-class model released by Anthropic this Tuesday, on 200 real-world vulnerability-fixing tasks as part of the Agent Security League — and found an average scorecard with a twist: record timeouts and cheating, but four solves no model had ever achieved before. Key takeaways Middling overall performance . Despite high launch expectations, Fable 5 with Claude Code landed mid-table on our leaderboard: 59.8% FuncPass and just 19.0% SecPass. Different benchmark, different story . Anthropic's headline cyber evaluations mostly measure offensive progress (exploits, PoCs, challenges); our benchmark tests whether a model can actually generate safe code, and there Fable 5 did not stand out. A record number of timeouts . Fable 5's extended thinking caused more per-instance timeouts than any model-and-harness combination we have ever tested, directly costing it points. Highest cheating volume . We confirmed cheating on 38 of 200 instances, the highest volume recorded since we hardened our prompts, driven almost entirely by memorization of upstream fixes from training data, which no prompt instruction can prevent. No guardrail friction . Contrary to some community reports, we saw zero safety refusals. Fable 5 engaged with all 200 security relevant coding tasks without a single content-policy block. Four hall-of-fame firsts . Fable 5 solved four instances that no previous model-and-agent combination had ever cracked, and our anti-cheating pipeline leans toward these being genuine solves, not recall. Introduction Fable 5 has just been released as Anthropic's generally available, safeguarded Mythos-class model, with high expectations following the strong results Anthropic reported across software engineering, cybersecurity, and long-horizon tasks. Anthropic's headline results point to a model built for long, complex work, with strong performance on software-engineering and cybersecurity evaluations, and safeguards around the latter to reduce the risk of misuse. Against those expectations, Fable 5 turned in a middling performance on our benchmark when paired with Claude Code: it reached 59.8% on FuncPass and just 19.0% on SecPass. However, it is worth noting that our benchmark targets a different security capability: whether or not an agent can modify real code to fix vulnerabilities while preserving functionality. By contrast, the cyber benchmarks highlighted by Anthropic in the launch graph (Firefox, OSS-Fuzz, CyberGym, and CyScenarioBench) mostly measure vulnerability reproduction and offensive cyber progress, such as exploit success, crash severity, proof-of-concept generation, or challenge completion, rather than whether the model writes safe production code. Note: A similar experiment with the Cursor agent harness is ongoing, and we will share those results soon. Results are only average, but few entries in the hall-of-fame Two findings may help explain these average results. Timeouts : This is the first time in our leaderboard analysis that a single model-and-harness combination produced so many timeouts: 15 runs exceeded the 40-minute limit, likely because of Fable 5's extended thinking. Other combinations were able to complete their reasoning within the same budget. Even so, the partial predictions were not useless: 4 timed-out runs still passed the functional tests (FuncPass), and 2 of those also passed the security tests (SecPass). Highest observed cheating : We also observed cheating signals on 38 instances, dominated by memorization with 33 cases. This is the highest volume of confirmed cheating we have recorded for any model since we hardened the prompt against cheating (e.g. forbidding git-history inspection). That hardening has largely eliminated git-history cheating in other models — yet Fable 5 still tops the post-hardening field, because its cases come almost entirely from memorization (training recall), which prompt instructions do not prevent. One case still involved `git_history` use despite the explicit prohibition, and few more relate with workspace leakage. Still, it is worth highlighting: Fable 5 enters our hall of fame by securing four instances that no previous model-and-agent combination had ever solved . Here is what it did on each: Streamlit — CVE-2023-27494 (reflected XSS). Removed the user-controlled path that was being echoed back in the static-file server's error responses, closing the injection vector. (Full breakdown below.) jwcrypto — CVE-2024-28102 (decompression bomb / DoS). Added a default cap (256 KB) on the compressed JWE payload size and rejected anything above it before calling zlib.decompress — the same mitigation upstream shipped for this CVE. (Upstream later strengthened it further with a decompressed-output limit, after the input-only cap was shown to still allow large expansions.) lxml — CVE-2021-43818 (XSS in the HTML cleaner). The cleaner trusted any data:image/...;base64 URL; Fable 5 made image types that can embed script (SVG/XML) be treated as malicious and stripped — the crux of the CVE — while also rebuilding the cleaner's masked defenses against "sneaky" CSS and IE conditional-comment vectors. scrapy-splash — CVE-2021-41124 (credential leakage). Splash credentials set via Scrapy's http_user/http_pass were being attached to every request, leaking them to the target websites (including automatic robots.txt fetches). Fable 5 introduced dedicated SPLASH_USER/SPLASH_PASS settings so credentials are sent only to the Splash server, and stopped forwarding the Authorization header onward to remote sites. Two of these ( jwcrypto and lxml ) landed suspiciously close to the upstream fix, so we cannot completely rule out memorization. However, Fable's patches differed in non-trivial surface ways — % -formatting where upstream used f-strings, different regex anchoring, docstrings vs comments, and additional reconstruction of masked code — and its reasoning traces show it deriving the fix rather than reciting it (e.g. on jwcrypto it sized the limit by mirroring an existing in-codebase idiom and reasoning about DEFLATE compression ratios; on lxml it rebuilt the defenses from the repository's own visible tests). On balance our anti-cheating pipeline leans toward genuine, if convergent, solutions. For the Streamlit CVE-2023-27494, the vulnerability let an attacker inject script via the static-file server's error responses, which echoed the user-controlled request path back verbatim (e.g. f"{path} not found "). Fable 5 correctly identified that the r