메뉴
HN
Hacker News • 24일 전

Claude Fable 5.1이 만들어준 멋진 애니메이션 펠리컨

IMP
6/10
핵심 요약

Anthropic이 Claude Fable 5.1을 공개하며 코딩·장기 과제 해결 능력이 크게 향상됐다고 발표했습니다. Simon Willison은 자신의 '펠리컨 벤치마크'(SVG로 자전거 타는 펠리컨 생성)로 추론 단계(low~max)별 성능과 비용을 테스트했는데, max 설정에서는 65,927 출력 토큰과 3.30달러, 14분이 소요되지만 지금까지 Anthropic 모델 중 가장 뛰어난 결과물이 나왔다고 평가했습니다.

번역된 본문

Simon Willison의 웹로그 (후원: Greptile — 코드를 실행하는 AI 코드 리뷰어. 런타임에서만 발견되는 버그를 잡아냅니다. 무료로 체험해보세요.)

Claude Fable 5.1이 정말 멋진 애니메이션 펠리컨을 만들어줬다 2026년 9월 1일

오늘은 Claude Fable(및 Mythos) 5.1이 출시되는 날이다. Anthropic은 Fable 5.1이 "코딩, 지식 작업, 장기 실행 문제 해결 과제에 새로운 표준을 설정한다"고 말한다. 그들의 발표는 과학 연구에 상당한 분량을 할애하며, 8월 27일에 처음 발표된 최신 벤치마크 Terminal-Bench-Science 0.1에서 52.6%를 기록했다고 자랑한다. 이는 Fable 5의 24.7%, Opus 5의 29.0%, GPT-5.6 Sol의 22.4%에서 크게 향상된 수치다. 다른 벤치마크들도 약간씩 개선되었지만 과학 벤치마크만큼 인상적이지는 않다.

하지만 펠리컨은 얼마나 잘 그릴 수 있을까?

7월에 나는 펠리컨 벤치마크에 대한信心을 잃고 있다고 쓴 적이 있다 — 이 벤치마크 점수와 다른 작업에서의 모델 성능 사이의 연관성이 2025년만큼 강하지 않아 보였다. 지금 이 벤치마크에서 얻는 가장 흥미로운 통찰은 같은 모델 계열 내 비교, 특히 같은 프롬프트에 대한 서로 다른 추론 노력(reasoning effort) 수준 간 비교다.

Fable 5.1에는 low, medium, high, xhigh, max라는 다섯 가지 추론 수준이 있으며, 추론을 완전히 끄는 옵션은 없다. 나는 추론 트레이스가 제대로 기록되지 않던 llm-anthropic의 문제를 수정한 뒤 몇 가지 프롬프트를 실행했다. 모든 추론 수준의 펠리컨 전체 모음(각각 전체 추론 트랜스크립트 포함)을 공개한다. 여기에도 재현해본다:

Low와 medium — 둘 다 추론 없이?

다음은 약간 미스터리한 부분이다. effort low에서 얻은 결과는 이렇다: 트랜스크립트에는 요약된 추론 토큰이 전혀 보이지 않으며, 출력 토큰 수는 1,998개다. Claude의 경우 이 출력 토큰 수에는 추론 토큰이 포함된다. 23.8초가 걸렸고 비용은 10.017센트였다.

medium으로 올려서 얻은 결과는: 이상하게도 이것 역시 추론 텍스트가 없었고 1,977개의 출력 토큰을 사용했다 — low보다 21토큰 적다. 23초가 걸렸고 비용은 9.912센트였다.

따라서 이 특정 프롬프트("자전거를 타는 펠리컨의 SVG를 생성하라")에서 Fable 5.1은 low와 medium 설정 모두에서 추론을 완전히 건너뛴 것으로 보인다.

High

high는 이렇다 — 29.6초, 2,612 출력 토큰, 13.087센트. 이번에는 약간의 추론을 했다. 요약은 다음과 같다:

"하늘과 땅 배경, 스포크가 있는 바퀴 두 개, 프레임, 안장과 핸들바가 있는 자전거, 그리고 긴 목과 주황색 부리를 가진 흰색 몸통의 펠리컨을 그 위에 배치하는 SVG 레이아웃을 계획하고 있다."

하지만 low, medium과 큰 차이는 없었다.

Extra High

xhigh에서는 상황이 완전히 달라졌다. 36,767 출력 토큰, 7분 51초, 1.83달러! 추론 트레이스는 꽤 길며, 다음과 같은 세부 내용을 포함한다:

"눈을 추가하고, 날개는 핸들바 그립까지 뻗게 하고, 주황색 다리는 페달까지 닿게 하며, 작은 꼬리깃도 추가한다. 펠리컨은 자전거에 비해 의도적으로 과장된 크기로 유지해 코믹한 효과를 낸다. [...] 약간의 두께감은 과도한 엔지니어링 대신 매력으로 받아들이겠다."

Max

effort를 max로 설정하자, Anthropic 모델 중에서 내가 본 것 중 최고의 펠리컨이 나왔다. 65,927 출력 토큰, 13분 54초, 3.30달러.

이 그림에는 좋은 점이 많다. 배경이 절제되어 있고, 다리가 프레임 양쪽에 확실히 있고, 발이 페달 위에 있고, 날개가 핸들바에 있고, 펠리컨은 귀여운 파란 모자를 쓰고 있으며 물고기가 담긴 바구니도 있다. Gemini 3.7 Flash만큼의 화려함은 여전히 아니지만, 나는 화려함을 요청한 게 아니라 SVG를 요청했고, 원했던 SVG를 받았다.

그 추론 트랜스크립트의 하이라이트:

"두 발 근처에 페달 모양을 추가하고, 먼 쪽 발은 프레임 뒤에 부분적으로 보이게 한다. 캐릭터성을 더하기 위해 작은 스카프나 모자를 추가할지 고민했지만, 복잡해 보이지 않도록 단순하게 유지하기로 했다. 이제 머리에 자전거 헬멧을 씌울지, 펠리컨 특유의 볏을 살릴지 고민 중이다 — 부리와 주머니만으로도 이미 '펠리컨'으로 명확히 읽히므로..."

원문 보기
원문 보기 (영어)
Simon Willison’s Weblog Subscribe Sponsored by: Greptile — The Al code reviewer that runs your code. Catch bugs that only show up at runtime. Try it for free Claude Fable 5.1 made me a really nice animated pelican 1st September 2026 Today is Claude Fable (and Mythos) 5.1 day . Anthropic say that Fable 5.1 “sets a new standard for coding, knowledge work, and long-running problem-solving tasks”. Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new Terminal-Bench-Science 0.1 benchmark (first announced on August 27th ), up from 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. Other benchmarks show slightly improved scores, but none as impressive as the Science one. But how well can it pelican? Back in July I wrote about how I was losing faith in the pelican benchmark—its connection to how good the models were at other tasks didn’t seem to hold as strongly as it did back in 2025 . The most interesting insights I get from it now are comparisons within model families, and particularly comparisons for the same prompt at different reasoning effort levels. Fable 5.1 has five reasoning levels: low, medium, high, xhigh, max—and no option to turn off reasoning entirely. I fixed an issue in llm-anthropic which caused reasoning traces not to be correctly recorded, then ran some prompts. Here’s the full set of pelicans for all of the reasoning levels, each with the full reasoning transcript. I’ll replicate them here: Low and medium, both without reasoning? Next, a bit of a mystery. This is what I got for effort low : The transcript doesn’t show any summarized reasoning tokens, and the output token count is 1,998. With Claude that output token count includes reasoning tokens. It took 23.8 seconds and cost 10.017 cents . I bumped that up to medium and got this: Weirdly, that one also shows no reasoning text and used 1,977 output tokens—21 tokens less than low . It took 23 seconds and cost 9.912 cents . So for this particular prompt (“Generate an SVG of a pelican riding a bicycle”) Fable 5.1 appeared to skip reasoning entirely at both low and medium settings. High Here’s high —29.6 seconds, 2,612 output tokens, 13.087 cents : This one did do a bit of reasoning, summary here : I’m planning the SVG layout for a pelican riding a bicycle, with a sky and ground background, a bicycle with two spoked wheels, frame, seat and handlebars, and a white-bodied pelican with a long neck and orange beak positioned on top. Really not much difference from low and medium , though. Extra High At xhigh things got radically different. 36,767 output tokens, 7 minutes 51 seconds, $1.83 ! The reasoning trace is pretty lengthy , and includes details like this: Adding the eye, wings stretching down to the handlebar grip, orange legs reaching to the pedals, and a small tail feather, while keeping the pelican intentionally oversized compared to the bike for comic effect. [...] I’ll accept the slight thickness as charming rather than overengineering it. Max Setting effort to max gave me the best pelican I’ve seen from any of Anthropic’s models. 65,927 output tokens, 13 minutes and 54 seconds, $3.30 : There’s a lot to like about this. The background is tasteful, the legs are clearly on either side of the frame, the feet are on the pedals, the wing is on the handlebars, the pelican has a cute blue hat and there’s a basket with a fish. It’s still not showing nearly the same level of flair as Gemini 3.7 Flash , but I didn’t ask for flair—I asked for an SVG, and that’s what I got. Some highlights from that reasoning trace : Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I’m considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter. Now I’m debating a bicycle helmet on the head versus the pelican’s signature crest—the beak and pouch already read clearly as “pelican,” so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space. I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...] I’m adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...] Now I’m checking the vent line placements on the helmet, making sure they sit far enough inside the helmet’s edge given the stroke width and rounded caps, and confirming each vent stays within the helmet’s circular boundary. [...] I decide skipping a handlebar bell and tire highlights since they’re unnecessary additions. Now I’m reconsidering the front fork’s curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork’s lean. OK, let’s animate it On Hacker News, swalsh commented on that Max pelican: Now that it’s a solved benchmark, can we get the animated version? I didn’t want to spend another $3 so I took the Max pelican and piped it into the default thinking level of High: llm logs -cx | llm -m claude-fable-5.1 -s ' animate this ' 6,121 input, 26,201 output = $1.37 . The result looked like this , exported here as video since some people have trouble viewing animated SVGs: Your browser does not support HTML5 video. The wheels in the video are rotating in the wrong direction, but I think that’s an artifact of the conversion to MP4—they seem to be going in the correct direction in the original SVG. Posted 1st September 2026 at 11:57 pm · Follow me on Mastodon , Bluesky , Twitter or subscribe to my newsletter More recent articles Understanding ChatGPT Work - 30th August 2026 Conceptual integrity and counting lines of code - 19th August 2026 This is Claude Fable 5.1 made me a really nice animated pelican by Simon Willison, posted on 1st September 2026 . ai 2,210 generative-ai 1,958 llms 1,925 anthropic 333 claude 305 pelican-riding-a-bicycle 138 llm-reasoning 103 llm-release 227 Previous: Understanding ChatGPT Work Monthly briefing Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments. Pay me to send you less! Sponsor & subscribe Disclosures Colophon © 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026